The Missing Layer in Healthcare Data Exchange
Healthcare learned how to move data between systems. It skipped the part where the data was never structured in the first place.
Every time someone sees a doctor, the visit turns into a record: a note, a referral, a discharge summary, often faxed or scanned. To a person, that record is easy to read. To a computer, it is just a picture of text, with no idea what any of it actually means.
That difference has a name. Data written for humans, with free-flowing sentences, abbreviations, and the occasional note scrawled in a margin, is called unstructured data. Data written for machines, with specific facts in labelled fields, like "diagnosis = type 2 diabetes" or "systolic blood pressure = 140", is called structured data. Software can sort, count, check, and act on structured data. It can do none of that with a paragraph of prose.
Here is the catch. Almost every part of healthcare that does something useful with your record, whether it is billing, quality reporting, risk scoring, or prior authorization, assumes the data already arrived in that clean, structured form. Most of the time, it did not.
About 80% of clinical data is unstructured. It lives in faxes, scanned PDFs, and handwritten notes. The work of turning it into data that software can actually use is the layer almost no one has built.
Someone has to translate the document
So who turns the messy document into clean data today? People do. The most important version of this work is called medical coding. A trained coder reads the clinical note and assigns standardized codes that summarize what happened during the visit: which diagnoses were made, which procedures were done, which medications were given.
Those codes come from official code sets. The best known is ICD-10, a catalogue of tens of thousands of diagnosis codes in which every condition has its own entry. Codes exist because the sentence "the patient has diabetes" means nothing to a billing system or a public-health database, but a specific code does. Codes are how healthcare gets paid, measured, and studied. They are the structured data that the document was supposed to become.
Why reading the words is not enough
The hard part is that the code has to match exactly what the document supports, not just what it loosely says. Most of the difficulty comes from three things ICD-10 forces you to pin down: specificity (how precisely the condition is described), severity (how advanced or complicated it is), and laterality (which side of the body it affects, left or right). The same diagnosis can map to dozens of codes depending on those details. To see how fine the line is, look at one diabetes diagnosis.
E11.9 means "type 2 diabetes without complications." E11.65 means "type 2 diabetes with hyperglycemia." They look almost identical, but E11.65 is only correct if the note actually documents the high blood sugar. Choose E11.65 when the hyperglycemia was never written down, and the code is unsupported. In an audit, it gets reversed.
That one decision turned on severity alone. Now multiply it by the tens of thousands of codes in ICD-10, layer on the choices about specificity and laterality, and add documentation rules that change from one insurer or state program to the next. That is why this has always been careful, expert work.
The demo is easy. Getting it right is not.
This is where modern AI is supposed to help, and in a quick demo it looks like it already has. Point a model at a clean PDF and it will pull the text out in seconds. But two things break in the real world:
- The documents are messy. Real records are faxes, photocopies of photocopies, handwriting, and layouts the model has never seen, where accuracy drops sharply.
- Reading is not the same as being right. Pulling the words off the page tells you nothing about whether the resulting code is clinically correct. That is the
E11.65problem all over again.
Why this matters now
For years this was a known problem with no deadline. That changed. Health data is moving to a standard digital format called FHIR, which you can think of as a common language that lets different health systems exchange data without translating it by hand. The federal Interoperability and Prior Authorization rule requires health plans to exchange clinical data through FHIR interfaces by 2027, including the clinical-note data defined in the national USCDI standard. National networks like TEFCA are already exchanging records across hundreds of organizations, and all of it assumes the data is already structured.
The cost of getting codes wrong rose at the same time. In one HHS Office of Inspector General audit of Medicare Advantage, the records for more than 61% of sampled cases did not support the diagnosis codes that had been submitted, and unsupported codes can mean repayments and penalties. As this work shifts to AI, that same liability comes with it, unless every code can be traced back to the exact words in the source document.
What NHance does
Curiflow builds that missing layer. NHance takes any clinical document, whether it is a PDF, a scan, a fax, or an HL7 or CCDA file, and turns it into structured, coded, FHIR-ready data. Every value it produces is linked back to the exact place in the source document it came from, so a person can check it and an auditor can trust it.
The same simple idea sits under every workflow: take a document, apply the right rule, and produce a decision. Get that one conversion right, and risk adjustment, prior authorization, eligibility, and quality reporting all run on data they can finally rely on.
Convert anything to FHIR.
NHance turns eFaxes, PDFs, CCDA, and CSVs into clean, structured, FHIR-ready data in under 3 minutes, not 3 months. HIPAA and SOC 2 compliant.