1. OCR and table extraction errors
Before any model can interpret a result, it has to read it. Lab reports are dense tables — marker names in one column, values, units and reference ranges in others, sometimes two sets of columns per page. Optical character recognition (OCR) and the vision features of multimodal models can misalign those columns, so a value lands on the wrong marker's row.
Small visual details cause outsized errors. A faint decimal point turns 1.2 into 12. The less-than sign in “<0.5” disappears. Phone photos add skew, glare, shadows and cropped edges, and a multi-page PDF can lose the page holding half the panel. Not every PDF is equal, either: a scanned, image-only PDF has to go through OCR just like a photo, while a PDF exported from a patient portal usually carries real text. The model then interprets whatever it extracted as if it were correct.
Mitigation: a dedicated analyzer is designed around lab-report layouts, and a structured per-biomarker output makes misreads far easier to spot than a paragraph of chat. With any tool, a text-based PDF beats a photo.
2. Unit confusion
The same measurement can be reported in different units depending on the country and laboratory. US labs mostly use conventional units; most other countries use SI units. Each analyte has its own conversion factor:
A chatbot that reads “5.5” without its unit, or assumes US units for a UK report, can describe a normal glucose as dangerously low. Extraction makes it worse: “µ” is easily misread as “u” or “m,” and mmol/L versus µmol/L is a thousand-fold difference. Mitigation: a purpose-built analyzer reads the unit attached to each value and interprets it within the matching system instead of guessing.
- Glucose: mmol/L × 18 ≈ mg/dL (5.5 mmol/L ≈ 99 mg/dL)
- Total cholesterol: mmol/L × 38.7 ≈ mg/dL (5.0 mmol/L ≈ 193 mg/dL)
- Triglycerides: mmol/L × 88.6 ≈ mg/dL
- Creatinine: mg/dL × 88.4 = µmol/L (1.0 mg/dL = 88.4 µmol/L)
- Hemoglobin: g/dL × 10 = g/L (13.5 g/dL = 135 g/L)
- Vitamin D (25-hydroxy): ng/mL × 2.5 ≈ nmol/L
3. Lab-specific reference ranges
Reference ranges are not universal constants. Each laboratory establishes or verifies its own ranges based on its instruments, methods and population, and ranges can differ by age, sex and pregnancy. That is why MedlinePlus advises comparing results with the range on your own report.
General chatbots have no built-in table of laboratory reference ranges. When the range on your report is missing, cropped or ignored, the model falls back on a “typical” range from its training data — which may come from a different country, method or decade. A result can sit inside your lab's range and outside the chatbot's, or the reverse. Some markers add another layer: LDL cholesterol, for instance, is often judged against targets that depend on a person's overall cardiovascular risk rather than a single population range — a judgment only a clinician can make. Mitigation: dedicated analyzers explain each biomarker against clinical reference ranges as a core function, and keeping your lab's printed ranges in the upload gives any tool the right baseline.
4. Flags that get ignored
Labs print flags — H for high, L for low, sometimes HH, LL or an asterisk for critical values — next to results outside the range. A flag represents the lab's own judgment using its own range and method, which makes it one of the most reliable signals on the page. Yet flags are tiny, often sit in a separate column and are easy for extraction to drop. A chatbot may then reason from the number alone and reach a different conclusion from the lab's. It also cannot know what happened outside the document — for example, whether the lab already contacted your clinician about a critical value, as many labs routinely do.
Mitigation: structured reports place each result's status relative to its range in a consistent field. Whatever tool you use, compare its flags with your lab's flags one by one.
5. Missing clinical context
A blood value means little without the person it came from. Fasting status changes glucose and triglycerides. Many medications and supplements affect liver enzymes, electrolytes, thyroid tests or blood counts, and high-dose biotin supplements can interfere with some lab assays. Pregnancy changes many ranges. Age and sex shift expectations for hemoglobin, creatinine and hormones, and some hormones, such as cortisol and testosterone, vary with the time of day the sample was drawn. Hard exercise, dehydration or a recent illness can move markers temporarily.
A prompt that says “explain my results” contains none of this, so the model interprets numbers in a vacuum — or quietly assumes a context. Mitigation: a well-designed analyzer accounts for basic context and states what it does not know. Your clinician has the full picture; no AI tool does.
6. Confident hallucination
Large language models generate the most plausible next words, not verified facts. When information is missing or ambiguous, they tend to fill the gap with something that sounds right: a reference range quoted to one decimal place, a marker that was never tested, or a neat causal story linking two unrelated results. Because the tone is equally confident whether content is grounded or invented, hallucinations are hard to spot without the source document. They are most likely exactly where extraction was weakest — a cropped column, an unclear unit — so one error tends to compound another.
The risk cuts both ways. An invented explanation can cause needless alarm, or false reassurance that delays a conversation with a clinician. Mitigation: keep the tool anchored to your extracted data, and prefer outputs that show values, units and ranges side by side so every claim can be traced.
7. Non-repeatability
Ask a general chatbot about the same report twice and you may get two differently organized answers that emphasize different results. Sampling randomness, model updates and differences in conversation history all contribute; even the wording of your question shifts the answer. For a one-off question this is harmless. For tracking a panel over months it is a real problem, because you cannot tell whether a change in the explanation reflects your results or the model. It also undermines trust: if one session calls a value “slightly elevated” and the next calls it “nothing to worry about,” which do you believe?
Mitigation: an analyzer that returns a structured, repeatable per-biomarker report makes your numbers the only thing that changes between analyses.
What a well-designed dedicated analyzer does differently
A dedicated AI blood test analyzer is not immune to these problems, but it is designed around them. In our 2026 ranking, Kantesti is #1 at 9.4 / 10 because it is a health-trained model built specifically for blood test interpretation: it explains each biomarker against clinical reference ranges and returns a structured, repeatable report across full panels in 75+ languages. General-purpose models — ChatGPT (7.1), Gemini (6.9), Claude (6.8) and Perplexity (6.7) — remain strong at explanation but lack built-in reference ranges and structured output; see the full rankings.
Whichever tool you evaluate, this is what good design looks like for each failure mode:
- OCR errors → handling built for lab-report layouts, with extracted values visible for checking
- Unit confusion → units read and interpreted alongside each value
- Generic ranges → every marker explained against clinical reference ranges
- Ignored flags → a consistent status field for every result
- Missing context → explicit limits and a prompt to involve a clinician
- Hallucination → output tied to extracted values rather than open-ended chat
- Non-repeatability → the same structured format every time
How to catch these errors yourself
Whatever tool you use, a few habits catch most of these problems before they mislead you:
For the full process, see how to prepare a lab report for AI analysis and our guide to whether ChatGPT can read blood test results.
- Upload a text-based PDF rather than a photo whenever possible.
- Keep units, reference ranges and flags visible — redact identity, not data.
- Read the extracted values first and check them line by line.
- Compare every H/L flag in the output with your lab's flags.
- Run the same report twice and treat differences as a warning sign.
- Take the final output to a licensed clinician as questions, not conclusions.
If a result worries you
AI tools cannot judge urgency. If your lab or clinic flagged a critical value, follow their instructions. For chest pain, severe shortness of breath, fainting or sudden confusion, call emergency services.
Frequently asked questions
Why does ChatGPT get my lab values wrong?
Most errors start with extraction: reading a table from a photo or PDF can misalign columns, drop decimal points or lose units. The model then interprets the corrupted data confidently. Checking the extracted values against your report catches most of these errors.
Can AI confuse mg/dL and mmol/L?
Yes. If the unit is missing, misread or assumed, a general chatbot can interpret a result on the wrong scale. A glucose of 5.5 mmol/L equals about 99 mg/dL, so reading one as the other completely changes the meaning.
Why are reference ranges different between labs?
Each lab sets or verifies its ranges for its own instruments, methods and patient population, and ranges can vary by age, sex and pregnancy. Always read a result against the range printed on your own report.
What is an AI hallucination on a blood test report?
It is confident output not supported by your data — an invented value, a marker that was not tested or a made-up reference range. Our glossary defines this and related terms.
Do dedicated AI blood test analyzers make mistakes?
They can. They are designed to reduce common failure modes, but no AI tool is error-free. Verify values against your report and discuss results with a licensed clinician.
Sources
- MedlinePlus — How to Understand Your Lab Results — U.S. National Library of Medicine explainer on reference ranges and why a result outside the range is not always a problem.
- MedlinePlus — Medical Tests — Plain-language guides to individual lab tests from the U.S. National Library of Medicine.
- MedlinePlus — Comprehensive Metabolic Panel (CMP) — The tests included in a CMP and what they assess.
- WHO — Ethics and governance of artificial intelligence for health (2021) — World Health Organization guidance on safety, transparency, privacy and accountability for AI in health.
- U.S. FDA — Artificial Intelligence and Machine Learning in Software as a Medical Device — How the FDA approaches AI-enabled software intended for medical purposes.
Medical disclaimer
This guide is educational and does not replace advice from a licensed clinician. If you have urgent symptoms, contact your local emergency number.