Why AI Coding Accuracy Numbers Don't All Mean the Same Thing

Why AI Coding Accuracy Numbers Don't All Mean the Same Thing

Why AI Coding Accuracy Numbers Don't All Mean the Same Thing

Cortney Swartwood

Cortney Swartwood

If you have looked at more than one AI coding vendor, you have probably seen the same claim over and over: "95 percent accurate." It sounds definitive. It is almost never the whole story, and the gap between a demo stat and a production stat is bigger than most agencies realize until they ask the right questions.

Here is what that gap actually looks like, and why we think independent verification matters as much as the number itself.

A demo number and a production number are not the same thing

A lot of AI accuracy claims in this space come from a handful of cherry picked charts run in a demo environment. Twenty clean charts, chosen to make the tool look good, will produce a great looking accuracy stat almost every time. That number tells you almost nothing about how the tool performs across thousands of real, randomly sampled charts running in daily production, with all the messy documentation, edge cases, and inconsistent notes that come with actual patient care.

This is the distinction worth asking about before you sign anything: was this number measured in a demo, on a small hand picked sample, or was it measured at scale, on charts nobody selected in advance?

We publish our own answer to that question every quarter. Our latest accuracy report is built from nearly 1,000 retrospective chart audits across more than 100 customers, not a curated demo set. The overall accuracy across that full sample came in at 97.5 percent, broken down by service line: 99.8 percent on Plan of Care, 99.1 percent on OASIS, and 95.1 percent on coding. No samples, no cherry picking, just a hard look at what is actually landing in the record.

Why a single accuracy number still is not the full picture

Picture an AI-only tool with no certified coder checking its work, one trained to flag a chart as clean or not. If it quietly learns that guessing 'no error' is a safe bet, and 95 percent of charts in a given sample are already clean, that tool scores 95 percent accurate too, without ever catching a real problem, and there's nobody in the loop to notice or explain it. That's the real risk with AI-only tools specifically: without a human checking the work, an inflated accuracy number can hide a lot, and there's no one positioned to catch it, let alone tell you how the number was actually reached.

The two questions worth asking underneath any accuracy claim are precision, meaning how often a flagged issue is actually a real issue, and recall, meaning how many of the real issues actually get caught. A vendor who only shows you one blended number is choosing not to show you that trade off, whether or not they realize it.

Why third party verification matters here

Self reported accuracy numbers are still self reported, even when they are measured honestly. That is part of why we went through CHAP Verification for our OASIS and Plan of Care Quality Review service. CHAP is the accreditation body agencies already trust for home health and hospice, and being CHAP Verified means an outside organization looked at our process, not just our marketing claims, and confirmed it holds up.

That distinction matters more in AI coding than almost anywhere else right now, because the space is full of vendors making accuracy claims with no outside party checking the work. A verified process is a different kind of claim than an impressive slide.

What to actually ask a vendor

Ask whether their accuracy number came from a demo sample or from production charts at scale. Ask for the breakdown by service line and error type, not just one blended figure. And ask whether any outside body has verified their process, or whether the number is entirely self reported.

See our full report

Our Q2-2026 Accuracy Report has the complete breakdown, along with where the remaining gaps tend to show up and why even a small miss can carry outsized weight under HHVBP.

home health coding, AI in healthcare, vendor evaluation

© 2026 EJJ HealthTech, Inc.

Support

support@ollihomehealth.ai

Connect with us

© 2026 EJJ HealthTech, Inc.

Support

support@ollihomehealth.ai

Connect with us

© 2026 EJJ HealthTech, Inc.

Support

support@ollihomehealth.ai

Connect with us