AI scribes are having a real moment in healthcare, and for good reason. They save clinicians real time and cut down on the documentation burden that pushes people out of the field. But there's a question that doesn't get asked enough: what happens when the scribe writes down something that doesn't match what a certified coder, or a careful second read of the actual encounter, would conclude?
That question matters more in OASIS specifically than in general medical documentation, and it's worth understanding why.
What the research actually shows
Published research on ambient AI scribes has started to paint a clearer picture of where these tools tend to fall short, and the pattern is worth paying attention to. One instrument validation study evaluating commercially available AI scribe products found real documentation errors worth flagging for patient safety, not just typos or phrasing issues. A separate randomized controlled trial found that when AI scribes get something wrong, the more common failure is an omission, something left out of the note entirely, rather than an outright inaccuracy. The AI is more likely to miss something than to state something false.
That distinction matters a lot for how you think about risk. A stated error is usually easier to catch, since it stands out as something specific to check. An omission is quieter. Nothing looks wrong on the page, because the thing that should've been there just isn't.
Why this is a bigger deal in OASIS than in a general visit note
A general medical note has some flexibility. If a scribe slightly under captures a detail of a conversation, a clinician reading it back can often fill in the gap from memory, and the consequence is usually limited to that one encounter.
OASIS doesn't work that way. Every functional item, every M-code answer, every piece of documentation supporting a diagnosis feeds directly into scoring that determines case mix weight and, increasingly, HHVBP performance. An omission isn't just a documentation gap, it's a scoring gap, and scoring gaps turn into payment gaps.
If an AI scribe captures a patient's self care description in a way that omits the specific level of assistance they needed, that's not a minor missed detail, that's the exact information a [GG0130 score depends on](insert GG0130 blog URL). The scribe didn't get anything wrong, technically. It just didn't capture the thing that mattered most for the score that came next.
Where the coder comes in
This is exactly the gap a certified coder is positioned to catch, if the workflow is actually built to let them. A coder reading a scribe generated note against the OASIS items it's supposed to support can catch the moment where the note is technically accurate but functionally incomplete, the omission the research keeps flagging as the more common failure mode. It's the same principle behind [how AI and certified coders work together in a chart review](insert AI plus certified coders blog URL): the technology handles the pattern matching and the coder brings the judgment to fill in what the pattern alone would miss.
That only works if the coder is actually reviewing against the specific OASIS items, not just accepting the scribe's note as a finished product. A scribe that saves time on the front end doesn't help if the coder on the back end isn't positioned to catch what got left out.
What we’re seeing in real charts using scribe
One of the biggest gaps in evaluating AI scribes is measuring true recommendation accuracy. It's not enough to report how many recommendations clinicians accept - you also need to measure false positives (incorrect recommendations) and false negatives (missed recommendations).
For example, some vendors report a 96% clinician acceptance rate, but what does that really mean? Are the recommendations truly accurate, or are clinicians accepting blindly and taking the "accept all approach" because it's faster? Most clinicians aren't OASIS experts and can't be expected to validate every recommendation against the OASIS guidance.
Our own retrospective reviews of scribe-generated OASIS assessments from existing vendors in the market continue to find meaningful documentation and coding errors. That's why clinician acceptance is a poor proxy for quality - true accuracy requires expert OASIS review and auditing.
How agencies should be evaluating scribes
If you're evaluating or already using an AI scribe, ask specifically how omissions get caught, not just how accuracy is measured. Ask whether the workflow includes a coder cross checking the note against the specific OASIS items it needs to support, or whether the note is treated as final once the scribe produces it.
The time savings from an AI scribe are real. The risk isn't that the tool lies to you, it's that it quietly leaves something out, and OASIS is exactly the kind of documentation where a small omission turns into a real scoring problem.
Want to see how this fits together
If you're exploring an AI scribe for your agency, we'd like to show you how Scribe is built alongside a coding review layer that catches exactly this kind of gap.
Schedule a call to learn more about Scribe.
AI Scribe, OASIS, accuracy, home health AI




