Access the original post here.
On the surface, AI scribes are meant to reduce the time we physicians spend typing in the EHR so we can focus more on patients. More or less, they are doing that.
But a second benefit is starting to get more airtime. Depending on who you ask, it is either the most exciting thing about these tools or the most troubling:
- Ambient scribes are great at coding accuracy.
In this article, I look at how AI scribes are changing coding accuracy and why that matters for physicians.
The Deets: Coding Accuracy
When we see patients, we document what we are treating. We do not naturally think in ICD-10 codes in the exam room (although, maybe some of us do). If a patient walks in with a chronic cough, mild hyponatremia, and a recently changed diuretic, we manage all three, but we might only formally code for one.
Clinical Documentation Integrity (CDI) teams exist because of this gap. Their job is to chase down every billable condition after the fact and make sure it gets captured. It is labor-intensive, inconsistent, and structurally slow. AI can close that gap in real time with ambient scribes.
Ambient scribes listen to the full encounter, capturing every diagnosis mentioned and every condition discussed. Some platforms go further by layering on retrospective coding AI that scans the full chart and surfaces additional billable diagnoses a physician referenced but never formally listed. The ROI numbers from early adopters are significant. One ambient vendor markets roughly $13,000 per clinician annually in recovered revenue.
From a pure documentation standpoint, this is an improvement in coding accuracy, since we were likely undercoding before and AI is correcting that.
Trilliant Health’s Data on Coding After AI Scribe Adoption
A recent analysis from Trilliant Health examined outpatient E/M billing patterns at six large health systems that have publicly adopted ambient AI scribing, using national all-payer claims data from 2018 to 2024. Across every system, visits shifted toward higher-intensity codes for both new and established patients.
For new patient visits, the share billed at the highest acuity levels (CPT 99204–99205) rose by 12 to 20 percentage points across all six systems (orange line). One health system saw 80% of new patient visits billed at high intensity by 2024. For established patients, the increase ranged from 7 to 12 percentage points. These changes were consistent across geographically and organizationally distinct institutions, all of which adopted ambient AI during the study period.
The increase appeared across nearly every diagnosis category. For factors influencing health status, respiratory disease, and mental and behavioral disorders, coding intensity rose across the board.
Trilliant's interpretation is measured. They argue the increase likely reflects better rules-based documentation rather than fraud. Ambient AI records every word spoken. A tool that captures everything clinically relevant will naturally generate more complete, higher-acuity documentation than a physician typing notes in Epic at 11 PM. And with the 2021 E/M coding revisions that shifted emphasis toward medical decision-making and total time, more encounters now legitimately qualify for higher codes.
Trilliant is also candid about the limits of what its analysis can show:
- No control group of health systems that did not adopt ambient AI, which means we can't isolate the technology's contribution from broader secular trends in coding intensity.
- It can't distinguish between documentation that accurately reflects clinical complexity and documentation that overstates it. That distinction between accurate coding and upcoding is precisely what's in dispute, and the data alone can't answer it.
A single outpatient visit coded one level higher might mean $10–40 more in reimbursement. Across millions of visits, that compounds fast.
Blue Cross Blue Shield’s Study
In my initial article, How AI Documentation Tools Are Making Upcoding Worse, I covered what Blue Cross Blue Shield's research arm found when it examined the same question across 62 million commercial members.
Their analysis flagged postpartum anemia as a case study: coding for that diagnosis tripled at the highest-growth hospitals between 2022 and 2025, a surrogate for AI adoption, while transfusion rates, the standard treatment for clinically significant postpartum anemia, stayed essentially flat. One BCBS plan audited a major outlier hospital system and found that fewer than 20% of cases coded with postpartum anemia actually met established clinical criteria.
BCBS estimated that coding intensity shifts in maternity admissions alone added roughly $22 million in spending over the study period.
The AI tools are not making clinical decisions, but they are doing exactly what they were designed to do: capture everything (Deming’s Principle!). In a fee-for-service system that rewards documentation volume, "capturing everything" and "maximizing revenue" can end up being the same thing.
