Home/AI/Autonomous Coding
92% tuned auto-accept · 78% day-one · coder-in-the-loop

Autonomous coding. The 92 percent honest number.

Every vendor demo of autonomous coding hits 99 percent. Every honest production number is lower. On tuned specialties, our auto-accept rate measured on a rolling 90-day window across the production book is 92 percent. On day one before tuning, it is 78 percent. We publish both because the second one is what a new client should plan on for the first 60 days, and pretending otherwise is the reason most autonomous coding rollouts collapse at month three.

CRC-certified coders in the loop MEAT evidence on every code Per-version accuracy tracking
The honest numbers

Three rates the vendor demo will not show you.

Every vendor leads with the highest number on their happiest day. We publish three because the gap between them is the real conversation. The auto-accept rate is a function of how long the model has seen your documentation, your payer mix, and your specialty edge cases. Numbers below are measured on our own production book.

Day-one baseline · pre-tuning
78%

Auto-accept rate on a new engagement before any tuning to your documentation patterns. This is the number a coder team should plan on for the first 60 days. Anyone showing you a higher day-one number is either cherry-picking the specialty or showing you their best client.

Tuned book · 90-day rolling
92%

Auto-accept rate after 60 days of tuning on a specialty we know well (ABA, common MA HCC, primary-care office E&M). Measured across the production book on a rolling 90-day window. The remaining 8 percent goes to a CRC, CPC, or CCS-certified coder for accept, modify, or reject.

Tuned book · clean-code rate
99.1%

Clean-code rate measured by first-pass acceptance at the clearinghouse on auto-accepted suggestions, before any payer touches them. The auto-accept rate is the upstream metric. The clean-code rate is what determines whether you get paid first pass. Both matter; only one shows up on the SLA.

Architecture

The coder-in-the-loop architecture.

The architecture has five steps. Each step is operationally tested, audited, and visible to the client through the same console our coders see. The 92 percent number is the output of all five steps working as designed. When the rate drops, we know which step is the cause because each one is independently measured.

The coder is not coding from scratch on the cases that route to them. They are reviewing a high-quality suggestion with MEAT evidence already extracted and the confidence score visible. Their time is spent on the cases where their judgment is the difference between captured and missed revenue, not on the cases where the answer was obvious.

Step 01 · Ingest
Chart, encounter, claim arrives
EHR feed normalized, PHI scrubbed for prompt input, payer context attached.
Step 02 · Model
Suggestion + MEAT evidence + score
Pinned model version, pinned prompt version. Output captured to audit log.
Step 03 · Decision
Auto-accept or coder route
Above the confidence floor and case is in-distribution: auto-accept. Otherwise route.
Step 04 · Validation
CRC, CPC, or CCS coder reviews
Accept, modify, or reject. Rejection logged against the model version. NPI captured.
Step 05 · Tuning
Rejection patterns feed the next version
Weekly tuning. Per-version accuracy tracked. Router rolls back on regression.
Five myths and the truth

Five things vendors say about autonomous coding. The truth on each one.

The autonomous coding category has its share of theater. Five of the most common claims we hear from prospects who have been pitched by other vendors, and what is actually true on each one. The truth is shorter and more boring than the myth.

Myth 01

99 percent of charts auto-coded, no coder needed.

Pitched as the future of coding. Demoed on a curated sample of straightforward office visits. Sold against a license-fee P&L.

The truth

Tuned book hits 92 percent. The 8 percent that escalates is where the revenue and risk live.

Coding is bimodal. The middle is easy. The edges are where the dollars and the audit exposure are. A coder on the edges is the right architecture; not having one is the wrong architecture.

Myth 02

The model learns your patterns in days, not months.

Sold as the reason rollout will be faster than the competition. Means the vendor is overfitting on a small sample.

The truth

60 days minimum to tune a specialty. The model gets there fast on the common cases and slowly on the long tail.

Days-not-months on common documentation patterns. Months on edge cases, unusual payer rules, and specialty-specific modifier patterns. The tuning rate is asymptotic, not linear.

Myth 03

RADV defense is automatic from the AI audit trail.

Implied to mean the auditor will accept the AI evidence without further review.

The truth

RADV defense is the MEAT documentation plus the certified-coder validation event. The AI is the upstream evidence pipeline.

The model extracts the MEAT and surfaces it; the certified coder validates and signs off; the audit trail captures both. The auditor reviews the chart, the MEAT, and the validation. The AI does not get the auditor to skip review.

Myth 04

A horizontal LLM with the right prompt is enough.

Said by every vendor who wrapped a foundation model in a CSS skin and called it a coding product.

The truth

The V28 condition map, the MEAT extraction rules, the three-way ABA match are domain primitives. The LLM is a layer in the stack, not the stack.

A horizontal model with no domain primitives will hit 60 percent and stop. The domain primitives plus the model plus the certified-coder loop are what get you to 92 percent and clean code.

Myth 05

Auto-accept rate is the metric that matters.

Heavily advertised because it is the easiest number to make look big in a slide deck.

The truth

Captured revenue, clean-code rate, and RAF lift are the metrics on the SLA. Auto-accept is upstream of them, not a substitute.

A 98 percent auto-accept with a 70 percent clean-code rate is worse than a 90 percent auto-accept with a 99 percent clean-code rate. The SLA measures the downstream outcome, not the upstream theater.

Common questions

Frequently asked questions: autonomous coding.

What does autonomous coding mean in 2026?
In 2026, autonomous coding means the model produces a coding suggestion (CPT, HCPCS, modifier, ICD, HCC) with MEAT evidence and a confidence score, and that suggestion either auto-accepts because it is high-confidence on a tuned case, or routes to a certified coder for the final accept. We do not call it autonomous when the rejection rate is 30 percent and the coder is reworking every chart. We do call it autonomous when the model is right often enough that the coder spends their time on the cases where their judgment is the value-add.
What is the honest 92 percent number?
On tuned specialties (ABA, common MA HCC chronic-condition documentation, primary-care office E&M), our auto-accept rate has been measured at 92 percent on a rolling 90-day window across the production book. Tuned means we have run the model on the specialty for at least 60 days, validated the coder-rejection patterns, and updated the prompt versions accordingly. On a generalist day-one engagement before tuning, the auto-accept rate is closer to 78 percent. We publish both numbers because the second one is what a new client should plan on for the first 60 days.
What is coder-in-the-loop?
Coder-in-the-loop is the architecture where every model suggestion either auto-accepts (the model is high-confidence and the case is in the trained distribution) or routes to a certified coder for accept, modify, or reject. The coder is not coding from scratch; they are reviewing the suggestion with the MEAT evidence already extracted, the confidence score visible, and the model version pinned. The coder rejection is logged against the specific model version and feeds the next tuning cycle.
How do you handle a case the model has never seen?
Three ways. First, the confidence score drops below the auto-accept floor, the case routes to a coder, and the coder codes it with the chart in front of them as they normally would. Second, if a new specialty or new payer pattern appears, the model runs in AI-suggestion mode for the first 60 days while we measure rejection patterns and tune. Third, when an entire encounter type is out of distribution, the model declines to suggest, the coder gets the chart, and the case is flagged for the next tuning batch. The model knows what it does not know.
What about RADV and risk-adjustment audit?
Every HCC code carries the MEAT documentation it was extracted from, the model version that produced it, the prompt version pinned at the time, the confidence score, and the certified coder who validated it. When CMS or a RADV contractor asks for support on a specific code, the audit response is a generated PDF with the chart pages, the highlighted MEAT, the validation event, and the coder NPI. The audit trail is the product.
Are you replacing the coder?
No. We are replacing the parts of the coder job that are repetitive, low-judgment, and low-defensibility. The coder still owns the rejection, the modification, the unusual case, and the final accept. The model gives them back the hours they were spending on charts where the answer was obvious. The coder spends those hours on charts where their judgment is the difference between captured and missed revenue. The economics work for the coder, for the practice, and for the auditor.
How is your autonomous coding different from competitor offerings?
Three structural differences. First, every coder is CPC, CRC, or CCS certified. We do not delegate the in-the-loop role to non-credentialed reviewers. Second, every code carries the source MEAT evidence and the model version. The audit trail is not a tab in a separate report; it is the primary view. Third, we run the AI as part of a managed service with a senior partner accountable for the SLA on RAF lift, denial rate, or first-pass acceptance. The economics do not work unless the AI actually produces clean, defensible code.
When should I expect to see the 92 percent number?
By the end of month two for tuned specialties; by the end of month four for everything else. The first 30 days, the model is being calibrated against your specific documentation patterns, payer mix, and EHR. The first 60 days, the auto-accept rate is climbing from the day-one 78 percent baseline. By month four, the auto-accept rate stabilizes at or above 92 percent on the tuned book. Throughout the journey, the SLA is measured on captured revenue and clean code rate, not on auto-accept percentage.

Run the model on your charts. Get the honest number.

Free 30-day coding audit on a 90-day encounter sample under a same-day BAA. We run the model on your real documentation, measure the day-one auto-accept rate, project the post-tuning rate, and deliver a four-page written report with sample evidence and the implementation roadmap.