How to evaluate medical coding AI vendors. The honest checklist.
Most vendor demos show you a beautifully tuned specialty on curated data. That is not what production looks like. This is the buyer's guide nobody else writes, because most people writing it are also selling. We are too. So we wrote down exactly what we would ask if we were on your side of the table, and the questions every honest vendor should be willing to answer in writing.
The 92 percent number is not the number.
Almost every medical coding AI vendor leads with a single accuracy figure. 92 percent. 94 percent. 96 percent. The figure is real in the narrow sense that the vendor measured something and got that result. The problem is what they measured.
Vendor accuracy claims are almost always pulled from a single tuned specialty, usually outpatient surgery, ED facility, or radiology professional, where coding patterns are highly templated and the model has been trained on tens of thousands of similar charts. Move the same model to ABA, behavioral health intake, FQHC sliding-scale primary care, hospital inpatient with comorbidities, or specialty drug J-codes, and the number falls. We have audited vendors who quote 94 percent in marketing and produce 68 percent code-level agreement on a buyer's own generalist data.
The defense is mechanical. Demand the test set composition: how many charts, what specialty mix, what payer mix, what date range, who coded the gold standard. Demand specialty-by-specialty breakouts on your own production data, not theirs. Demand code-level agreement (ICD-10-CM, CPT, HCPCS, modifiers, sequence) rather than encounter-level agreement, which can hide a missing modifier that costs you 18 percent of the reimbursement. AHIMA and AAPC both publish coder accuracy benchmarks (95 to 97 percent for credentialed coders on production work), and any AI claim above that floor on generalist data should trigger a hard look at the methodology.
Ten questions every vendor must answer in writing.
Print this. Send it before the demo. The vendors who respond clearly and in writing are the ones worth the second meeting. The vendors who deflect to a slide deck are the ones who will deflect again at month six when accuracy slips and you need answers.
What is your accuracy on my data, not your demo data?
The single most important question. Demand a POC on 200 to 500 of your own de-identified charts, stratified by your specialty and payer mix. Demand code-level agreement, not encounter-level. Demand specialty-by-specialty breakouts.
→ Red flag: vendor refuses or wants to charge for the POCWhere is the audit trail for every code decision?
For every code the AI assigns, you need the source text it cited, the rule or guideline it applied, the confidence score, and the time stamp. Without an audit trail you cannot defend a payer audit or a RAC review. Compliance is yours, not the vendor's.
→ Red flag: explanations are post-hoc or genericIs this a coder-in-loop workflow or a black box?
The best systems route low-confidence codes to a human coder with the AI's reasoning attached, and learn from every override. The worst systems output a final claim and expect you to trust it. Ask to see the coder UI in production, not a demo.
→ Red flag: no override mechanic or override is invisible to the modelHow does it integrate with my EHR?
HL7 v2, FHIR R4, direct API, SFTP file drop, or a screen-scrape browser extension. Each has cost, latency, and reliability tradeoffs. Get the integration approach in writing, with the named EHRs they have shipped to production, the engineering hours estimated for your stack, and who owns the connector long term.
→ Red flag: "we will figure it out during onboarding"What specialties are tuned vs untuned?
Every vendor has a tuned set and an untuned set. The tuned set is where their training data is dense; the untuned set is where you should expect 15 to 25 points lower accuracy. Demand the list. If your specialty mix overlaps the untuned set by more than 30 percent, that is the wrong vendor for you.
→ Red flag: "we work across all specialties equally"What is the support model when a coder disagrees?
Coder pushback is the loudest signal you will get in the first 90 days. If a coder marks 30 charts as miscoded, who reviews the dispute, on what cadence, and does the model retrain on the override? Vendors who have no answer here will lose coder adoption and quietly disappear from the workflow.
→ Red flag: support is a ticketing queue with no SLAWhat is the pricing structure?
Per-chart, per-coder, per-encounter, per-month flat, or per-percentage-of-collections. Each has incentive distortions. Per-chart pricing punishes complex cases. Per-percentage aligns vendor incentives but caps your margin. Demand a written quote that survives volume growth and specialty expansion without renegotiation.
→ Red flag: pricing only "on request"How fast is the feedback loop on payer policy changes?
CMS publishes thousands of LCD and NCD updates per year. Commercial payers publish hundreds more. When a policy changes, how long until the model reflects it, and who is responsible for catching the change? AHIMA and HFMA both track this as a top compliance risk for autonomous coding.
→ Red flag: no defined retraining cadence or policy monitoring teamWhat is your hallucination rate and how do you measure it?
LLM-based coding tools can fabricate codes that look plausible but are not in the documentation. The measure is rate of codes assigned where no source text justifies them. Less than 0.5 percent is acceptable for a coder-assist tool; less than 0.1 percent for an autonomous one. If the vendor does not measure it, assume it is high.
→ Red flag: "we do not use LLMs" with no architectural detailWho owns the data and what happens at contract end?
Three sub-questions. Does the vendor train on your data and benefit other customers from your patterns? On termination, do you get a complete export of every coded chart in a usable format? Is the model itself escrowed so a vendor failure does not strand your operation? These are the terms you negotiate before signing, not after.
→ Red flag: data ownership ambiguous or training rights asymmetricHow to score vendors side by side.
The matrix below is the scorecard structure we recommend. Vendor A, B, and C represent typical archetypes we have audited on behalf of buyers. We list ourselves in the last column because we are also a vendor and the buyer's framing is to evaluate everyone, including us, against the same bar.
We are a vendor. Here is what we publish honestly.
We are not above the comparison. We are in it. ASP-RCM sells medical coding AI as part of our broader revenue cycle stack, and the same checklist applies to us. If we cannot give you specialty-by-specialty accuracy on your data, written pricing, an audit trail screenshot, and a 90-day pilot structure inside the first 48 hours, you should drop us from the shortlist.
What we publish: the engine architecture (deterministic HCC engine with a curated V28 crosswalk and hierarchies, not a black-box LLM that hallucinates plausible codes), the tuned specialty list (ABA, HCC risk adjustment, behavioral health, FQHC primary care, hospital outpatient), the coder-in-loop workflow with override telemetry, and the audit trail design that ties every code to its source text, rule citation, and confidence score.
What we do not claim: that we are best on every specialty, that accuracy is 95 percent across the board, or that any AI removes the need for credentialed coders. AHIMA and AAPC are right that human coder review remains the standard of care for production claims. We built our workflow around that, not against it.
Our chart to codes engine.
The actual production page for our HCC and risk-adjustment coding engine. Architecture, V28 crosswalk, UAT results, and the audit trail screenshots.
Product · ABA billing AIABA-specific coding and billing AI.
The tuned ABA workflow: 97151 to 97158 code assignment, supervision ratio enforcement, three-way match between auth, session, and claim.
Free buyer scorecardGet the weighted scorecard.
Send your last 30 days of charts and the vendor list you are evaluating. We return the scorecard filled in for every vendor including us, with evidence per criterion.
Frequently asked questions: evaluating coding AI.
How do I run a fair test of a medical coding AI vendor?
What accuracy number is realistic for medical coding AI?
Should I ask for a proof of concept or a production pilot?
What about HIPAA and PHI in test data?
How long should a medical coding AI pilot run?
How do I score vendors against each other?
What happens if AI accuracy degrades after go-live?
Who in my organization should own vendor evaluation?
Get the buyer scorecard. Free.
Send the names of the vendors you are evaluating and 30 days of de-identified charts. We return the weighted scorecard filled in for every vendor on the list, including us, with evidence per criterion and the questions still open against each. A senior coder on the review call. If we are not the right fit, we will tell you.