How to evaluate HCC coding AI for risk adjustment. The V28 lens.
Risk adjustment is the highest-stakes coding category in healthcare. RAF scores drive Medicare Advantage payment, ACO REACH benchmarks, and a growing share of Medicaid managed-care revenue. CMS-HCC v28 rolled out across 2024 to 2026 and restructured the condition map. Vendors who are still operating on v24 logic are mis-coding now. Before any pricing conversation, run a V28 currency check. This is the buyer-side guide for everything that comes after.
The V28 currency check. Run it first.
CMS-HCC v28 is not a small update. CMS phased it in across payment years 2024, 2025, and 2026. The model retired and consolidated condition categories, the depression and substance-use buckets were restructured, several diabetes complication codes were resequenced, and coefficient weights shifted across the entire map. Vendors who built their HCC engine on v24 and never refreshed it are mis-coding today. They are over-coding conditions that v28 retired and under-coding conditions that v28 added. The downstream effect is twofold. The RAF numbers they report are wrong, and the codes they push into your submissions create RADV exposure when CMS audits.
Before any pricing conversation, ask three questions. What model version are you current on. Do you support a v24 and v28 blended workflow for retrospective coding on older dates of service while running pure v28 for current dates of service. What is your refresh cadence when CMS issues mid-year clarifications and what is the SLA on getting that update into production. A vendor who cannot answer all three precisely is not ready for a Medicare Advantage book.
One more practical filter. Ask the vendor to walk you through ten randomly sampled charts and show you which HCC categories changed between v24 and v28 on those specific charts. The vendors who actually run a current engine can do that in the demo. The vendors who cannot are running marketing slides, not software.
Ten questions. Ask them in order.
These are the ten qualifying questions that every HCC coding AI vendor must answer with specificity before pricing matters. They are ordered. If the vendor stumbles on question one, the answer to question ten is irrelevant. Print this list. Bring it to every vendor demo.
What CMS-HCC version are you current on?
Pure v28 for current dates of service. Blended v24 and v28 for retrospective. Refresh SLA in days, not quarters. Anything less means obsolete coding shipped to your submissions.
Do you check for MEAT documentation in every encounter?
The condition must be Monitored, Evaluated, Assessed, or Treated in the current encounter note. Vendors who pull from problem lists without MEAT evidence create RADV risk.
What is your RADV defensibility methodology?
Ask for a simulated RADV pass rate on a sample of your own charts. The right metric is audit-defensible RAF lift, not raw lift. Vendors who do not quote a defensibility rate are selling lift you may pay back.
Can you run recapture campaigns for prior-year unconfirmed HCCs?
Conditions that were coded last year but not confirmed this year drop off the RAF. The vendor must run recapture as a continuous workflow with clinician-confirmation surfaces, not as an annual report.
How do you handle chronic vs acute distinction?
Many conditions code as one or the other depending on documentation language. The AI must surface the linguistic cues correctly and not roll an acute condition into a chronic HCC. Test it on diabetes complications and CHF.
What is your false-positive rate?
Over-coding is a RADV nightmare and a clinician-trust killer. Ask for the precision metric (percentage of AI-suggested HCCs that the coder accepts after review). Below 80 percent means too many false positives.
Do you support prospective and retrospective workflows?
Prospective surfaces suspected HCCs to the clinician inside the encounter so the documentation is born compliant. Retrospective catches missed HCCs after the encounter closes. You need both, not one.
How does it integrate with the EHR problem list?
The AI must reconcile its suggested HCCs against the patient's active problem list, the prior encounter notes, and the current encounter document. Problem-list integration is required, not optional. Ask for the supported EHRs by name.
What is your audit trail when CMS comes asking?
For every coded HCC, the system must produce the MEAT sentence, the source document, the timestamp, the coder who confirmed, and the version of the engine that produced the suggestion. Anything less fails RADV.
Who owns the coded data and RAF lift attribution?
Read the contract. Some vendors retain rights to the encoded outputs. Some bill on lift attribution they cannot independently verify. You need contractual data ownership and a transparent attribution methodology.
MEAT. The four pillars of RADV defensibility.
CMS requires that an HCC condition be Monitored, Evaluated, Assessed, or Treated in the encounter note for the code to be RADV defensible. AI that surfaces a condition from a problem list without MEAT evidence is creating audit risk, not revenue. Real HCC coding AI walks the chart and surfaces the MEAT sentence for every coded HCC.
The MEAT standard is the line between defensible coding and audit risk. A condition listed on the problem list without any sentence in the current encounter note that monitors, evaluates, assesses, or treats it does not survive a RADV pull. The auditor will demand documentation contemporaneous with the date of service. If the documentation does not exist, CMS recoups the payment with interest, and pattern-of-conduct violations cascade across the entire plan.
The vendor evaluation question is whether the AI presents the MEAT evidence sentence inline with every suggested HCC. Not whether the AI claims to check MEAT. Not whether MEAT is in the vendor's marketing deck. Whether, on a sample of your own charts, the AI surfaces the actual sentence from the actual encounter document that proves the condition meets one of the four pillars. The vendors who do this are running real natural-language processing against the clinical text. The vendors who do not are pulling from problem lists and calling it AI.
One useful test. Pick a chart where a condition is on the problem list but is not mentioned anywhere in the current encounter note. Does the vendor suggest the HCC anyway? If yes, that vendor is going to fail your next RADV audit.
Monitored
Status checked. Ordered labs, reviewed vitals, tracked symptoms. The note shows ongoing surveillance of the condition.
Evaluated
Reviewed for change. Considered against history, prior results, current findings. The note shows clinical judgment was applied.
Assessed
Documented in the assessment. Carries a status (stable, worsening, improved). The note shows the condition was clinically reasoned about.
Treated
Active intervention. Medication adjusted, referral placed, procedure planned. The note shows the clinician acted on the condition.
RADV. The only RAF lift that matters.
CMS Risk Adjustment Data Validation audits sample charts every payment year and demand contemporaneous documentation for every HCC submitted. The defensibility metric is the percentage of submitted HCCs that survive the simulated audit. Vendors who quote raw RAF lift without a defensibility rate are selling lift you may have to pay back with interest. Vendors who cannot produce a RADV simulation on a sample of your own charts are not ready for an MA contract.
The right buyer-side conversation does not start at RAF lift. It starts at defensibility rate. Ask the vendor to show you a RADV simulation pass rate on a stratified sample of your charts before signing. A reputable vendor will run the simulation, surface the HCCs that would fail audit, explain why each one failed (typically missing MEAT, problem-list-only sourcing, or specificity mismatch), and quote the defensible RAF lift after the failing HCCs are pulled out.
That is the number you negotiate against. Defensible RAF lift, not raw RAF lift. The difference between those two numbers is the recoupment risk you would otherwise carry on the plan books.
Recapture campaigns. Not an annual report.
Prior-year HCCs that are not confirmed in the current calendar year drop off the RAF score. CMS does not carry chronic conditions forward. Every condition must be re-documented every year, with MEAT evidence, in a fresh encounter, on a fresh date of service. The plans that win on RAF score are the plans that run a continuous recapture campaign, not the plans that run an annual sweep in November.
Real HCC AI surfaces the prior-year unconfirmed HCCs to the clinician inside the workflow. Before the encounter starts, the AI presents a list of conditions that need re-documentation, with the suggested MEAT prompts. After the encounter, the AI verifies that each prior-year HCC was either re-documented with MEAT, ruled out clinically, or flagged for follow-up. This is the workflow that separates real HCC coding AI from a report-running vendor.
The vendor question is whether the recapture workflow runs continuously or whether it dumps a list of unconfirmed conditions on the coder once a year. If the answer is the latter, the vendor is selling a spreadsheet, not software.
Pre-encounter surface.
Prior-year unconfirmed HCCs presented to the clinician before the visit starts, with suggested MEAT prompts and current chronicity evidence.
In-encounter capture.
As the clinician documents, the AI flags MEAT-compliant language for prior-year conditions and prompts when a condition is mentioned without sufficient documentation.
Post-encounter verify.
For every prior-year HCC, the AI confirms one of three outcomes: re-documented with MEAT, ruled out clinically with documentation, or flagged for follow-up at the next visit.
Quarterly reconciliation.
The recapture rate, the rule-out rate, and the still-open rate roll up to a panel-level dashboard. Open conditions get re-queued for the next encounter automatically.
We are a vendor too. Honest framing.
This page is a buyer-side guide and we are also an HCC coding AI vendor. We built this guide because we wanted prospective buyers to be able to compare any vendor, including us, against a real specification. The ten questions, the V28 currency check, the MEAT framework, the RADV defensibility rate, and the recapture workflow are the same evaluation criteria we hold ourselves to. If you run the guide against us, we expect you to find some things you like and some things you want to improve. That is the right buyer posture.
Our HCC coding engine runs pure v28 with v24 blend support for retrospective dates. It surfaces MEAT evidence inline with every suggested HCC. The RADV simulation runs on a sample of your charts before contracting. The recapture workflow is continuous, not annual. We integrate with the major ambulatory EHRs and we publish the precision and recall numbers per condition family. Our product page is linked below. Read it after you have read this guide, not before.
Frequently asked questions: HCC coding AI vendor evaluation.
What is CMS-HCC v28 and why does it matter for vendor selection?
What is MEAT documentation and why do AI vendors need to surface it?
How is RADV defensibility actually measured?
What ROI is realistic on HCC coding AI for a Medicare Advantage population?
Should I focus on prospective or retrospective HCC workflows first?
How do I prove RAF lift attribution to the AI vendor?
Does HCC coding AI matter outside Medicare Advantage?
How does HCC coding AI connect to clinical documentation improvement (CDI)?
Get the buyer scorecard. Free.
The full evaluation worksheet with the V28 currency check, the 10 questions, the MEAT framework test, the RADV simulation request template, and the recapture workflow checklist. Bring it to every HCC coding AI vendor demo. A senior partner on the call to walk through scoring with you.