AI coding implementation: the 90-day playbook.
A real implementation calendar for AI medical coding. Twelve weeks from kickoff to cutover. Four phases, named deliverables per phase, a training plan that touches three roles, a scorecard the CFO can read, and the five mistakes that kill these projects before they finish. Built from actual client rollouts, not from a vendor brochure.
The thesisAI coding is a workflow project. Not a software install.
Every AI coding implementation that we have seen fail has failed for the same reason. The practice treated it as a software install, the vendor was happy to be treated as a software install, and the operating team discovered six weeks after go-live that nobody had actually done the workflow design. The chart kept arriving. The AI kept coding. The reconciliation queue grew. The denial rate kept climbing. By month four, the executive sponsor was looking for someone to blame.
This playbook treats AI coding implementation the way we run them on our active book: as a structured 90-day workflow project with four named phases, a baseline that gets measured before anything else happens, a parallel run that is non-negotiable, and a cutover with a hyper-care window. The total elapsed time is 12 weeks. The total dollar lift over the following 12 months on a representative book is between 4 and 9 percent of net revenue. The implementations that miss that lift are usually missing one of the four phases below.
Compressing the parallel run is the single most predictable way to miss the dollar lift. Skip nothing.
Twelve weeks. Four phases.
A pragmatic 90-day rollout for a practice that wants the lift without disrupting cash flow. The phases below are sequential. Compressing or skipping any one is the most reliable predictor of a failed implementation.
Measure before you build.
- Pull 6 months of prior coding output (chart, ICD-10, CPT, modifier, RVU, charge, denial).
- Capture payer mix and top-50 CPT volume.
- Record current coder productivity (charts/day, accuracy, time per chart).
- Document current denial rate by reason and dollar.
- Lock the success metrics scorecard with CFO.
- Name the implementation team: project lead, coding lead, CFO sponsor, IT lead.
Build the rule pack.
- Configure AI on a sandbox copy of historical charts.
- Tune the specialty rule pack against the baseline coder output.
- Map the reconciliation queue UI and workflow.
- Build coder review screen with override logging.
- Define AI confidence-score thresholds for auto-bill versus human review.
- Train coders on the workflow (4 hours per coder).
Both code. Reconcile daily.
- AI codes every chart. Human coder also codes every chart.
- Senior auditor reconciles differences daily.
- Track AI accuracy, override rate, confidence-score distribution.
- Tune rule pack based on actual differences, not anticipated ones.
- Identify code categories AI is trusted on and categories that still need human-first.
- Daily 30-min reconciliation huddle.
Flip with safety net.
- AI codes trusted categories with human review on confidence-low outputs.
- Daily reconciliation continues with 25 percent random sample.
- Senior auditor on call for escalation.
- CFO scorecard reviewed weekly.
- Day 75: first formal scorecard review with all metrics.
- Day 90: transition project governance to operating cadence.
Required data points.
The baseline is the most under-resourced phase of every failed implementation. Pulling these data points in the first seven days is the single best predictor of a clean cutover at day 90.
Six months of coded charts
Chart text or chart ID, coded ICD-10, coded CPT, modifier set, RVU, charge amount, payer of record, denial outcome. The training set.
Payer mix breakdown
Commercial, Medicare, Medicare Advantage, Medicaid, Medicaid managed care, self-pay. By volume and by dollar.
Top-50 CPT volume list
By volume across the trailing 6 months. The list anchors the rule-pack priority and the parallel-run reconciliation focus.
Denial reason breakdown
By CPT, by reason code, by dollar. The post-cutover scorecard compares against this. Without it, success is unprovable.
Coder productivity baseline
Charts per day per FTE, accuracy rate against chart audit, average time per chart. The productivity lift gets measured against this.
Audit defensibility documentation
How the practice currently documents coding decisions for audit defense. The AI workflow has to preserve or improve this. Skipping it triggers compliance risk.
HCC and RAF capture (if applicable)
For HCC and value-based books, the V28 RAF baseline by patient. The implementation has to improve specificity capture without inflating risk inappropriately.
Time-to-final-bill
From DOS to claim drop, by service line. AI coding typically halves this number. Without the baseline, the improvement is invisible.
Team training: three roles, four hours each.
Twelve total hours of structured training across the implementation team. None of the three roles can be skipped. The role-specific content matters; generic "AI overview" training does not change behavior.
Workflow training.
- How the AI review screen works in the EHR or coding tool
- When to accept the AI suggestion versus override
- How to document the override (free-text plus structured reason)
- How AI confidence scores map to required human review
- Reading the daily reconciliation summary
- Escalation path for ambiguous cases
Rule-pack governance.
- How to file a rule-pack update request
- How rule updates flow through QA and production
- Reading the AI confidence-score distribution
- Running the daily 30-minute reconciliation huddle
- Producing the weekly scorecard for CFO review
- Quarterly payer-policy refresh workflow
Scorecard interpretation.
- The six KPIs on the scorecard and what each measures
- What "good" looks like for each KPI
- Action to take when each KPI moves the wrong direction
- The relationship between coder override rate and rule-pack tuning
- Audit defensibility under the AI workflow
- How to brief the board on the implementation outcomes
Six metrics. Weekly review.
The scorecard is the implementation's contract with the CFO. Each metric has a baseline (from Phase 01), a target for day 90, and a target for day 180. The CFO sees this every week from day 30 onward.
| Metric | Baseline | Day 90 target | Day 180 target | Why it matters |
|---|---|---|---|---|
| Coding accuracy vs chart audit | 92-95% | ≥ baseline | +1.5 pts | Audit defensibility. AI must not move accuracy backward. |
| Initial denial rate | 11-13% | ≤ baseline | -2 pts | The dollar lift comes from here. Material movement is required. |
| Time-to-final-bill | 4-6 days | -30% | -50% | AR aging compresses. Cash velocity improves. |
| Coder productivity (charts/FTE/day) | 40-55 | +40% | +90% | Workflow lift. Volume per FTE roughly doubles by day 180. |
| RAF capture (HCC books) | Specialty-dependent | +3-5% | +5-8% | Specificity prompts at coding time drive lift. Defensible. |
| Audit defensibility (sampled) | Baseline pass rate | ≥ baseline | +5 pts | Compliance risk does not increase. Structured override log helps. |
Five mistakes to avoid.
Every failed implementation we have seen made at least one of these. Most made two. Reading them upfront is materially cheaper than learning them in production.
Skipping the baseline measurement.
The practice opens with "we know our numbers, let's just install the AI." Three months later the scorecard is empty and nobody can prove the implementation helped, because there is nothing to compare against. The vendor declares success. The CFO is unconvinced. The implementation gets quietly rolled back.
Compressing the parallel run.
The parallel run is expensive. Two sets of coding cost more than one, and the daily reconciliation eats senior auditor time. Practices try to cut it from 30 days to 10. The 10-day parallel does not surface the AI's actual failure modes on real production patterns. Cutover lands on a brittle workflow that breaks in the first month.
Treating AI as a black box without rule-pack governance.
The vendor ships the model, the practice installs it, nobody owns the rule pack. Payer policies change. The rule pack does not. Denials climb. The coding lead does not know how to file a rule-pack update. The vendor does not know the practice's payer mix matters. Three months later the implementation is underperforming for entirely fixable reasons.
Picking AI on demo accuracy rather than the specialty rule library.
Every AI coding vendor demos at 95 percent accuracy. The number is real and the number is misleading. The accuracy that matters is on your specialty, your payer mix, your charts. Vendors with strong demos and weak specialty rule libraries miss on the codes that drive your dollars. The selection criterion should be the rule library depth, not the demo number.
Cutting over without a hyper-care window.
Cutover lands on a Friday. The next 30 days are nominally "post-implementation." Nobody is watching the AI output in real time. By the time the first weekly scorecard surfaces a problem, three weeks of claims have shipped on broken output. The fix is operationally painful and politically expensive.
Day 91 and beyondFrom project to operating cadence.
Day 90 is not the end of the implementation. It is the transition from project governance to operating cadence. The project lead steps back. The coding lead takes ongoing ownership of the rule pack. The senior auditor moves from daily reconciliation to a 10 percent sampling cadence. The CFO scorecard shifts from weekly to monthly. The implementation becomes part of how the practice runs, not something the practice is doing alongside its normal operations.
The single most important post-day-90 commitment is the quarterly rule-pack refresh. Payer policies will continue to update. CPT additions will continue to land annually. State Medicaid rules will continue to drift. The rule pack that was tuned on day 30 will be stale by day 180 unless it is maintained. Build the maintenance cadence before day 90 ends, while the implementation muscle is still strong, not three months later when the team has moved on.
The honest summaryThe lift is real. The discipline is the cost.
Across the implementations we have run, the dollar lift over the 12 months after cutover lands between 4 and 9 percent of net revenue on a representative book. The lift comes from denial reduction, faster cash velocity, productivity gains, and (for HCC books) RAF capture. The implementations that hit the high end of the range followed this playbook. The implementations that landed at the low end skipped one of the four phases. The implementations that failed skipped two.
The discipline is the cost. Baseline measurement is unglamorous. Parallel run is expensive. Rule-pack governance is a permanent role, not a one-time setup. Hyper-care eats senior auditor time. CFO scorecard reviews are recurring meetings. Practices that view this playbook as overhead are reading it backward. The lift comes from the discipline. There is no version of AI coding implementation where the lift shows up without the discipline behind it.
AI coding implementation, frequently asked questions.
The questions coding leads and CFOs most often ask when scoping or running an AI coding implementation.
How long does an AI coding implementation actually take?
What baseline data do we need before starting?
What is the parallel run phase and why is it non-negotiable?
What are the success metrics for AI coding?
What roles need training and how much?
What are the five mistakes that doom AI coding implementations?
How does AI coding affect coder roles?
What does day 91 look like once the implementation is done?
Want a 90-day implementation plan on your data?
A free 30-day implementation audit under a same-day BAA. The output is a written report covering your baseline data readiness, recommended rule-pack scope, expected dollar lift on your payer mix, named risks, and a week-by-week project plan. A senior partner on the call.