Home/Blog/AI coding implementation playbook
Implementation guide · 90-day playbook

AI coding implementation: the 90-day playbook.

A real implementation calendar for AI medical coding. Twelve weeks from kickoff to cutover. Four phases, named deliverables per phase, a training plan that touches three roles, a scorecard the CFO can read, and the five mistakes that kill these projects before they finish. Built from actual client rollouts, not from a vendor brochure.

90 daysEnd-to-end timeline 4 phasesBaseline · Sandbox · Parallel · Cutover 3 rolesTrained at 4 hours each 6 KPIsScorecard from day 1

The thesisAI coding is a workflow project. Not a software install.

Every AI coding implementation that we have seen fail has failed for the same reason. The practice treated it as a software install, the vendor was happy to be treated as a software install, and the operating team discovered six weeks after go-live that nobody had actually done the workflow design. The chart kept arriving. The AI kept coding. The reconciliation queue grew. The denial rate kept climbing. By month four, the executive sponsor was looking for someone to blame.

This playbook treats AI coding implementation the way we run them on our active book: as a structured 90-day workflow project with four named phases, a baseline that gets measured before anything else happens, a parallel run that is non-negotiable, and a cutover with a hyper-care window. The total elapsed time is 12 weeks. The total dollar lift over the following 12 months on a representative book is between 4 and 9 percent of net revenue. The implementations that miss that lift are usually missing one of the four phases below.

Compressing the parallel run is the single most predictable way to miss the dollar lift. Skip nothing.

The week-by-week timeline

Twelve weeks. Four phases.

A pragmatic 90-day rollout for a practice that wants the lift without disrupting cash flow. The phases below are sequential. Compressing or skipping any one is the most reliable predictor of a failed implementation.

D0-7 D7-30 D30-60 D60-90 BASELINE SANDBOX PARALLEL RUN CUTOVER + HYPER-CARE DAY 0 DAY 30 DAY 60 DAY 90
Days 0 to 7
Phase 01 · Baseline

Measure before you build.

  • Pull 6 months of prior coding output (chart, ICD-10, CPT, modifier, RVU, charge, denial).
  • Capture payer mix and top-50 CPT volume.
  • Record current coder productivity (charts/day, accuracy, time per chart).
  • Document current denial rate by reason and dollar.
  • Lock the success metrics scorecard with CFO.
  • Name the implementation team: project lead, coding lead, CFO sponsor, IT lead.
Days 7 to 30
Phase 02 · Sandbox

Build the rule pack.

  • Configure AI on a sandbox copy of historical charts.
  • Tune the specialty rule pack against the baseline coder output.
  • Map the reconciliation queue UI and workflow.
  • Build coder review screen with override logging.
  • Define AI confidence-score thresholds for auto-bill versus human review.
  • Train coders on the workflow (4 hours per coder).
Days 30 to 60
Phase 03 · Parallel run

Both code. Reconcile daily.

  • AI codes every chart. Human coder also codes every chart.
  • Senior auditor reconciles differences daily.
  • Track AI accuracy, override rate, confidence-score distribution.
  • Tune rule pack based on actual differences, not anticipated ones.
  • Identify code categories AI is trusted on and categories that still need human-first.
  • Daily 30-min reconciliation huddle.
Days 60 to 90
Phase 04 · Cutover + hyper-care

Flip with safety net.

  • AI codes trusted categories with human review on confidence-low outputs.
  • Daily reconciliation continues with 25 percent random sample.
  • Senior auditor on call for escalation.
  • CFO scorecard reviewed weekly.
  • Day 75: first formal scorecard review with all metrics.
  • Day 90: transition project governance to operating cadence.

Required data points.

The baseline is the most under-resourced phase of every failed implementation. Pulling these data points in the first seven days is the single best predictor of a clean cutover at day 90.

Data 01

Six months of coded charts

Chart text or chart ID, coded ICD-10, coded CPT, modifier set, RVU, charge amount, payer of record, denial outcome. The training set.

Data 02

Payer mix breakdown

Commercial, Medicare, Medicare Advantage, Medicaid, Medicaid managed care, self-pay. By volume and by dollar.

Data 03

Top-50 CPT volume list

By volume across the trailing 6 months. The list anchors the rule-pack priority and the parallel-run reconciliation focus.

Data 04

Denial reason breakdown

By CPT, by reason code, by dollar. The post-cutover scorecard compares against this. Without it, success is unprovable.

Data 05

Coder productivity baseline

Charts per day per FTE, accuracy rate against chart audit, average time per chart. The productivity lift gets measured against this.

Data 06

Audit defensibility documentation

How the practice currently documents coding decisions for audit defense. The AI workflow has to preserve or improve this. Skipping it triggers compliance risk.

Data 07

HCC and RAF capture (if applicable)

For HCC and value-based books, the V28 RAF baseline by patient. The implementation has to improve specificity capture without inflating risk inappropriately.

Data 08

Time-to-final-bill

From DOS to claim drop, by service line. AI coding typically halves this number. Without the baseline, the improvement is invisible.

Team training: three roles, four hours each.

Twelve total hours of structured training across the implementation team. None of the three roles can be skipped. The role-specific content matters; generic "AI overview" training does not change behavior.

Role 01 · Coders

Workflow training.

4hrs
Per coder, before parallel run starts
  • How the AI review screen works in the EHR or coding tool
  • When to accept the AI suggestion versus override
  • How to document the override (free-text plus structured reason)
  • How AI confidence scores map to required human review
  • Reading the daily reconciliation summary
  • Escalation path for ambiguous cases
Role 02 · Coding leads

Rule-pack governance.

4hrs
Per lead, during sandbox phase
  • How to file a rule-pack update request
  • How rule updates flow through QA and production
  • Reading the AI confidence-score distribution
  • Running the daily 30-minute reconciliation huddle
  • Producing the weekly scorecard for CFO review
  • Quarterly payer-policy refresh workflow
Role 03 · CFO + RC leadership

Scorecard interpretation.

4hrs
CFO, VP RC, and CMO if HCC
  • The six KPIs on the scorecard and what each measures
  • What "good" looks like for each KPI
  • Action to take when each KPI moves the wrong direction
  • The relationship between coder override rate and rule-pack tuning
  • Audit defensibility under the AI workflow
  • How to brief the board on the implementation outcomes
The success scorecard

Six metrics. Weekly review.

The scorecard is the implementation's contract with the CFO. Each metric has a baseline (from Phase 01), a target for day 90, and a target for day 180. The CFO sees this every week from day 30 onward.

MetricBaselineDay 90 targetDay 180 targetWhy it matters
Coding accuracy vs chart audit 92-95% ≥ baseline +1.5 pts Audit defensibility. AI must not move accuracy backward.
Initial denial rate 11-13% ≤ baseline -2 pts The dollar lift comes from here. Material movement is required.
Time-to-final-bill 4-6 days -30% -50% AR aging compresses. Cash velocity improves.
Coder productivity (charts/FTE/day) 40-55 +40% +90% Workflow lift. Volume per FTE roughly doubles by day 180.
RAF capture (HCC books) Specialty-dependent +3-5% +5-8% Specificity prompts at coding time drive lift. Defensible.
Audit defensibility (sampled) Baseline pass rate ≥ baseline +5 pts Compliance risk does not increase. Structured override log helps.

Five mistakes to avoid.

Every failed implementation we have seen made at least one of these. Most made two. Reading them upfront is materially cheaper than learning them in production.

01

Skipping the baseline measurement.

The practice opens with "we know our numbers, let's just install the AI." Three months later the scorecard is empty and nobody can prove the implementation helped, because there is nothing to compare against. The vendor declares success. The CFO is unconvinced. The implementation gets quietly rolled back.

Real exampleA 60-provider medical group skipped baseline. At day 90 the AI vendor said denial rate "improved to 11 percent." The CFO had no idea what the pre-implementation rate was. The contract was not renewed.
02

Compressing the parallel run.

The parallel run is expensive. Two sets of coding cost more than one, and the daily reconciliation eats senior auditor time. Practices try to cut it from 30 days to 10. The 10-day parallel does not surface the AI's actual failure modes on real production patterns. Cutover lands on a brittle workflow that breaks in the first month.

Real exampleA behavioral health practice compressed parallel to 7 days. Within 2 weeks of cutover the AI was missing modifier 95 on telehealth claims at scale. Three weeks of payer denials before the rule pack got tuned. $86K in delayed cash.
03

Treating AI as a black box without rule-pack governance.

The vendor ships the model, the practice installs it, nobody owns the rule pack. Payer policies change. The rule pack does not. Denials climb. The coding lead does not know how to file a rule-pack update. The vendor does not know the practice's payer mix matters. Three months later the implementation is underperforming for entirely fixable reasons.

Real exampleA multi-specialty group treated AI as install-and-forget. A commercial payer updated its prior-auth requirement on imaging. Eight weeks of denials before anyone connected the dots. The rule update was a 20-minute fix.
04

Picking AI on demo accuracy rather than the specialty rule library.

Every AI coding vendor demos at 95 percent accuracy. The number is real and the number is misleading. The accuracy that matters is on your specialty, your payer mix, your charts. Vendors with strong demos and weak specialty rule libraries miss on the codes that drive your dollars. The selection criterion should be the rule library depth, not the demo number.

Real exampleAn ABA practice picked an AI on a 96 percent demo accuracy. The vendor's ABA rule pack had no authorization-tracking logic. The practice had to build it themselves at 14 weeks of internal engineering cost.
05

Cutting over without a hyper-care window.

Cutover lands on a Friday. The next 30 days are nominally "post-implementation." Nobody is watching the AI output in real time. By the time the first weekly scorecard surfaces a problem, three weeks of claims have shipped on broken output. The fix is operationally painful and politically expensive.

Real exampleA hospital outpatient group skipped hyper-care. A logic gap in inpatient observation coding shipped 4 weeks of incorrect claims. $340K in rework, plus a downstream audit risk that took 3 months to close out.

Day 91 and beyondFrom project to operating cadence.

Day 90 is not the end of the implementation. It is the transition from project governance to operating cadence. The project lead steps back. The coding lead takes ongoing ownership of the rule pack. The senior auditor moves from daily reconciliation to a 10 percent sampling cadence. The CFO scorecard shifts from weekly to monthly. The implementation becomes part of how the practice runs, not something the practice is doing alongside its normal operations.

The single most important post-day-90 commitment is the quarterly rule-pack refresh. Payer policies will continue to update. CPT additions will continue to land annually. State Medicaid rules will continue to drift. The rule pack that was tuned on day 30 will be stale by day 180 unless it is maintained. Build the maintenance cadence before day 90 ends, while the implementation muscle is still strong, not three months later when the team has moved on.

The honest summaryThe lift is real. The discipline is the cost.

Across the implementations we have run, the dollar lift over the 12 months after cutover lands between 4 and 9 percent of net revenue on a representative book. The lift comes from denial reduction, faster cash velocity, productivity gains, and (for HCC books) RAF capture. The implementations that hit the high end of the range followed this playbook. The implementations that landed at the low end skipped one of the four phases. The implementations that failed skipped two.

The discipline is the cost. Baseline measurement is unglamorous. Parallel run is expensive. Rule-pack governance is a permanent role, not a one-time setup. Hyper-care eats senior auditor time. CFO scorecard reviews are recurring meetings. Practices that view this playbook as overhead are reading it backward. The lift comes from the discipline. There is no version of AI coding implementation where the lift shows up without the discipline behind it.

AI coding implementation, frequently asked questions.

The questions coding leads and CFOs most often ask when scoping or running an AI coding implementation.

How long does an AI coding implementation actually take?
A well-run AI coding implementation runs 90 days from kickoff to fully operational cutover. Days 0 to 7 are baseline measurement. Days 7 to 30 are sandbox build and rule-pack tuning. Days 30 to 60 are parallel run where AI codes alongside the human coder and outputs are reconciled. Days 60 to 90 are cutover and hyper-care. Practices that try to compress this to 45 days consistently miss baseline measurement or skip the parallel run; those are the two phases that are non-negotiable.
What baseline data do we need before starting?
Six months of prior coding output is the minimum: chart, coded ICD-10, coded CPT, modifier set, RVU, charge amount, and denial outcome. Plus a payer mix breakdown, a top-50 CPT code volume list, a denial reason breakdown by code, and a current coder productivity baseline (charts per day, accuracy rate, time per chart). The baseline is what makes the success metrics scorecard meaningful; without it the implementation is operating in the dark.
What is the parallel run phase and why is it non-negotiable?
Parallel run is the 30-day window where AI codes every chart and the human coder also codes every chart, with a senior auditor reconciling differences daily. The parallel run is non-negotiable because it surfaces three things: the AI's actual accuracy on this practice's specific clinical documentation patterns, the human coder's actual baseline (often different from claimed baseline), and the categories where AI is and is not yet trusted. Without parallel run, cutover is a blind risk.
What are the success metrics for AI coding?
The scorecard covers coding accuracy versus chart audit, denial rate change versus pre-implementation baseline, time-to-final-bill, coder productivity (charts per day per FTE), RAF capture for HCC books, and audit defensibility. The implementation should not be considered successful until coding accuracy is at or above pre-implementation, denial rate is at or below pre-implementation, and time-to-final-bill is materially better. Any metric that moves the wrong direction triggers a structured root-cause investigation.
What roles need training and how much?
Three roles, four hours each. Coders need workflow training on how to review AI output, when to override, and how to document overrides. Coding leads need rule-pack governance training on how to file rule updates, how to read AI confidence scores, and how to run the daily reconciliation. CFO and revenue cycle leadership need scorecard interpretation training on what the metrics mean and what action to take when they move. None of the three roles can be skipped.
What are the five mistakes that doom AI coding implementations?
First, skipping baseline measurement. Second, compressing the parallel run. Third, treating AI as a black box and not building rule-pack governance. Fourth, picking AI vendors based on demo accuracy rather than on the specialty rule library. Fifth, cutting over without a hyper-care window. Each of these mistakes has a story behind it; the playbook lays them out so you can avoid them rather than learn them the expensive way.
How does AI coding affect coder roles?
Coder roles shift from data entry to clinical-billing interface. Volume per FTE roughly doubles. The work itself becomes harder and more interesting; coders spend their time on the cases that need clinical judgment, modifier disputes, and documentation queries to physicians rather than on the straightforward charts. Coder retention typically improves post-implementation because the role becomes more skilled. Headcount reduction is usually achieved by attrition rather than layoff.
What does day 91 look like once the implementation is done?
Daily reconciliation continues but moves to a senior auditor sampling 10 percent of AI-coded charts. Weekly scorecard review with coding lead and CFO. Monthly rule-pack review against published payer policy updates and quarterly rule library refresh. Quarterly coder calibration session. The implementation does not end at day 90; it transitions from project to operating cadence.

Want a 90-day implementation plan on your data?

A free 30-day implementation audit under a same-day BAA. The output is a written report covering your baseline data readiness, recommended rule-pack scope, expected dollar lift on your payer mix, named risks, and a week-by-week project plan. A senior partner on the call.