Home/Technology/SLA & Uptime
99.9% in writing · 99.97% measured · quarterly DR

99.9% uptime SLA, in writing.

An SLA matters only when it is enforceable. Ours is appended to every engagement contract: 99.9 percent uptime floor on every production service, 12-month measured uptime of 99.97 percent, quarterly DR drills with documented RTO/RPO, daily encrypted backups with monthly restore tests, 24/7 monitoring with under-5-minute mean time to acknowledge, and a public status page that stays live even when production is degraded. Every commitment maps to a control. Every control maps to an audit trail.

99.9% SLA floor in writing 99.97% measured 12-month Quarterly DR drill documented
The SLA matrix

Every commitment maps to a control.

The SLA matrix to the right is the actual matrix appended to every engagement contract. Each row defines the commitment in writing, the measurement methodology, the service credit on breach, and the audit-trail evidence. Uptime is measured per minute. P95 latency is measured per request. Ticket response is measured from first acknowledgement. Incident mean time to resolution is measured from page to recovery.

The matrix is not a marketing artifact. It is a SOC 2 control surface, a HIPAA technical-safeguard evidence source, and a contractual commitment. The same matrix anchors the quarterly business review and feeds the client director’s standing scorecard.

Production SLA matrix · written, measured, credited
99.9%
Production uptime floorPer calendar month, per covered service, measured per minute
<400ms
P95 API latency budgetPer request, real user monitoring, alert on breach
<4 hrs
Ticket response SLABusiness-hour acknowledge on P3/P4, 24/7 on P1/P2
<1 hr
Incident MTTR (P1)Page to recovery on production-critical services
Service credits on breach · quarterly scorecard · SOC 2 + HIPAA evidence
The capabilities

Seven controls. From uptime to incident.

Each control is a production engineering investment, not a marketing claim. Every control has a documented runbook, a scheduled exercise cadence, and an evidence chain. The SLA is what the controls add up to.

01 · Uptime SLA

99.9% floor, in writing.

Measured per minute per covered service. 12-month measured uptime is 99.97 percent. The floor is the contractual commitment with service credits on breach. We over-engineer to the measured number because the floor allows 43 minutes per month and the measured number is 13.

02 · DR drills

Quarterly, documented.

Every calendar quarter: full failover exercise against the documented runbook. RTO and RPO measured against the test. Post-drill report shared with clients under NDA. Runbook updated with any deviation. The DR is real; the drill is the test that proves it.

03 · RTO & RPO

1 hour / 15 min.

RTO of 1 hour on production-critical services. RPO of 15 minutes for transactional data. Point-in-time recovery on the database tier across the 35-day retention window. Both numbers tested quarterly. Both numbers documented in the SLA matrix.

04 · Encrypted backups

Daily full, monthly restore test.

Daily full database backup encrypted with AES-256-GCM and rotating per-backup keys. 35-day retention. Cross-region replication. Monthly restore test verifies the backup is actually restorable. Restore test is part of the runbook and the result is logged.

05 · Status page

Public · 30-day uptime.

Component-level status per production service. 30-day uptime grid per component. Incident history with detailed post-mortem. Scheduled maintenance posted 7 days in advance. Hosted on an independent surface so it stays up when production is degraded. Subscribed clients get real-time notification.

06 · 24/7 monitoring

Synthetic + RUM, under 5 min MTTA.

Synthetic transactions every 60 seconds against every production surface. Real user monitoring on every client session. Alert paging fires within 90 seconds of confirmed failure. Paired on-call rotation, primary plus secondary, covers every minute. Mean time to acknowledge under 5 minutes.

07 · Incident runbook

Version-controlled per failure mode.

Documented runbook per failure mode: database failover, region failover, integration outage, partial degradation, DDoS, security incident. Named primary owner per runbook. Documented escalation path. Post-incident lessons feed back to the runbook within one week. Runbook exercised every quarter via DR drill and via real incidents.

How incidents run

Four phases. Page to post-mortem.

When an incident fires, this is the workflow. Every phase is a defined gate with a named owner, a measured duration target, and a logged evidence chain. The MTTR number is the sum of these phases against the active book.

Phase 01

Detect.

Synthetic transaction or RUM signal trips. Alert paging fires within 90 seconds of confirmed failure. Primary on-call acknowledges within 5 minutes. Detection-to-acknowledge logged for every incident. Mean time to acknowledge under 5 minutes against the active book.

Phase 02

Triage.

On-call opens the incident channel and pulls the runbook for the matching failure mode. Severity classified (P1 production critical, P2 partial degradation, P3/P4 lower urgency). Secondary on-call joins on P1/P2. Status page updated within 10 minutes of confirmed P1/P2.

Phase 03

Resolve.

Runbook execution: database failover, region failover, integration restore, mitigation deploy, or rollback as the failure mode requires. Recovery confirmed via synthetic transaction and RUM signal. Mean time to resolution under 35 minutes for production-critical incidents against the active book.

Phase 04

Post-mortem.

Blameless post-mortem within 5 business days. Root cause documented, contributing factors mapped, remediation items assigned with owners and due dates. Runbook updated within 1 week. Status page incident detail updated. Affected clients receive direct communication with the post-mortem summary.

What clients see

Measured outcomes from the SLA matrix.

Across the active book, last 12 calendar months. Numbers are what we measure, not what we hope.

99.97%
Measured 12-month uptime
Across every covered production service: AR engine, Credential OS, eligibility, discovery, authorization, dashboard. SLA floor is 99.9 percent. The 99.97 measured number is what we engineer to. The gap between the two is intentional headroom for the once-a-quarter unexpected.
35min
P1 incident MTTR
Page to recovery on production-critical incidents. Achieved through paired on-call coverage, runbook-driven response, and pre-staged failover paths. Mean time to acknowledge under 5 minutes feeds directly into the resolution number.
4
DR drills last 12 months
Every calendar quarter, executed against the documented runbook. RTO and RPO measured against the test. Post-drill report shared with clients under NDA. Runbook updated with deviations. The DR is real because we drill it.
Common questions

Frequently asked questions: SLA & Uptime.

What does the 99.9% uptime SLA actually cover?
The SLA covers the production surface every client touches: the AR workflow engine, Credential OS, the eligibility verification service, the discovery engine, the authorization tracker, and the client director dashboard. Uptime is measured as the percentage of minutes per calendar month during which every covered service is responsive within the documented P95 latency budget. 99.9 percent is the floor; the 12-month rolling measured number is 99.97 percent. The SLA matrix is appended to every engagement contract.
What is the difference between 99.9% SLA and 99.97% measured?
99.9 percent is the contractual floor: the number we owe you in writing. If we miss the floor in any calendar month, the service credit terms in the SLA matrix apply. 99.97 percent is the actual measured number across our active production over the last 12 months. We over-engineer to the measured number, not the contractual floor, because the contractual floor still allows roughly 43 minutes of downtime per month. The 99.97 measured number is roughly 13 minutes per month.
What are RTO and RPO?
Recovery Time Objective is how fast we restore service after a major incident. Recovery Point Objective is how much data loss the recovery can tolerate. Our documented RTO is one hour for production-critical services (AR engine, Credential OS, eligibility, discovery, authorization). Our documented RPO is 15 minutes for transactional data, backed by point-in-time recovery on the database tier. Both numbers are tested quarterly via DR drill and the test results are shared with clients on request.
What does the quarterly DR drill look like?
A documented disaster recovery exercise run every calendar quarter. The drill simulates a region-level outage on production. The runbook executes: traffic fails over to standby, database promotes from the warm replica, integration endpoints re-establish, monitoring confirms green. The drill is timed end-to-end against the documented RTO. A post-drill report documents the time per phase, any deviation from the runbook, and the remediation items. The report is available to clients under NDA.
What backups run and how are they protected?
Daily full encrypted backups on the database tier, with 35-day retention. Point-in-time recovery available across the retention window with 15-minute granularity. Backups are encrypted at rest using AES-256-GCM with rotating per-backup keys. Backups are stored in a region separate from the primary production region. A monthly restore test verifies that the backup is restorable; the restore test is part of the standard runbook and the result is logged.
What does the status page show?
Component-level uptime for every production service: AR engine, Credential OS, eligibility, discovery, authorization, dashboard. 30-day uptime per component. Incident history with detailed post-mortem on any incident exceeding the documented thresholds. Scheduled maintenance windows posted 7 days in advance. Subscribed clients receive real-time notification on any status-page change. The page is hosted on an independent surface so it stays up even if production is degraded.
What does 24/7 monitoring actually mean?
Synthetic transactions run every 60 seconds against every production service surface. Real user monitoring runs against every client session. Alert paging fires within 90 seconds of a confirmed failure. The on-call rotation covers every minute of every day across paired on-call engineers with primary and secondary coverage. Mean time to acknowledge is under 5 minutes; mean time to resolution against the active book is under 35 minutes for production-critical incidents.
Where does the incident runbook live and what does it cover?
The runbook is version-controlled internally and covers every documented failure mode: database failover, region failover, integration endpoint outage, partial-degradation scenarios, DDoS, security incident. Every runbook has a named primary owner, a defined success criterion, a documented escalation path, and a post-incident communication template. The runbook is exercised through the quarterly DR drill and through real production incidents. Lessons learned feed back to the runbook within one week of every post-mortem.

Get the SLA matrix. Read before you sign.

The full production SLA matrix: uptime floor, P95 latency budget, ticket response targets, incident MTTR, RTO/RPO, service credit schedule, DR drill cadence, status page link, evidence chain for SOC 2 and HIPAA. Six pages. Read it before you sign anything. A senior partner on the walkthrough call.