99.9% uptime SLA, in writing.
An SLA matters only when it is enforceable. Ours is appended to every engagement contract: 99.9 percent uptime floor on every production service, 12-month measured uptime of 99.97 percent, quarterly DR drills with documented RTO/RPO, daily encrypted backups with monthly restore tests, 24/7 monitoring with under-5-minute mean time to acknowledge, and a public status page that stays live even when production is degraded. Every commitment maps to a control. Every control maps to an audit trail.
Every commitment maps to a control.
The SLA matrix to the right is the actual matrix appended to every engagement contract. Each row defines the commitment in writing, the measurement methodology, the service credit on breach, and the audit-trail evidence. Uptime is measured per minute. P95 latency is measured per request. Ticket response is measured from first acknowledgement. Incident mean time to resolution is measured from page to recovery.
The matrix is not a marketing artifact. It is a SOC 2 control surface, a HIPAA technical-safeguard evidence source, and a contractual commitment. The same matrix anchors the quarterly business review and feeds the client director’s standing scorecard.
Seven controls. From uptime to incident.
Each control is a production engineering investment, not a marketing claim. Every control has a documented runbook, a scheduled exercise cadence, and an evidence chain. The SLA is what the controls add up to.
99.9% floor, in writing.
Measured per minute per covered service. 12-month measured uptime is 99.97 percent. The floor is the contractual commitment with service credits on breach. We over-engineer to the measured number because the floor allows 43 minutes per month and the measured number is 13.
Quarterly, documented.
Every calendar quarter: full failover exercise against the documented runbook. RTO and RPO measured against the test. Post-drill report shared with clients under NDA. Runbook updated with any deviation. The DR is real; the drill is the test that proves it.
1 hour / 15 min.
RTO of 1 hour on production-critical services. RPO of 15 minutes for transactional data. Point-in-time recovery on the database tier across the 35-day retention window. Both numbers tested quarterly. Both numbers documented in the SLA matrix.
Daily full, monthly restore test.
Daily full database backup encrypted with AES-256-GCM and rotating per-backup keys. 35-day retention. Cross-region replication. Monthly restore test verifies the backup is actually restorable. Restore test is part of the runbook and the result is logged.
Public · 30-day uptime.
Component-level status per production service. 30-day uptime grid per component. Incident history with detailed post-mortem. Scheduled maintenance posted 7 days in advance. Hosted on an independent surface so it stays up when production is degraded. Subscribed clients get real-time notification.
Synthetic + RUM, under 5 min MTTA.
Synthetic transactions every 60 seconds against every production surface. Real user monitoring on every client session. Alert paging fires within 90 seconds of confirmed failure. Paired on-call rotation, primary plus secondary, covers every minute. Mean time to acknowledge under 5 minutes.
Version-controlled per failure mode.
Documented runbook per failure mode: database failover, region failover, integration outage, partial degradation, DDoS, security incident. Named primary owner per runbook. Documented escalation path. Post-incident lessons feed back to the runbook within one week. Runbook exercised every quarter via DR drill and via real incidents.
Four phases. Page to post-mortem.
When an incident fires, this is the workflow. Every phase is a defined gate with a named owner, a measured duration target, and a logged evidence chain. The MTTR number is the sum of these phases against the active book.
Detect.
Synthetic transaction or RUM signal trips. Alert paging fires within 90 seconds of confirmed failure. Primary on-call acknowledges within 5 minutes. Detection-to-acknowledge logged for every incident. Mean time to acknowledge under 5 minutes against the active book.
Triage.
On-call opens the incident channel and pulls the runbook for the matching failure mode. Severity classified (P1 production critical, P2 partial degradation, P3/P4 lower urgency). Secondary on-call joins on P1/P2. Status page updated within 10 minutes of confirmed P1/P2.
Resolve.
Runbook execution: database failover, region failover, integration restore, mitigation deploy, or rollback as the failure mode requires. Recovery confirmed via synthetic transaction and RUM signal. Mean time to resolution under 35 minutes for production-critical incidents against the active book.
Post-mortem.
Blameless post-mortem within 5 business days. Root cause documented, contributing factors mapped, remediation items assigned with owners and due dates. Runbook updated within 1 week. Status page incident detail updated. Affected clients receive direct communication with the post-mortem summary.
Measured outcomes from the SLA matrix.
Across the active book, last 12 calendar months. Numbers are what we measure, not what we hope.
Frequently asked questions: SLA & Uptime.
What does the 99.9% uptime SLA actually cover?
What is the difference between 99.9% SLA and 99.97% measured?
What are RTO and RPO?
What does the quarterly DR drill look like?
What backups run and how are they protected?
What does the status page show?
What does 24/7 monitoring actually mean?
Where does the incident runbook live and what does it cover?
SLA compounds with the stack.
An SLA is only as good as the services it protects. The credentialing platform, the AR engine, and the 14-day implementation are the three production guarantees that the SLA matrix backstops.
Credentialing under SLA.
HIPAA-hardened multi-tenant credentialing with AES-256-GCM PHI-at-rest encryption, RS256 passports, security headers, audited reveal, and the phi_access_log audit trail. Credential OS is on the 99.9 percent SLA and the quarterly DR drill scope.
AR engine backstopped.
AR triage by recoverable dollars on the same 99.9 percent SLA. Specialist worklists never go cold on a deploy or an incident. Paired on-call and runbook-driven response keep the engine running through the worst hour of the worst day.
14-day to production SLA.
The 14-day implementation hands directly into the production SLA. Day 14 go-live triggers the 99.9 percent uptime commitment. Hyper-care plus standard SLA cover the first 60 days post go-live with zero gap.
Get the SLA matrix. Read before you sign.
The full production SLA matrix: uptime floor, P95 latency budget, ticket response targets, incident MTTR, RTO/RPO, service credit schedule, DR drill cadence, status page link, evidence chain for SOC 2 and HIPAA. Six pages. Read it before you sign anything. A senior partner on the walkthrough call.