What we commit to when a managed workflow fails.
Real commitments, not adjectives. Every one of these is measured against actual runs, and the SLA credits below trigger automatically when a threshold is breached.
What does Phronimos commit to for detection speed?
For every workflow under a Fractional AI Officer retainer, Phronimos aims to detect a silent failure within a target window measured from the moment the failure occurs to the moment an alert fires on our side. Detection targets vary by workflow severity tier.
| Severity | Detection target | Definition |
|---|---|---|
| Critical | < 15 minutes | Customer-facing workflow, revenue-impacting, or compliance-adjacent. |
| High | < 2 hours | Internal workflow with downstream dependencies; delay creates rework. |
| Medium | < 24 hours | Isolated internal workflow; delay costs time but not customers. |
Who gets paged when detection fires?
On detection, Phronimos pages an on-call operator through the primary channel (Slack, SMS, or email — client picks at onboarding). The client's designated contact is looped in within the same window if we can't confirm the failure is contained.
- Critical: on-call paged within 5 minutes of detection; client contact looped in within 30 minutes if unresolved.
- High: on-call paged within 30 minutes; client contact notified within 4 hours if unresolved.
- Medium: on-call notified next business hour; client contact notified in the next weekly reliability digest.
What response times does Phronimos target?
Response time = detection to first mitigating action. Not the same as resolution time (which depends on root cause and often on a third party). Response is what we control.
| Severity | Response target |
|---|---|
| Critical | < 30 minutes |
| High | < 4 business hours |
| Medium | < 1 business day |
What happens when an agent fails?
Every managed workflow ships with a rollback plan, a documented human takeover procedure, and a freeze switch. On failure, Phronimos executes the least-disruptive available response in that order.
- Rollback to the last known-good state if the workflow has a checkpoint. Preferred response — the client's downstream state stays consistent.
- Human takeover by a Phronimos operator or the client's designated user, following the workflow's documented takeover procedure. Applied when rollback isn't possible or wouldn't help.
- Freeze — the workflow is paused, the trigger disabled, and the client is notified with a plain-English summary of what stopped and what didn't. Applied when we're not yet sure the failure is contained.
What are the SLA credits when Phronimos misses these commitments?
When we breach a detection or response target on a Critical or High severity workflow, the client's next retainer invoice is credited automatically. Credits are proportional to the miss, capped at one month of retainer, and don't require the client to file anything.
| Breach | Credit |
|---|---|
| Critical detection miss (> 15 min) | 10% of next monthly retainer |
| Critical response miss (> 30 min) | 10% of next monthly retainer |
| High detection miss (> 2 hours) | 5% of next monthly retainer |
| High response miss (> 4 business hours) | 5% of next monthly retainer |
| Any combined miss on a single incident | Stacks to a per-incident cap of 25% |
| Monthly cap | 100% of that month's retainer |
What isn't covered by these commitments?
The Agent Reliability Review (the $999 diagnostic) is not covered by these SLAs — it's a one-time review, not a managed engagement. Implementation Sprint deliverables come with a 30-day quality warranty on the specific workflow shipped. Everything else — third-party outages, client-side auth changes, changes we weren't told about — is triaged the same way but not counted as a breach.
See a redacted example of how a real incident gets reported: Sample incident report.