Reliability

What we commit to when a managed workflow fails.

Real commitments, not adjectives. Every one of these is measured against actual runs, and the SLA credits below trigger automatically when a threshold is breached.

What does Phronimos commit to for detection speed?

For every workflow under a Fractional AI Officer retainer, Phronimos aims to detect a silent failure within a target window measured from the moment the failure occurs to the moment an alert fires on our side. Detection targets vary by workflow severity tier.

SeverityDetection targetDefinition
Critical< 15 minutesCustomer-facing workflow, revenue-impacting, or compliance-adjacent.
High< 2 hoursInternal workflow with downstream dependencies; delay creates rework.
Medium< 24 hoursIsolated internal workflow; delay costs time but not customers.

Who gets paged when detection fires?

On detection, Phronimos pages an on-call operator through the primary channel (Slack, SMS, or email — client picks at onboarding). The client's designated contact is looped in within the same window if we can't confirm the failure is contained.

What response times does Phronimos target?

Response time = detection to first mitigating action. Not the same as resolution time (which depends on root cause and often on a third party). Response is what we control.

SeverityResponse target
Critical< 30 minutes
High< 4 business hours
Medium< 1 business day

What happens when an agent fails?

Every managed workflow ships with a rollback plan, a documented human takeover procedure, and a freeze switch. On failure, Phronimos executes the least-disruptive available response in that order.

  1. Rollback to the last known-good state if the workflow has a checkpoint. Preferred response — the client's downstream state stays consistent.
  2. Human takeover by a Phronimos operator or the client's designated user, following the workflow's documented takeover procedure. Applied when rollback isn't possible or wouldn't help.
  3. Freeze — the workflow is paused, the trigger disabled, and the client is notified with a plain-English summary of what stopped and what didn't. Applied when we're not yet sure the failure is contained.

What are the SLA credits when Phronimos misses these commitments?

When we breach a detection or response target on a Critical or High severity workflow, the client's next retainer invoice is credited automatically. Credits are proportional to the miss, capped at one month of retainer, and don't require the client to file anything.

BreachCredit
Critical detection miss (> 15 min)10% of next monthly retainer
Critical response miss (> 30 min)10% of next monthly retainer
High detection miss (> 2 hours)5% of next monthly retainer
High response miss (> 4 business hours)5% of next monthly retainer
Any combined miss on a single incidentStacks to a per-incident cap of 25%
Monthly cap100% of that month's retainer

What isn't covered by these commitments?

The Agent Reliability Review (the $999 diagnostic) is not covered by these SLAs — it's a one-time review, not a managed engagement. Implementation Sprint deliverables come with a 30-day quality warranty on the specific workflow shipped. Everything else — third-party outages, client-side auth changes, changes we weren't told about — is triaged the same way but not counted as a breach.

See a redacted example of how a real incident gets reported: Sample incident report.