Runbook: chaos p99 latency breach
Alert source: packages/testkit/alerts/chaos-assertions.rules.yaml
Canonical markdown: docs/runbooks/chaos-p99-latency-breach.md
Severity: critical (fires only during a chaos run).
Symptom
During the current chaos run, p99 channel.outbound.latency_ms
exceeds the 10s SLO.
Likely cause
- A
partition-channelfault longer than the adapter's retry budget pushed queued sends past the SLO. - The bus backpressure listener paused a channel adapter and queued sends didn't drain fast enough on resume.
- An upstream platform's own latency happened to coincide with the fault window.
Immediate mitigation
Stop the chaos run and review assertions:
# The chaos harness is the testkit chaos driver (packages/testkit/src/chaos)
# — stop the run by interrupting the harness process (Ctrl-C / SIGINT);
# the driver's stop() emits the run report on the way out.
Root-cause investigation
Correlate the chaos report's fault timeline with the latency histogram samples. A p99 breach that started within 60s of a fault fire is almost always caused by that fault.
Post-incident
- Capture: fault → breach timeline, recovery duration.
- Close when: a targeted test reproduces the breach and a fix stabilizes p99 under the same fault.