The release gate needs 7 of 10. It scores 6.
Ten fault-injected Kubernetes failures with known root causes, each run three times against a real EKS cluster where Prometheus notices and Alertmanager fires naturally. A human scores every run against the stored trace. We are shipping at six and publishing the failures, because a benchmark you can only read when it flatters you measures nothing.
What this measures that a pass rate does not
- Evidence faithfulness gates the run — a right answer with fabricated citations fails.
- A conviction ladder instead of a confidence number.
- Controls that pass on an admission of ignorance — 20% of the gate.
- Unrun scenarios published with their blockers, not omitted.
- Per-model numbers, stated at the top of every artifact.
What this number is not
- One model (deepseek.v3-v1:0 via Bedrock Converse). These are this model's results on Keryx, not “Keryx's accuracy”.
- Three runs per scenario is the gate floor; published benchmarks use five to ten.
- Human-scored by the maintainer. An automated judge is future work.
- One cluster topology, one demo application.
- The prompt-injection cases are specified and have not been run.