Keryx

The release gate needs 7 of 10. It scores 6.

Ten fault-injected Kubernetes failures with known root causes, each run three times against a real EKS cluster where Prometheus notices and Alertmanager fires naturally. A human scores every run against the stored trace. We are shipping at six and publishing the failures, because a benchmark you can only read when it flatters you measures nothing.

What this measures that a pass rate does not

  • Evidence faithfulness gates the run — a right answer with fabricated citations fails.
  • A conviction ladder instead of a confidence number.
  • Controls that pass on an admission of ignorance — 20% of the gate.
  • Unrun scenarios published with their blockers, not omitted.
  • Per-model numbers, stated at the top of every artifact.

What this number is not

  • One model (deepseek.v3-v1:0 via Bedrock Converse). These are this model's results on Keryx, not “Keryx's accuracy”.
  • Three runs per scenario is the gate floor; published benchmarks use five to ten.
  • Human-scored by the maintainer. An automated judge is future work.
  • One cluster topology, one demo application.
  • The prompt-injection cases are specified and have not been run.