Kubernetes incident investigation
The evidence-first Kubernetes incident investigator that never touches your cluster and gets smarter after every incident.
Keryx's only write surface is Git PRs that a human reviews and merges; the cluster changes only through your existing GitOps pipeline.
Apache-2.0. Container images are not public yet — there is nothing to helm install today, and we would rather say so than ship a quickstart that fails on the first pull.
What is actually different
In-cluster, read-only, BYO-LLM, Alertmanager and Slack — several projects claim that sentence, and one of them is a CNCF Sandbox project. The tagline is not the differentiator. These four are what a source audit of the leading open-source incumbents found unclaimed.
Per-claim citations, verified by re-execution
Every claim in a finding carries the ID of the tool call that produced it, and core-api rejects a finding that cites a step which does not exist or does not contain the quoted excerpt. Citations are structural, not a thing the prompt asks for politely.
Re-execution is designed and specified; today the honesty axis is a human spot-check that each cited excerpt supports its claim (evals.md §7).
A conviction ladder, not a confidence score
Findings publish at a named rung — speculation, pattern match, supported by context, validated by system state, alternatives ruled out — computed in code from the structure of the evidence. The model may argue a rung down, never up.
Conviction ships as presentation. Rung accuracy measured against outcomes is future work, and no accuracy claim is made from rungs until it is measured.
"Insufficient evidence" is a valid answer
Two of the ten gate scenarios are controls: alerts engineered so no cause is discoverable from where Keryx is allowed to look. The only passing answer is a variant of “I cannot see the cause.” Confabulating a plausible local explanation fails the run.
Deploy correlation through the Flux reconcile chain
The triage sweep always asks what actually rolled out in the alert window — Flux revisions and release tags back to the GitHub PR — instead of leaving the model to guess whether a deploy was involved.
The release gate needs 7 of 10. It scores 6.
Ten fault-injected Kubernetes failures with known root causes, each run three times against a real EKS cluster where Prometheus notices and Alertmanager fires naturally. A human scores every run against the stored trace. We are shipping at six and publishing the failures, because a benchmark you can only read when it flatters you measures nothing.
Security posture
Read-only enforcement is the primary claim, so it is the one with the most written down about it — including what is not yet proven.
Read the threat model summaryPricing
The OSS tier is free forever and full-featured on a single cluster. When the hosted control plane ships, pricing will be per-cluster or per-seat, published on the site, with no “call us.”