Keryx

Kubernetes incident investigation

The evidence-first Kubernetes incident investigator that never touches your cluster and gets smarter after every incident.

Keryx's only write surface is Git PRs that a human reviews and merges; the cluster changes only through your existing GitOps pipeline.

Apache-2.0. Container images are not public yet — there is nothing to helm install today, and we would rather say so than ship a quickstart that fails on the first pull.

What is actually different

In-cluster, read-only, BYO-LLM, Alertmanager and Slack — several projects claim that sentence, and one of them is a CNCF Sandbox project. The tagline is not the differentiator. These four are what a source audit of the leading open-source incumbents found unclaimed.

  • Per-claim citations, verified by re-execution

    Every claim in a finding carries the ID of the tool call that produced it, and core-api rejects a finding that cites a step which does not exist or does not contain the quoted excerpt. Citations are structural, not a thing the prompt asks for politely.

    Re-execution is designed and specified; today the honesty axis is a human spot-check that each cited excerpt supports its claim (evals.md §7).

  • A conviction ladder, not a confidence score

    Findings publish at a named rung — speculation, pattern match, supported by context, validated by system state, alternatives ruled out — computed in code from the structure of the evidence. The model may argue a rung down, never up.

    Conviction ships as presentation. Rung accuracy measured against outcomes is future work, and no accuracy claim is made from rungs until it is measured.

  • "Insufficient evidence" is a valid answer

    Two of the ten gate scenarios are controls: alerts engineered so no cause is discoverable from where Keryx is allowed to look. The only passing answer is a variant of “I cannot see the cause.” Confabulating a plausible local explanation fails the run.

  • Deploy correlation through the Flux reconcile chain

    The triage sweep always asks what actually rolled out in the alert window — Flux revisions and release tags back to the GitHub PR — instead of leaving the model to guess whether a deploy was involved.

The release gate needs 7 of 10. It scores 6.

Ten fault-injected Kubernetes failures with known root causes, each run three times against a real EKS cluster where Prometheus notices and Alertmanager fires naturally. A human scores every run against the stored trace. We are shipping at six and publishing the failures, because a benchmark you can only read when it flatters you measures nothing.

Security posture

Read-only enforcement is the primary claim, so it is the one with the most written down about it — including what is not yet proven.

Read the threat model summary

Pricing

The OSS tier is free forever and full-featured on a single cluster. When the hosted control plane ships, pricing will be per-cluster or per-seat, published on the site, with no “call us.”