Outside-in (Uptime Kuma)
Synthetic probes check public reachability, TLS, and expected content. A failure triggers the shared triage pipeline.
Interactive Simulation
Incidents can begin through two verified signal paths—an external synthetic failure from Uptime Kuma, or an internal infrastructure alert from Alertmanager. Both join a common normalized flow: deduplication, HARP evidence retrieval, Nero-Camp task creation, operator notification, and human closeout.
Synthetic probes check public reachability, TLS, and expected content. A failure triggers the shared triage pipeline.
Prometheus rules fire on cluster and infrastructure thresholds—disk pressure, pod crashes, node health. Alerts normalize into the same downstream flow.
"Notice that every AI-facing step is constrained. HARP retrieves evidence and safe diagnostics. Nero tracks the decision. LiteLLM can summarize in the wider architecture, but it does not replace citations. No mutation occurs automatically—the human records the result, and the RCA becomes future evidence."
Why This Works
Uptime Kuma or Alertmanager produces a signal.
Synthetic Monitoring gateway deduplicates and maps to a shared contract.
HARP finds prior context from runbooks and RCAs.
Evidence pack contains citations, snippets, confidence, and gaps.
Nero-Camp records human decisions. No mutation occurs automatically.
Outcome becomes an RCA, runbook update, or session log for future retrieval.