← Back to repo
Interactive demo · real evaluation data

When your server breaks,
this AI tells you what broke, why, and how to fix it — in minutes.

Think of it as an AI on-call detective: a monitoring alert fires, the agent investigates metrics, logs and runbooks — then hands you a verdict with proof you can click.

9/9fault scenarios diagnosed
100%evidence grounded — zero fabrication
7/7unit tests green
gpt-4o-miniAI Model
CLICK ONE
▼
and let the AI do the rest

What just happened?

The same pipeline runs every time an alert fires in production.

STEP 1

An alert arrives

ProviderPrometheus
RouterAlertmanager
Deliverywebhook → agent
STEP 2

The agent investigates

MetricsPrometheus
LogsLoki (tool call)
Knowledgerunbooks (RAG)
STEP 3

A verdict, with proof

Outputroot cause + actions
Confidence0.8–0.9, calibrated
Evidencereal queries + links