Diagnostics
One endpoint that answers: "what is every subsystem of this process doing right now?"
curl localhost:3000/dashboard/api/diagnostics
curl localhost:3000/dashboard/api/diagnostics?subsystem=events {
"total": 14,
"subsystems": [
{
"subsystem": "events",
"status": "ok",
"summary": "3 event types, 7 listeners, 1204 dispatches, avg 0.12 ms",
"detail": {"event_types": 3, "listeners": 7, "avg_ms": 0.12}
},
{
"subsystem": "compression",
"status": "off",
"summary": "not installed",
"detail": {"reason": "call router.Use(middleware.CompressionMiddleware())"}
}
]
} Why it is a leaf package
diag imports nothing but the standard library, and that is a requirement,
not a virtue. The root breeze package imports binding, rpc, scalar and internal/mcp; dashboard imports breeze, events, observability and video; fleet imports dashboard. A registry living in breeze could not be used by events or workflow, which sit below it; a registry
in dashboard could not be used by almost anything. So the registry is a
leaf, importable from every layer including the lowest.
Zero cost
Nothing here runs on a request path. Registration happens once, while an
application is being wired, and costs one slice append. Reading happens
only when someone asks, and that is the only time a probe function is
invoked. The registry is copy-on-write behind an atomic pointer, so Snapshot is a single atomic load with no lock — chosen not for
throughput (registration is rare, reads rarer) but so reading diagnostics
can never contend with anything, including another reader, on a process
already in trouble.
The four statuses
| Status | Means | A reader should |
|---|---|---|
ok | running and healthy | nothing |
degraded | running and unhappy | read the summary — it names what is wrong |
off | not installed, or deliberately disabled | nothing — not a fault |
unknown | registered but unable to answer; a probe that panicked | suspect the probe, not the subsystem |
off is separate from degraded because a dashboard rendering "compression
is not installed" in red would train its reader to ignore red. An empty
status is normalised to unknown rather than trusted, so a probe that
forgets to set one cannot be misread as healthy.
The registry keys
Fourteen subsystems register, and the key is the same string used on the
command line wherever one exists (router, workerpool, templates, i18n, websocket, auto-mcp, static, events, workflow, observability, dashboard, fleet, video, docs, jsonrpc, migrate, mcp, oauth2, plus every middleware). video registers twice
— once under the shared key, once under video:<prefix> — because a
multi-mount process needs the per-mount answer.
The endpoint that reports on itself
The mcp probe is the one an agent needs most, because breeze_diagnose_service reads this same registry — so the endpoint
serving the call was previously the one subsystem missing from its own
report. It surfaces three states no error anywhere else reveals: a scope
withholding a tool, generator mode running with no source tree, and AllowWorkspaceTools left on in production.
Writing a probe
func (t *Tracer) probe() diag.Report {
if !t.cfg.Enabled {
return diag.Off("tracing is disabled; set TracerConfig.Enabled")
}
exported, failed := t.exported.Load(), t.failed.Load()
detail := map[string]any{
"service": t.cfg.ServiceName,
"exported": exported,
"failures": failed,
}
if failed > 0 {
return diag.Degraded(
fmt.Sprintf("%d spans exported, %d export failures", exported, failed),
detail,
).WithNotes("the aggregator was unreachable at " + t.cfg.AggregatorURL)
}
return diag.OK(fmt.Sprintf("%d spans exported", exported), detail)
}
func (t *Tracer) registerDiagnostics() {
diag.Register("fleet", t.probe)
} Rules the built-in probes hold to:
- The summary names numbers, not the status — "412 spans exported, 3 export failures" beats "tracer is degraded".
Offnames the call that would turn it on.- Notes record what the report could not determine — a reader who sees no note is entitled to treat the numbers as complete.
Detailholds JSON-encodable scalars, slices and maps only — never a live handle.- Answer from state you already hold — a probe must be cheap and non-blocking; it must never dial, read a file, or take a lock that a request path uses.
Register replaces rather than appends, so installing the dashboard twice
produces one report, not three. Registering a nil probe unregisters. Snapshot is sorted by subsystem name, so two reads of an unchanged
process are comparable. A panicking probe is recovered individually and
reported as unknown with the panic value in a note — one broken probe
cannot hide the other thirteen.
Counted diagnostics
Most subsystems answer for free from state they already hold. A few don't — compression doesn't know how many responses it compressed, the rate limiter knows its client map but not how many requests it rejected — and those facts can only be recovered by counting them as they happen.
Why there is a gate
A shared counter incremented by every core moves a cache line between cores on every request; under real concurrency that coherence traffic is the cost, not the increment. So counting is off by default and the hot path reads a gate first:
if !counting.Load() { return } That load sits in every core's cache in shared state, costs no coherence traffic, and predicts perfectly.
diag.EnableCounters() // idempotent, safe at any time
diag.DisableCounters() // keeps what was counted
since, on := diag.CountersSince() dashboard.Install and mcp.ServeInProcess both call EnableCounters —
their presence already means the process accepted observability cost. A
bare application that installed neither pays nothing for counters it has
no reader for.
Every counted number says whether it was counted: CounterSnapshot.Counting travels with the numbers deliberately, since
zeroes with counting: false mean "not measured", not "did not happen".