Tracey / Documentation
Operate production agents with evidence.
Connect your existing agents, observability backend, and infrastructure. Tracey helps you observe runs, investigate failures, prepare safe changes, and verify recovery.
Start with one agent
The quickest path is a read-only setup. Connect SigNoz, register one agent, emit a root run span, and confirm the complete execution appears in Tracey.
- Getting started: connect the stack in one guided path
- Agent runs: understand the telemetry model
- Custom OpenTelemetry: instrument any framework
The operating loop
Tracey is built around a durable workflow rather than a chat window: observe, investigate, decide, execute, verify, and recover.
- Investigations across runs, traces, logs, metrics, and Kubernetes
- Approval workflow with typed actions and policy evaluation
- Verification based on health evidence, not API acceptance
Connectors and boundaries
SigNoz remains the telemetry system of record. Agents stay independently deployed. Tracey adds the semantic, policy, and operator layer above them.
- SigNoz for traces, logs, and metrics
- Kubernetes through separate investigator and executor identities
- Codex, Claude Code, custom OTel agents, and allowlisted MCP
Read this before production access
Start in Observe or Recommend mode. Keep production approval-first until your policies, verification criteria, audit history, and recovery paths have been exercised in a non-production environment.
Tracey never gives a model a shell or unrestricted cloud credentials. Missing telemetry is reported as missing evidence.