Every agent,
on the record.
Trace every step an agent takes, score it against your own judges and roll back a prompt before a bad run reaches a customer.
OpenTelemetry native · 14-day trial · No card
Run #48213
support-agent · v14 prompt
- 00.00PlanPlan refund request120 ms
- 00.12Toollookup_order("A-2291")340 ms
- 00.46Toolpolicy_search("refund window")210 ms
- 00.67LLMDraft reply · 1,204 tok1.2 s
- 01.87JudgeJudge · tone 0.92pass
- 01.90JudgeJudge · PII redactedflag
- 12M
- traces ingested a day
- 38 ms
- median ingest latency
- 240+
- judges in the library
- 9 / 10
- rollbacks run without a human
Instrument once.
Watch every run.
One SDK call wraps your agent. Every plan, tool call and model response becomes a span you can search, score and replay.
Wrap the agent
Add rheostat.trace() around your agent loop. Works with any framework.
trace(agent)Attach judges
Pick from 240 judges or write your own in a few lines.
judge("pii")Set the dial
Route 5% of traffic to a new prompt; roll back on any failed judge.
rollout 5%Everything an agent did,
and why.
Replay any run.
Scrub through a run step by step, swap the prompt and re-run it against the same inputs.
Cost per run.
Tokens and dollars per run, per prompt version.
$0.0041 median · −31% since v12
Judges you trust.
Score tone, accuracy, safety and cost on live traffic.
- Accuracy0.94
- Tone0.92
- PII0.99
Guardrails, not hopes.
Block, redact or roll back the moment a rule fires.
when judge.pii < 0.95 redact(output) dial(prompt, 0%)
Humans in the loop.
Low-confidence runs go to a person, with the full trace attached.
Plugs into the stack
you already run.
SDKs for four languages, OpenTelemetry in and out, and a warehouse export so your data team is never left out.
- Python SDK
- TypeScript SDK
- Go SDK
- Java SDK
- OpenTelemetry
- Webhooks
- REST API
- gRPC
- Terraform provider
- CLI
- MCP server
- Warehouse export
Teams that stopped guessing.
“We caught a prompt that leaked order numbers in the first hour. It never reached a customer.”
“Replay turned a two-day incident review into a twenty-minute one.”
“We ship a new prompt every day now. The dial means nobody is scared of it.”
Put your agents
on the record.
Free for your first million spans. Most teams see their first trace in under ten minutes.