Features
SRE-grade monitoring for every model call
Trace latency, tokens, cost, and errors with a one-line wrapper. Alerts fire when error rates climb or spend spikes — not when your users complain.
One-line auto-instrumentation
Wrap once. Trace everything.
Wrap your OpenAI or Anthropic client once. Every call — including streams — is traced with model, tokens, latency, and cost. No proxy, no gateway in your request path.
Streaming-aware cost tracking
Know spend mid-stream.
Token usage is captured from streamed responses as they flow, priced against a per-model table you control. Know what every agent run cost, even mid-stream failures.
Health dashboard
See the spike first.
Requests, error rate, latency percentiles, and spend — per project, charted over 24 hours or 7 days. See the spike before your users do.
Alerts that page you, not spam you
Pager-style reliability.
Error-rate and cost-spike detection with tunable thresholds, delivered to Slack. Built for reliability, not vanity metrics.
Errors with context
3am becomes actionable.
Failed calls keep their error message, model, and latency — so 'it broke at 3am' becomes 'rate-limited on claude-sonnet for 12 minutes'.
Self-hostable, tiny footprint
Your traces stay yours.
A Fastify ingest service, a Next.js dashboard, Postgres. Runs on one small VPS. Your traces never leave your infrastructure.
See it in action
Instrument your first project in under two minutes. Streaming included.
Go to quickstart →