Features

SRE-grade monitoring for every model call

Trace latency, tokens, cost, and errors with a one-line wrapper. Alerts fire when error rates climb or spend spikes — not when your users complain.

One-line auto-instrumentation

Wrap once. Trace everything.

Wrap your OpenAI or Anthropic client once. Every call — including streams — is traced with model, tokens, latency, and cost. No proxy, no gateway in your request path.

Streaming-aware cost tracking

Know spend mid-stream.

Token usage is captured from streamed responses as they flow, priced against a per-model table you control. Know what every agent run cost, even mid-stream failures.

Health dashboard

See the spike first.

Requests, error rate, latency percentiles, and spend — per project, charted over 24 hours or 7 days. See the spike before your users do.

Alerts that page you, not spam you

Pager-style reliability.

Error-rate and cost-spike detection with tunable thresholds, delivered to Slack. Built for reliability, not vanity metrics.

Errors with context

3am becomes actionable.

Failed calls keep their error message, model, and latency — so 'it broke at 3am' becomes 'rate-limited on claude-sonnet for 12 minutes'.

Self-hostable, tiny footprint

Your traces stay yours.

A Fastify ingest service, a Next.js dashboard, Postgres. Runs on one small VPS. Your traces never leave your infrastructure.

See it in action

Instrument your first project in under two minutes. Streaming included.

Go to quickstart →