About
SRE for AI agents
We built Puse after watching agents fail silently — a provider hiccup, a rate limit, a retry loop quietly burning money. Existing LLM observability leans toward prompt engineering and evals; we wanted pager-style reliability.
Puse traces every LLM call your app makes — model, tokens, latency, cost, errors — with a one-line SDK wrapper that handles streaming. A health dashboard shows error rates and spend per project; Slack alerts fire on error spikes and cost regressions.
The name comes from the pulse you want on every model call. Your AI agents fail quietly. Puse makes it loud.
We're self-hostable by design: Fastify ingest, Next.js dashboard, Postgres on one small VPS. Your traces never have to leave your infrastructure.