Blog
Reliability for AI teams
Notes on agent failures, cost regressions, and building the SRE layer for LLM apps.
SRE for AI agents: why uptime beats prompt engineering
LLM observability tools focus on evals. Production teams need error rates, cost regressions, and alerts that page someone.
Read →
Streaming cost tracking without guesswork
Token counts from streamed responses are tricky. Here's how Puse captures usage mid-stream and prices it against your model table.
Read →
Silent agent failures: the $40 retry loop
A provider hiccup, a rate limit, a retry loop quietly burning money — and nothing paged anyone. Sound familiar?
Read →