LLM Observability

LLM observability is the practice of recording what an LLM application does at runtime so that its behavior, cost, and quality can be inspected after the fact. It captures traces made of spans for each model call, retrieval, and tool call, with the prompt, response, token counts, latency, and errors attached to each span. Classic application monitoring answers whether a service is up and fast, while LLM observability also has to answer whether a nondeterministic output was correct, which requires storing content and attaching evaluation scores. In agentic systems, a single user request can fan out into many model and tool calls, so a trace is often the only way to see which step went wrong.

OpenTelemetry publishes semantic conventions for generative AI that standardize attribute names for the model, token usage, and operation type, so traces can move between open-source and commercial backends. Typical uses include finding the prompt version behind a regression, attributing token spend to features or customers, replaying failed runs, and feeding sampled traces into LLM-as-a-judge or human review. Traces contain user input and model output, so retention limits and redaction of personal data belong in the initial setup.

Work with us

Ready to put agentic AI to work?

Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.