LLM observability answers “is the model behaving”, for engineers, on storage built to be cheap and short-lived. That is a good product. It is not the artefact you hand a regulator.
Both watch the same agent do the same thing. They are built to answer different people, and it shows in every design decision underneath.
The shortest-lived trace store in a typical stack keeps roughly seventy times less than the shortest obligation it would have to satisfy.
If you already emit traces, you have done most of the instrumentation work. We read the same spans.
OpenTelemetry spans, or the OpenInference instrumentors for OpenAI, Anthropic and LangChain.
Model calls become recorded actions, tool executions become tool calls, and payloads become fingerprints.
Sealed, countersigned, and checkable by someone who has never heard of your stack.
Including the one where the answer is that you may not need us yet.
Keep the traces. Add the record.
Two lines in one agent, free while you evaluate — and it reads the spans you already emit.