ENGINEERING BLOG · AI AGENT OBSERVABILITY
What Is AI Agent Observability? The Complete Guide to Seeing What Your Agents Do
An AI agent can return a perfect answer after calling the wrong tool, retrying three times, and spending ten times its budget. Your dashboards stay green. AI agent observability is how you see the run behind the answer: every model call, tool call, token, dollar, and error, stitched into one trace.
Key takeaways
- AI agent observability is the ability to reconstruct and explain an agent's full run, not just its final response.
- The unit is the trace. One run becomes a tree of spans: the agent, each model call, each tool call, with tokens, cost, timing, and errors attached.
- Logs and APM miss agent failures because the worst ones throw no exception: a wrong tool, a silent retry loop, a quietly expensive run.
- OpenTelemetry keeps you portable. Agent telemetry should follow the same standard as the rest of your stack.
- Observability tells you what happened. Stopping what should not happen is governance, and it needs the same trace underneath.
A support agent answers a customer in four seconds. The reply is polite and correct.
Behind it, the agent looked up the wrong order first, called the same tool twice, sent a 9,000-token prompt where 900 would have done, and only then found the right record. Nothing failed. No alert fired. The only signal was a bill at the end of the month.
That gap between the answer looked fine and the run was fine is what AI agent observability closes. This guide explains what it is, which signals matter, why OpenTelemetry is the right foundation, and exactly how Traccia's open-source Python SDK captures an agent run.
What Is AI Agent Observability?
The word comes from control theory: a system is observable when you can infer its internal state from its outputs. For software, it means you can answer a question you did not plan for in advance, from data you already collected.
LLM observability usually means watching single model calls: the prompt, the response, the latency, the token count. That is necessary, but an agent is not one call. It plans, picks tools, reads their output, changes course, and sometimes hands off to another agent. AI agent observability follows the whole run, so you can see not only what each call returned but why the next step happened.
| LLM observability | AI agent observability | |
|---|---|---|
| Unit of analysis | One model call | One run: a tree of model calls, tool calls, and decisions |
| Main question | Was this response good, fast, and affordable? | What did the agent do, in what order, and why? |
| Typical failure | A slow or low-quality response | A wrong tool, a loop, a skipped step, a runaway cost |
| Cost view | Tokens per call | Cost per run, per agent, per environment |
Why Traditional Monitoring Falls Short for Agents
Application performance monitoring was built for deterministic services: the same request takes the same path, and a failure shows up as an error code or a slow endpoint. Agents break both assumptions.
The path is decided at runtime
The same request can take a different route on every run. You cannot instrument a fixed call graph because there isn't one.
The worst failures are silent
A wrong tool or a bad argument rarely raises an exception. The agent keeps going and returns something plausible.
Cost varies by the run, not the request
Two identical questions can differ in cost tenfold, depending on how many steps and tokens the agent decides to use.
Actions have side effects
An agent that refunds, deletes, or emails changes the world. You need to know which agent did it, in which environment, and what it was given.
| Question | Logs and APM | Agent observability |
|---|---|---|
| Did the request succeed? | Yes | Yes |
| Which tools did the agent call, with which arguments? | Only if you logged it by hand | Every tool span carries its inputs and result |
| Why did this run cost more than the last one? | No | Tokens and cost on every model call, rolled up per run |
| Which agent and environment produced it? | Rarely | Agent and environment attributes on every span |
| Where did the time go? | Per endpoint | Per step, inside the run |
The Signals Worth Capturing
Good agent observability is not "log everything." It is a small set of signals, attached to the right step, that let you answer real questions later. Here is what to capture, and the attribute Traccia's SDK uses for it.
| Signal | The question it answers | Traccia attribute |
|---|---|---|
| Step type | Is this an agent step, a model call, or a tool call? | span.type |
| Model | Which model handled this step? | llm.model |
| Prompt and completion | What did the model see, and what did it say? | llm.prompt, llm.completion |
| Tokens | How much context went in and came out? | llm.usage.prompt_tokens, llm.usage.completion_tokens |
| Cost | What did this step cost? | llm.cost.usd |
| Tool inputs and result | What did the agent ask the tool, and what came back? | function arguments, result |
| Errors | What failed, and where? | error.type, error.message, error.stack_trace |
| Identity | Which agent, in which environment? | agent.id, agent.name, env |
| Guardrails | Did a safety check fire? | guardrail.triggered, guardrail.findings |
Alongside spans, aggregate metrics answer fleet-wide questions such as token spend per model per day, without reading individual traces.
Anatomy of an Agent Trace
A trace is one run. A span is one step inside it, with a start time, a duration, a parent, and attributes. Nesting is what turns a pile of events into a story: the tool call and the model call sit under the agent step that made them.
With that tree, the questions from the opening story have answers. The duplicate tool call is two sibling spans. The oversized prompt is a token count on one model span. The cost of the run is the sum of the cost on its model spans.
Why OpenTelemetry Matters for Agent Observability
OpenTelemetry (OTel) is the open standard for traces, metrics, and logs, supported by every major observability backend. Building agent observability on it has three practical benefits.
Spans are standard OTLP. You can send them to Traccia, to your own collector, or to Jaeger or Grafana Tempo, and change your mind later without re-instrumenting.
OpenTelemetry's generative-AI semantic conventions name the common metrics. Traccia's SDK emits them under those names, such as gen_ai.client.token.usage and gen_ai.client.operation.duration.
W3C Trace Context headers carry the trace from your agent into the services it calls, so a tool that hits an internal API still lands in the same trace.
How Traccia Observes AI Agents
Traccia's observability starts in your process, with the open-source traccia-py SDK. It is built on the OpenTelemetry SDK, so everything below produces standard spans and metrics. Here is what it does, step by step.
1. Initialize once, and model calls are traced
Install the SDK and call init() before your agent runs.
pip install traccia openaifrom openai import OpenAIfrom traccia import init
init(api_key="YOUR_TRACCIA_KEY", agent_id="support-agent", env="production")
client = OpenAI()resp = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": "Where is order ORD-1?"}],)That chat completion is now a span named llm.openai.chat.completions carrying the model, the prompt, the completion and its finish reason, prompt and completion token counts from the provider's usage data, and the cost in USD. With patching on, which is the default, the SDK instruments:
- OpenAI chat completions made with the synchronous Python client
- Anthropic Messages API calls, with input and output tokens, cost, and the stop reason
- Google Gemini calls through the google-genai Interactions API
- Outbound HTTP made with
requests: each call gets a client span, and a W3Ctraceparentheader carries the trace to the service you call
Framework integrations cover the agent layer:
- OpenAI Agents SDK: registered automatically when the package is installed; records agent, generation, tool, handoff, and guardrail spans with model and token usage
- CrewAI: registered automatically; records crew kickoff, task, agent, and LLM call spans
- LangChain: add Traccia's callback handler to trace LLM and chat-model runs
2. Turn your own functions into spans with @observe
Auto-instrumentation sees the model. It cannot see your agent loop or your tools. The @observe decorator makes any function a span, and nesting follows your call stack.
from openai import OpenAIfrom traccia import init, observe
init(api_key="YOUR_TRACCIA_KEY", agent_id="support-agent")client = OpenAI()
@observe(name="lookup_order", as_type="tool")def lookup_order(order_id: str) -> dict: return {"order_id": order_id, "status": "shipped"}
@observe(name="support_agent")def run_agent(question: str) -> str: order = lookup_order("ORD-1") resp = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": f"{question} {order}"}], ) return resp.choices[0].message.contentThis produces the trace in the diagram above. For each decorated function, the SDK records:
- The function's arguments as attributes, and its return value as
result - The step type from
as_type:"tool","llm", or"guardrail" - Any exception, with
error.type,error.message, and a stack trace, and the span marked as an error - Both regular and
asyncfunctions
3. Give every run an identity
Every span carries agent.id, agent.name, and env, so you can filter one agent in one environment. Set them in init() for a single-agent service, or per run with run_identity when one process serves several agents.
from traccia import init, runtime_config
# Each request becomes its own trace instead of joining one process-wide trace.init(api_key="YOUR_TRACCIA_KEY", auto_start_trace=False)
with runtime_config.run_identity( agent_id="billing-agent", agent_name="Billing Agent", env="production"): run_agent("Why was I charged twice?")By default init() opens a root span for the whole process, which suits scripts. In a server, pass auto_start_trace=False so each request starts its own trace.
4. Cost on every model call
The SDK prices each model span from a bundled snapshot covering more than 2,000 models, and writes the result to llm.cost.usd. It also records where the price came from (llm.pricing.source, llm.pricing.model_key, llm.pricing.snapshot_version) and where the token counts came from: the provider's own usage data first, then a local tokenizer, then an estimate, each labeled in llm.usage.source.
You can override prices for negotiated rates, and traccia pricing refresh pulls a newer snapshot. On the platform, Traccia recomputes cost against its own up-to-date catalog, so dashboards stay current even when an agent runs an older SDK.
5. Metrics for the whole fleet
Traces explain one run. Metrics show trends across all of them. The SDK records OpenTelemetry metrics at the moment each model call completes:
| Metric | What it measures |
|---|---|
| gen_ai.client.token.usage | Input and output tokens per call |
| gen_ai.client.operation.duration | Model call latency, in seconds |
| gen_ai.client.operation.cost | Cost per call, in USD |
| gen_ai.client.completions.exceptions | Failed model calls |
| gen_ai.agent.runs | Agent runs, for the framework integrations |
Each data point carries the provider, the model, and the agent, so you can break spend down by any of them.
6. Sample traces without losing the totals
At high volume you may not want every trace. init(sample_rate=0.1) keeps one in ten traces, decided once at the root, so a kept trace is always complete and a dropped one leaves no orphan spans.
Sampling does not touch the numbers that matter for cost. Token, cost, and latency metrics are recorded inline, at the moment each model call completes, and never pass through the trace sampler. Keep 10% of traces for debugging and your spend totals still count 100% of calls.
7. Keep sensitive data out of telemetry
Prompts and tool arguments can contain personal data. Two controls run inside your process, before anything is exported:
from traccia import init, observe
init(api_key="YOUR_TRACCIA_KEY", redact_pii=True)
@observe(skip_args=["password"], skip_result=True)def login(user: str, password: str) -> str: ...redact_pii=Truemasks email addresses, US phone numbers, and SSN-like patterns in prompt, completion, and message attributes, and marks the span withgovernance.redaction_applied. It is pattern-based, not a classifier.skip_argsandskip_resultkeep chosen arguments and return values out of the span entirely. Use them for anything redaction's patterns would not catch.
8. Send it anywhere
By default the SDK exports to Traccia over OTLP/HTTP. Point it at your own collector instead, or print spans to the console while you develop. The SDK does not require a Traccia account.
from traccia import init
# Send spans and metrics to your own OpenTelemetry collector (Jaeger, Tempo, ...)init( endpoint="http://localhost:4318/v1/traces", metrics_endpoint="http://localhost:4318/v1/metrics", agent_id="support-agent", env="dev",)
# Or keep everything on your machine while you develop:# init(use_otlp=False, enable_console_exporter=True, enable_metrics=False)9. Guardrail signals in the same trace
Mark your own checks with @observe(as_type="guardrail"); a check that returns True sets guardrail.triggered. The SDK also picks up the provider's own safety signals, such as an OpenAI completion that ends with a content_filter finish reason, and flags tool calls that fail with permission errors. It collects these findings per trace and summarizes which guardrail categories it saw and which it would expect to see given what the agent did.
From Observability to Control
Observability answers what happened. It cannot stop the next run from repeating it. For that you need governance: rules checked before an action executes. The two are not separate systems; governance runs on the same trace. When a call runs under Traccia's @govern(), the policy decision is written onto the span it applies to, as traccia.policy.decision_id and traccia.policy.effect. The evidence of what was allowed or blocked sits next to the evidence of what happened, and every span also carries a governance.integrity_hash for audit evidence.
That is the idea behind an AI agent control plane: one layer that observes agents, evaluates them, governs what they may do, and proves it afterwards. For the difference between watching and stopping, see AI Observability vs AI Governance.
An AI Agent Observability Checklist
| Practice | Why it matters |
|---|---|
| One trace per run | You can read a run end to end instead of stitching logs together |
| Every tool is a span | Wrong tools and bad arguments become visible |
| Tokens and cost on every model call | Expensive runs are explained, not just noticed |
| Agent and environment on every span | You can isolate one agent in production |
| Errors recorded on the step that failed | Debugging starts at the cause, not the symptom |
| Sensitive data redacted or skipped at the source | Telemetry never becomes a data leak |
| Sample traces, never the cost metrics | Volume stays manageable while spend totals stay exact |
| Standard OpenTelemetry export | You keep your data and your options |
| The same trace feeds evaluation and governance | Observing, judging, and controlling agree on what happened |
Frequently Asked Questions
What is the difference between LLM observability and AI agent observability?
LLM observability looks at individual model calls. AI agent observability looks at the whole run: the sequence of model calls, tool calls, and decisions, nested in one trace, with cost and identity attached.
Is OpenTelemetry enough for AI agents?
OpenTelemetry provides the transport, the data model, and shared metric names. Agents also need agent-specific attributes such as tool inputs, token usage, cost, and agent identity. Traccia's SDK adds those on top of standard OpenTelemetry spans.
Do I need a Traccia account to use the SDK?
No. The SDK is open source. Point it at your own OpenTelemetry collector, or print spans to the console. A Traccia account adds the hosted trace explorer, cost views, evaluation, and policies.
Can observability stop an agent from doing something harmful?
No. Observability records what happened. Stopping an action before it runs is governance, which Traccia applies through policies that run on the same trace.
References
- traccia-py: the open-source Traccia Python SDK (https://github.com/traccia-ai/traccia-py)
- OpenTelemetry semantic conventions for generative AI (https://opentelemetry.io/docs/specs/semconv/gen-ai/)
- W3C Trace Context (https://www.w3.org/TR/trace-context/)
- Traccia SDK quickstart (https://traccia.ai/docs/sdk/quickstart/)
- The @observe decorator (https://traccia.ai/docs/sdk/observe-decorator/)
- Auto-instrumentation (https://traccia.ai/docs/sdk/auto-instrumentation/)
- SDK metrics (https://traccia.ai/docs/sdk/metrics/)
- Exporters (https://traccia.ai/docs/sdk/exporters/)
- Traces in the Traccia platform (https://traccia.ai/docs/platform/traces/)
- AI Observability vs AI Governance (https://traccia.ai/blog/ai-observability-vs-ai-governance/)
- What Is an AI Agent Control Plane? (https://traccia.ai/blog/what-is-an-ai-agent-control-plane/)
See Every Step Your Agents Take.
Add Traccia's open-source SDK in a few lines and get one OpenTelemetry trace per agent run, with tools, model calls, tokens, and cost.