ENGINEERING BLOG · AI AGENT OBSERVABILITY

What Is AI Agent Observability? The Complete Guide to Seeing What Your Agents Do

An AI agent can return a perfect answer after calling the wrong tool, retrying three times, and spending ten times its budget. Your dashboards stay green. AI agent observability is how you see the run behind the answer: every model call, tool call, token, dollar, and error, stitched into one trace.

·14 min read·AI Agent Observability
Read the SDK Quickstart

Key takeaways

  • AI agent observability is the ability to reconstruct and explain an agent's full run, not just its final response.
  • The unit is the trace. One run becomes a tree of spans: the agent, each model call, each tool call, with tokens, cost, timing, and errors attached.
  • Logs and APM miss agent failures because the worst ones throw no exception: a wrong tool, a silent retry loop, a quietly expensive run.
  • OpenTelemetry keeps you portable. Agent telemetry should follow the same standard as the rest of your stack.
  • Observability tells you what happened. Stopping what should not happen is governance, and it needs the same trace underneath.

A support agent answers a customer in four seconds. The reply is polite and correct.

Behind it, the agent looked up the wrong order first, called the same tool twice, sent a 9,000-token prompt where 900 would have done, and only then found the right record. Nothing failed. No alert fired. The only signal was a bill at the end of the month.

That gap between the answer looked fine and the run was fine is what AI agent observability closes. This guide explains what it is, which signals matter, why OpenTelemetry is the right foundation, and exactly how Traccia's open-source Python SDK captures an agent run.

What Is AI Agent Observability?

AI agent observability is the ability to reconstruct and explain what an autonomous agent did, step by step, from the telemetry it emits: every model call, tool call, decision, token, cost, and error in a run.

The word comes from control theory: a system is observable when you can infer its internal state from its outputs. For software, it means you can answer a question you did not plan for in advance, from data you already collected.

LLM observability usually means watching single model calls: the prompt, the response, the latency, the token count. That is necessary, but an agent is not one call. It plans, picks tools, reads their output, changes course, and sometimes hands off to another agent. AI agent observability follows the whole run, so you can see not only what each call returned but why the next step happened.

LLM observabilityAI agent observability
Unit of analysisOne model callOne run: a tree of model calls, tool calls, and decisions
Main questionWas this response good, fast, and affordable?What did the agent do, in what order, and why?
Typical failureA slow or low-quality responseA wrong tool, a loop, a skipped step, a runaway cost
Cost viewTokens per callCost per run, per agent, per environment

Why Traditional Monitoring Falls Short for Agents

Application performance monitoring was built for deterministic services: the same request takes the same path, and a failure shows up as an error code or a slow endpoint. Agents break both assumptions.

The path is decided at runtime

The same request can take a different route on every run. You cannot instrument a fixed call graph because there isn't one.

The worst failures are silent

A wrong tool or a bad argument rarely raises an exception. The agent keeps going and returns something plausible.

Cost varies by the run, not the request

Two identical questions can differ in cost tenfold, depending on how many steps and tokens the agent decides to use.

Actions have side effects

An agent that refunds, deletes, or emails changes the world. You need to know which agent did it, in which environment, and what it was given.

QuestionLogs and APMAgent observability
Did the request succeed?YesYes
Which tools did the agent call, with which arguments?Only if you logged it by handEvery tool span carries its inputs and result
Why did this run cost more than the last one?NoTokens and cost on every model call, rolled up per run
Which agent and environment produced it?RarelyAgent and environment attributes on every span
Where did the time go?Per endpointPer step, inside the run

The Signals Worth Capturing

Good agent observability is not "log everything." It is a small set of signals, attached to the right step, that let you answer real questions later. Here is what to capture, and the attribute Traccia's SDK uses for it.

SignalThe question it answersTraccia attribute
Step typeIs this an agent step, a model call, or a tool call?span.type
ModelWhich model handled this step?llm.model
Prompt and completionWhat did the model see, and what did it say?llm.prompt, llm.completion
TokensHow much context went in and came out?llm.usage.prompt_tokens, llm.usage.completion_tokens
CostWhat did this step cost?llm.cost.usd
Tool inputs and resultWhat did the agent ask the tool, and what came back?function arguments, result
ErrorsWhat failed, and where?error.type, error.message, error.stack_trace
IdentityWhich agent, in which environment?agent.id, agent.name, env
GuardrailsDid a safety check fire?guardrail.triggered, guardrail.findings

Alongside spans, aggregate metrics answer fleet-wide questions such as token spend per model per day, without reading individual traces.

Anatomy of an Agent Trace

A trace is one run. A span is one step inside it, with a start time, a duration, a parent, and attributes. Nesting is what turns a pile of events into a story: the tool call and the model call sit under the agent step that made them.

Anatomy of an agent traceA support_agent span contains a lookup_order tool span and an OpenAI chat completion span. Each span carries its own attributes.ONE RUN = ONE TRACEEach step is a span. Children nest under the step that called them.AGENT · support_agentagent.id = support-agent · env = production · question = "Where is my order?"TOOL · lookup_orderspan.type = tool · order_id = ORD-1 · result = {status: shipped}LLM · llm.openai.chat.completionsllm.model = gpt-4o-mini · llm.usage.total_tokens · llm.cost.usdRead top to bottom: what the agent was asked, what it looked up, what the model cost.
The trace Traccia records for the support agent in the code below: one agent span with a tool span and a model span nested under it.

With that tree, the questions from the opening story have answers. The duplicate tool call is two sibling spans. The oversized prompt is a token count on one model span. The cost of the run is the sum of the cost on its model spans.

Why OpenTelemetry Matters for Agent Observability

OpenTelemetry (OTel) is the open standard for traces, metrics, and logs, supported by every major observability backend. Building agent observability on it has three practical benefits.

No lock-in

Spans are standard OTLP. You can send them to Traccia, to your own collector, or to Jaeger or Grafana Tempo, and change your mind later without re-instrumenting.

Shared vocabulary

OpenTelemetry's generative-AI semantic conventions name the common metrics. Traccia's SDK emits them under those names, such as gen_ai.client.token.usage and gen_ai.client.operation.duration.

One trace across services

W3C Trace Context headers carry the trace from your agent into the services it calls, so a tool that hits an internal API still lands in the same trace.

How Traccia Observes AI Agents

Traccia's observability starts in your process, with the open-source traccia-py SDK. It is built on the OpenTelemetry SDK, so everything below produces standard spans and metrics. Here is what it does, step by step.

1. Initialize once, and model calls are traced

Install the SDK and call init() before your agent runs.

bash
pip install traccia openai
agent.py
python
from openai import OpenAI
from traccia import init
init(api_key="YOUR_TRACCIA_KEY", agent_id="support-agent", env="production")
client = OpenAI()
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Where is order ORD-1?"}],
)

That chat completion is now a span named llm.openai.chat.completions carrying the model, the prompt, the completion and its finish reason, prompt and completion token counts from the provider's usage data, and the cost in USD. With patching on, which is the default, the SDK instruments:

  • OpenAI chat completions made with the synchronous Python client
  • Anthropic Messages API calls, with input and output tokens, cost, and the stop reason
  • Google Gemini calls through the google-genai Interactions API
  • Outbound HTTP made with requests: each call gets a client span, and a W3C traceparent header carries the trace to the service you call

Framework integrations cover the agent layer:

  • OpenAI Agents SDK: registered automatically when the package is installed; records agent, generation, tool, handoff, and guardrail spans with model and token usage
  • CrewAI: registered automatically; records crew kickoff, task, agent, and LLM call spans
  • LangChain: add Traccia's callback handler to trace LLM and chat-model runs

2. Turn your own functions into spans with @observe

Auto-instrumentation sees the model. It cannot see your agent loop or your tools. The @observe decorator makes any function a span, and nesting follows your call stack.

agent.py
python
from openai import OpenAI
from traccia import init, observe
init(api_key="YOUR_TRACCIA_KEY", agent_id="support-agent")
client = OpenAI()
@observe(name="lookup_order", as_type="tool")
def lookup_order(order_id: str) -> dict:
return {"order_id": order_id, "status": "shipped"}
@observe(name="support_agent")
def run_agent(question: str) -> str:
order = lookup_order("ORD-1")
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": f"{question} {order}"}],
)
return resp.choices[0].message.content

This produces the trace in the diagram above. For each decorated function, the SDK records:

  • The function's arguments as attributes, and its return value as result
  • The step type from as_type: "tool", "llm", or "guardrail"
  • Any exception, with error.type, error.message, and a stack trace, and the span marked as an error
  • Both regular and async functions

3. Give every run an identity

Every span carries agent.id, agent.name, and env, so you can filter one agent in one environment. Set them in init() for a single-agent service, or per run with run_identity when one process serves several agents.

server.py
python
from traccia import init, runtime_config
# Each request becomes its own trace instead of joining one process-wide trace.
init(api_key="YOUR_TRACCIA_KEY", auto_start_trace=False)
with runtime_config.run_identity(
agent_id="billing-agent", agent_name="Billing Agent", env="production"
):
run_agent("Why was I charged twice?")
One trace per request

By default init() opens a root span for the whole process, which suits scripts. In a server, pass auto_start_trace=False so each request starts its own trace.

4. Cost on every model call

The SDK prices each model span from a bundled snapshot covering more than 2,000 models, and writes the result to llm.cost.usd. It also records where the price came from (llm.pricing.source, llm.pricing.model_key, llm.pricing.snapshot_version) and where the token counts came from: the provider's own usage data first, then a local tokenizer, then an estimate, each labeled in llm.usage.source.

You can override prices for negotiated rates, and traccia pricing refresh pulls a newer snapshot. On the platform, Traccia recomputes cost against its own up-to-date catalog, so dashboards stay current even when an agent runs an older SDK.

5. Metrics for the whole fleet

Traces explain one run. Metrics show trends across all of them. The SDK records OpenTelemetry metrics at the moment each model call completes:

MetricWhat it measures
gen_ai.client.token.usageInput and output tokens per call
gen_ai.client.operation.durationModel call latency, in seconds
gen_ai.client.operation.costCost per call, in USD
gen_ai.client.completions.exceptionsFailed model calls
gen_ai.agent.runsAgent runs, for the framework integrations

Each data point carries the provider, the model, and the agent, so you can break spend down by any of them.

6. Sample traces without losing the totals

At high volume you may not want every trace. init(sample_rate=0.1) keeps one in ten traces, decided once at the root, so a kept trace is always complete and a dropped one leaves no orphan spans.

Sampling does not touch the numbers that matter for cost. Token, cost, and latency metrics are recorded inline, at the moment each model call completes, and never pass through the trace sampler. Keep 10% of traces for debugging and your spend totals still count 100% of calls.

7. Keep sensitive data out of telemetry

Prompts and tool arguments can contain personal data. Two controls run inside your process, before anything is exported:

agent.py
python
from traccia import init, observe
init(api_key="YOUR_TRACCIA_KEY", redact_pii=True)
@observe(skip_args=["password"], skip_result=True)
def login(user: str, password: str) -> str:
...
  • redact_pii=True masks email addresses, US phone numbers, and SSN-like patterns in prompt, completion, and message attributes, and marks the span with governance.redaction_applied. It is pattern-based, not a classifier.
  • skip_args and skip_result keep chosen arguments and return values out of the span entirely. Use them for anything redaction's patterns would not catch.

8. Send it anywhere

By default the SDK exports to Traccia over OTLP/HTTP. Point it at your own collector instead, or print spans to the console while you develop. The SDK does not require a Traccia account.

local.py
python
from traccia import init
# Send spans and metrics to your own OpenTelemetry collector (Jaeger, Tempo, ...)
init(
endpoint="http://localhost:4318/v1/traces",
metrics_endpoint="http://localhost:4318/v1/metrics",
agent_id="support-agent",
env="dev",
)
# Or keep everything on your machine while you develop:
# init(use_otlp=False, enable_console_exporter=True, enable_metrics=False)

9. Guardrail signals in the same trace

Mark your own checks with @observe(as_type="guardrail"); a check that returns True sets guardrail.triggered. The SDK also picks up the provider's own safety signals, such as an OpenAI completion that ends with a content_filter finish reason, and flags tool calls that fail with permission errors. It collects these findings per trace and summarizes which guardrail categories it saw and which it would expect to see given what the agent did.

The result: one OpenTelemetry trace per run, with the agent, its tools, its model calls, their cost, and its guardrails in one place, and nothing locked into a proprietary format.

From Observability to Control

Observability answers what happened. It cannot stop the next run from repeating it. For that you need governance: rules checked before an action executes. The two are not separate systems; governance runs on the same trace. When a call runs under Traccia's @govern(), the policy decision is written onto the span it applies to, as traccia.policy.decision_id and traccia.policy.effect. The evidence of what was allowed or blocked sits next to the evidence of what happened, and every span also carries a governance.integrity_hash for audit evidence.

The AI agent control plane loopObserve asks what happened, evaluate asks whether it was correct, govern asks whether it was allowed, and prove asks whether you can show it. All four run on the same trace.OBSERVEWhat happened?EVALUATEWas it correct?GOVERNWas it allowed?PROVECan we show it?All four questions are answered from the same trace.Observability is the foundation of the AI agent control plane, not a separate tool.
Observability is the first layer of an AI agent control plane. Evaluation, governance, and evidence all read the same trace.

That is the idea behind an AI agent control plane: one layer that observes agents, evaluates them, governs what they may do, and proves it afterwards. For the difference between watching and stopping, see AI Observability vs AI Governance.

An AI Agent Observability Checklist

PracticeWhy it matters
One trace per runYou can read a run end to end instead of stitching logs together
Every tool is a spanWrong tools and bad arguments become visible
Tokens and cost on every model callExpensive runs are explained, not just noticed
Agent and environment on every spanYou can isolate one agent in production
Errors recorded on the step that failedDebugging starts at the cause, not the symptom
Sensitive data redacted or skipped at the sourceTelemetry never becomes a data leak
Sample traces, never the cost metricsVolume stays manageable while spend totals stay exact
Standard OpenTelemetry exportYou keep your data and your options
The same trace feeds evaluation and governanceObserving, judging, and controlling agree on what happened

Frequently Asked Questions

What is the difference between LLM observability and AI agent observability?

LLM observability looks at individual model calls. AI agent observability looks at the whole run: the sequence of model calls, tool calls, and decisions, nested in one trace, with cost and identity attached.

Is OpenTelemetry enough for AI agents?

OpenTelemetry provides the transport, the data model, and shared metric names. Agents also need agent-specific attributes such as tool inputs, token usage, cost, and agent identity. Traccia's SDK adds those on top of standard OpenTelemetry spans.

Do I need a Traccia account to use the SDK?

No. The SDK is open source. Point it at your own OpenTelemetry collector, or print spans to the console. A Traccia account adds the hosted trace explorer, cost views, evaluation, and policies.

Can observability stop an agent from doing something harmful?

No. Observability records what happened. Stopping an action before it runs is governance, which Traccia applies through policies that run on the same trace.

References

See Every Step Your Agents Take.

Add Traccia's open-source SDK in a few lines and get one OpenTelemetry trace per agent run, with tools, model calls, tokens, and cost.

Read the SDK Quickstart