Output Token Limit

Get alerted when an AI agent's responses run longer than your token budget. Output Token Limit watches output tokens over a window and opens a violation when it passes your limit.

After Run(Detective)Checked after the trace lands, so it records and alerts

Long answers cost more and slow users down. If a prompt change makes your agent ramble, you want to know the same day. Output Token Limit watches how many tokens the model writes and opens a violation when they run past your budget.

Runs On
After traces land, on an Hourly, Daily, or Weekly window
You Set
An output token limit, how it is combined (Avg by default), and how often to check.
Your Agent Sends
Nothing extra. Normal tracing records token counts.

The template starts at Avg of output tokens above 4000, checked Daily, with Warning severity. You can change every value.

What Happens

What Traccia seesResult
The window is at or under your limitNo violation.
The window goes over your limitA violation opens. With Warn or Block, the next governed run is warned or stopped.
It drops back under the limit in the same windowThe violation resolves on its own (Auto Resolved).

This runs after the trace is accepted, so the run that triggered it is never undone. Observe records the violation. Warn also warns the next governed run. Block stops the next governed run. Ingest never rejects a trace because of a policy.

Set It Up

  1. In the app, open Policies and click Create Policy.
  2. Choose Output Token Limit under After Run.
  3. Give the policy a name and pick its scope: the organization, a workspace, or one Agent ID.
  4. Leave Metric on Output Token Count. Pick an Aggregation: Avg for typical response length, Max for the single longest, Sum for total tokens written. Set the Threshold in tokens and the Frequency.
  5. Pick a mode. Observe only records. Warn and Block act on the next governed run.
  6. Turn on Email Alerts if you want mail when a violation opens.
  7. On Review, run Simulate Last 7 Days to see how often it would have matched, then click Activate Policy.

Good To Know

  • Output Token Limit looks at finished traces. It does not cut a response short.
  • Spend Cap limits what a call costs in dollars. It does not limit tokens, so use this policy when token count itself is what you care about.
  • Pair it with Input Token Limit to watch both sides of the conversation.

Where You See Results

Open the policy's Violations tab for what needs attention and Decision Log for every check. See Violations for the full lifecycle.

Next Steps

© 2026 Traccia.