Skip to main content

Stop runaway workloads

Enterprise Tier

A rate limit holds one consumer to a token ceiling per hour, so a runaway workload hits a wall at a chosen volume instead of at three weeks of spend. Ceilings are set through the management API on an API key, a project, a tag, or the whole deployment, and each key's owner can additionally cap their own key from the Console.


Request ceilings and per-minute windows arrive in the 2026 Q4 release

A rate limit today is a token ceiling per hour. The 2026 Q4 release adds request ceilings, per-minute windows, and the Admin Console screen that sets both on a project or an API key. Those are documented ahead of release in Set rate limits on projects and keys. Everything on this page is available now.

A rate limit is flow control, not a spend cap. It protects shared capacity by bounding the volume of traffic; it does not stop a consumer from spending a budget over the course of a month. The two are complementary, and choosing between them is covered in Choose the right cost control for each workload; the spend caps themselves are set in Set a budget and track spend against it.

Persona: Platform operator working through the management API, with the Console for monitoring.

Estimated time: 15 to 25 minutes for an initial pass; ongoing as workloads evolve.

When this guide applies

Rate limits are the right concern in any of these situations:

SituationControl that fits
A consumer drives large token volumes (long contexts and verbose completions) that strain upstream capacityAn hourly token ceiling on its key
One team's traffic should never crowd out another team sharing the same deploymentA ceiling on the tag that team's keys carry
A new or untrusted integration should be bounded before its real traffic shape is knownA deliberately tight ceiling at the key scope
Aggregate volume across every consumer should be boundedA ceiling on the whole deployment
A retry loop or runaway agent is sending far more requests than a healthy client wouldA token ceiling bounds the damage; ending the loop means addressing the key, as in Contain a leaked key before it drains the budget. Request ceilings arrive in the 2026 Q4 release
Spend over a month should be capped regardless of traffic volumeA budget, not a rate limit. See Set a budget and track spend against it

Outcomes

By the end of this guide:

  • The available counters (input, output, or total tokens per hour) and subjects (an API key, a project, a tag, or the whole deployment) are understood.
  • At least one rate limit is applied at the subject appropriate to the workload it protects.
  • The caller-side experience of a hit limit, an HTTP 429 response, is understood, along with the back-off behavior applications are expected to implement.
  • A way to confirm whether limits are being hit is in place.

Prerequisites

  • Administrator access to the management API: typically the super_admin role, or a role with permission to manage limits.
  • At least one provider and model already provisioned. See Provision models and providers.
  • For per-tag limits, the governed tag catalog entries that the affected keys carry. See Know what every app and project actually costs.
  • Coordination with the developer teams that own the affected keys. Limits are enforced inline, so a misconfigured limit produces visible production impact.

Step 1: choose the counter

A rate limit counts tokens over an hour. The counter can cover input tokens, output tokens, or the two together.

CounterWhat it caps
Input tokens per hourPrompt volume: long contexts and large documents
Output tokens per hourCompletion volume: verbose or unbounded generation
Total tokens per hourBoth together, as one figure

Total tokens is the general-purpose ceiling. The separate input and output counters fit workloads whose load is lopsided: a summarization pipeline that sends large prompts and returns short answers is bounded by its input volume, and a generation workload by its output.

There is no request counter and no per-minute window today; both arrive in the 2026 Q4 release. A flood of small calls accumulates tokens slowly, so a token ceiling catches a retry storm late. Where the runaway is a misbehaving or compromised key rather than a legitimately heavy workload, the faster tool is containment.

Step 2: choose the subject

The subject determines which traffic the ceiling is measured against.

SubjectWhat the limit governsTypical use
An API keyAll traffic presented with one keyBounding a specific application or integration to its expected envelope
A projectTraffic under one projectBounding a piece of work whose keys share a project
A tagTraffic across every key carrying the tagReserving a fair share for a business function whose keys span projects
The whole deploymentAll traffic through the gatewayA backstop against aggregate volume, independent of who is driving it

Per-key limits are the most precise and the most common, because a key usually maps to a single workload. Tag limits sit above the key scope and follow the governed tag catalog, so a ceiling lands on every key that carries the tag, including keys issued after the limit is set.

Alongside the operator-set limits, each API key also carries Console-side hourly token limits (Total, Input, and Output tokens per rolling hour), configured by the developer who owns the key. The operator-side decision about which keys should carry which limits is what makes that mechanism useful: the operator establishes the policy ("research keys cap at 100 K tokens per hour") and the developer applies it. The developer-side flow is documented in Route requests across providers under What the API Keys page shows.

Step 3: set the limit through the management API

Rate limits are created and adjusted through the management API; there is no Admin Console screen for them until the 2026 Q4 release, when the endpoints also join the published API reference. Two behaviors of the current API are worth knowing:

  • The API accepts a period value of daily, weekly, or monthly, and the gateway ignores it: the count is always kept per hour.
  • A limit left unset for a given subject is not enforced for that subject. A key with no ceiling of its own is still covered by any project, tag, or deployment-wide ceiling that applies to it.

A changed limit reaches the gateway the way any configuration change does; the delivery window is described in Configuration propagation.

Step 4: size and adjust limits

Sizing is the part of this work that takes the most judgment. Two failure modes are worth avoiding.

  • Limits set too tight. Normal traffic hits the ceiling, healthy clients start receiving 429 responses, and the workload looks, from the consumer's side, as though Agent Router is failing. The cost of this failure mode is immediate and visible.
  • Limits set too loose. A runaway consumer is never actually constrained, and the limit becomes a number that never fires. The cost surfaces later, as overloaded upstreams or a degraded experience for other consumers sharing the same capacity.

The dependable approach is to size each limit from observed traffic. The expected peak is read from the usage surface, see Monitor traffic and usage, the highest legitimate hourly token volume over a representative window is taken as the baseline, and the limit is set comfortably above that baseline but well below any level that would constitute overload. Setting a limit at roughly two to three times the observed peak is a common starting point, adjusted for how bursty the workload is.

For a brand-new workload with no traffic history, a deliberately tight limit is the safer starting point. Real traffic surfaces the true shape quickly, and the limit is relaxed once the baseline is known. Relaxing a limit that proved too tight is a low-risk adjustment; discovering that a generous limit allowed an overload is harder to recover from.

Some workloads are quiet most of the time and very loud occasionally: end-of-month batch runs and scheduled report generation. A limit sized for the quiet baseline fires on the burst. The cleanest answer is usually to isolate the bursty work on its own key with its own, more generous limit, which keeps the steady-state key tight and the usage attribution clean. The operational complication of issuing a second key is typically smaller than the complication of explaining why a single key's traffic shape is irregular.

Step 5: monitor whether limits are being hit

A limit that never fires and a limit that fires constantly are both worth knowing about: the first may be set too loosely to matter, and the second is likely throttling healthy traffic. Both are visible in Agent Router's usage and traffic surfaces.

The signal to watch for is the rate of 429 responses, broken down by the subject the limit is set on. A workflow that pairs well with this guide:

  1. Open the usage and traffic view. See Monitor traffic and usage.
  2. Apply a time range that matches the cadence of the workload under review.
  3. Break the traffic down by the subject the limit is set on: by key, by project, or by tag.
  4. Look at the proportion of requests returning 429. A small, occasional fraction during peaks is normal; a sustained high fraction indicates the limit is too tight for the legitimate workload.

When a limit is firing more than expected, the choices are to raise it if the traffic is legitimate, see Step 4, or to address the consumer if the traffic is not, by working with the owning team or, in the case of a suspected-compromised key, following Contain a leaked key before it drains the budget.

What a caller experiences when a limit is hit

When a request would breach an active limit, the gateway rejects it with an HTTP 429 Too Many Requests response rather than forwarding it upstream. The rejection applies only to the request that crossed the threshold; once the window advances, traffic flows again. A 429 is therefore a transient, recoverable signal, not a hard failure, though to a client that does not handle it, it presents the same way as an error.

Applications are expected to back off and retry rather than fail or retry immediately:

  • Retry the request after a short delay rather than treating the 429 as fatal.
  • Use exponential back-off with jitter (increasing the delay on each successive 429, with a small random offset) so that many clients hitting the limit at once do not retry in lockstep and re-create the burst.
  • Respect any retry-after guidance the response carries, where the client library surfaces it.
  • Cap the number of retries so that a sustained limit does not turn into an unbounded retry loop, which is itself a source of the pressure the limit exists to contain.