Skip to main content

Routing policy reference

Enterprise Tier

Smart routing in Tetrate Agent Router covers two mechanisms, model aliases and attribute-based routing. The object model behind the attribute-based half is a small set of objects, presented in the order they relate to one another: the request attribute an application declares, the routing policy that holds an ordered list of rules, the scope that decides which policy applies to a key, the mode that decides whether it acts, and the record every request leaves behind. Each entry states what the object holds and how it is resolved. There are no procedures here: the scenarios that show these objects in use are on the smart routing page, and each links back here.


Scope and availability

Attribute-based routing is scheduled for the v0.7.0 release. It is not available yet, and is documented ahead of the release so that the design can be reviewed and planned against.

The shapes shown in text blocks on this page are illustrative, to show what an object holds. None is a schema.

The X-Tars-Metadata header format and its limits are the exception. They are fixed, and an application team can build against them now: the header is the half of this feature an application ships, and it is published ahead of the release so that work can start before a policy exists to match it.

Persona: Platform operator working in the Admin Console.

Estimated time: 15 minutes to read.

The objects​

Three objects, and each is defined below in this order.

ObjectWhat it isWhere it lives
Request attributeA named thing an application may declare about a requestDefined centrally in the Admin Console
Routing policyAn ordered list of routing rules with a default, attached to a scopeThe Policies area
Routing ruleOne condition and one destination, inside a policyInside a policy

The division is between the attribute definition, which is reusable and referenced by many policies, and the policy itself, which is an attachment: its rule list is its configuration, so it holds no reusable definition of its own.

Request attribute​

A request attribute is a reusable definition that many policies reference, defined centrally in the Admin Console. Defining one changes no traffic.

Request attribute
name workload, task, plan
source where the value comes from
value type one of a named set, or free text
allowed values when the value type is a named set
FieldHolds
NameThe attribute's identifier, used unprefixed by the calling application and prefixed by its source in a rule
SourceWhere the value comes from, and therefore whether the caller can set it
Value typeEither a named set of permitted values, or free text
Allowed valuesThe permitted values, when the value type is a named set

An attribute whose value type is a named set can be matched only against a value from that set. A request may still carry a value outside it: the value is recorded, matches no condition, and the request takes the policy's default.

Source, and the name a rule uses​

Source records where an attribute's value comes from. One source exists today, metadata, which means the value is set by the calling application on each request. A rule names an attribute by its source, so a rule reads metadata.workload rather than workload. The attribute definition and the rule builder state the source on every attribute, so a rule author can see what a condition depends on.

The prefix distinguishes values the caller sets from values the caller does not, so a source that is not caller-supplied can be added later without an existing rule becoming ambiguous.

The name is unprefixed in the header the application sends, because the header is the source. See how a routing rule knows what a request is for the wire form.

Attribute values reach the gateway in one request header. That header is the only input attribute-based routing reads: nothing is taken from the request body, and nothing is inferred.

X-Tars-Metadata: workload=batch,task=summarization

The format is a comma-separated list of key=value pairs. Each key is an attribute name as defined in the Admin Console, written unprefixed, because the header itself is the source.

Limits​

These limits are fixed, and they are applied to every request before any rule is evaluated.

LimitValueWhat happens outside it
Entries per request8The request fails with 400, naming the first entry past the limit
Key[a-z][a-z0-9_]{0,31}, so 1 to 32 characters beginning with a lowercase letterThe request fails with 400, naming the entry
Value[A-Za-z0-9._:-]{1,64}, so 1 to 64 charactersThe request fails with 400, naming the entry
Whole header2 KiB, measured before parsingThe request fails with 400

An entry with no = is malformed in the same way, and fails the request with 400. These failures apply only where a routing policy governs the request; see a malformed header fails the request.

Five parsing rules complete the format.

  1. An uppercase ASCII letter in a key is folded to lowercase before the key is validated, so Workload=batch matches the workload attribute. Values are never folded: they are compared byte for byte against the attribute's allowed values.
  2. Whitespace around = and around , is trimmed.
  3. Where a request carries the header on more than one line, the lines are joined with a comma and parsed as one, so two lines and a single line of two entries are equivalent.
  4. Where a key appears more than once, the last occurrence wins. A repeated key still counts toward the 8.
  5. An empty segment, such as a trailing comma or a blank header line, is ignored. It is not malformed.

A value cannot itself contain , or =. Both are separators, and the value grammar excludes them.

A malformed header fails the request, and nothing else does​

A header that breaks the format above is a bug in the calling application, so the request fails with 400. The error has type invalid_request_error, code invalid_metadata and param X-Tars-Metadata, and its message names the first bad entry so that it can be found in the calling code. Malformed means an entry with no =, a key or value outside its grammar, more than 8 entries, or a header over 2 KiB. A developer sees the failure on the first test request, the same way a malformed JSON body is caught.

The 400 applies only where a routing policy governs the request, in monitor as well as in enforce. A request to a project and key that no routing policy covers is never affected by what the header contains. Every rejection is recorded in the request log, so an administrator can see which callers send a bad header, not only the caller that received the error.

Nothing else about this header can fail a request. A well-formed header is never rejected, whatever it says. A key naming an attribute that is not defined, a value outside an attribute's allowed set, an attribute the caller did not send, no rule matching, a paused policy: each sends the request to the policy's default, exactly as a request with no header at all would be. Routing itself never refuses a request.

The practical consequence is that a well-formed mistake is silent. A team that ships workload=btch, or a key nobody defined, sees its requests served by the default, which is indistinguishable from a team that has not shipped the header at all. Reviewing a policy in monitor before it is enforced is how the two are told apart.

Attribute values are visible to administrators​

The attribute values a policy matched on are written to the request log, where any administrator of the organization can read them. They are visible per request in the Admin Console alongside the rest of the decision record.

Attribute values are therefore not a place for user identifiers, account numbers, email addresses, or any other personal or sensitive data. An attribute is a low-cardinality label describing a class of request, such as workload=batch or plan=wealth. Attributing spend to a person or a team is a different mechanism: see Know what every app and project costs.

Only the attributes the winning policy names are recorded, not everything the caller sent. Telemetry exported over OpenTelemetry carries the policy identifier, the position of the matched rule, and whether any attribute was present at all, never the attribute names or their values.

Routing metadata is also a separate namespace from the tags an administrator assigns to an API key. A calling application cannot set a tag, and routing metadata is never read as one, so a caller cannot make its requests look as though they carry an administrator's tag.

Routing policy​

A routing policy is an ordered list of routing rules with a required default, attached to a scope and held in one of two modes.

Routing policy
scope project | key
mode monitor | enforce
rules an ordered list, evaluated top to bottom
default required, applied when no rule matches

A routing policy is a policy type, so it is created, targeted, monitored, enforced, and listed the same way every other policy type is.

Routing rule​

Routing rule
condition attribute, operator, value
destination a model available to the project

A condition compares one attribute against one value, or against a set of values, using one of four operators.

OperatorMatches when the value declared is
EqualsThe single value the condition names
Does not equalAnything other than the single value the condition names
InOne of the set of values the condition names
Not inAbsent from the set of values the condition names

Conditions may be combined within a rule, and all of them must hold for the rule to match, so one rule can require both that metadata.workload is batch and that metadata.task is summarization.

Every rule carries at least one condition. A rule with none would match every request and shadow everything below it, which is what the default already does, so a rule without a condition cannot be saved. A policy that should send every request from its target to one model is written as a default with no rules at all, which is 9: a degraded provider.

A rule's destination is a model available to the project. Routing chooses between the models a project already has, and cannot reach one the project does not.

The model the caller named​

A matched rule overrides the model named in the request. An application that asks for opus in its request body and matches a rule whose destination is haiku is served by haiku. The model the caller named is not a condition, and it is not consulted when a rule is chosen.

That is the behavior the feature exists for: the applications worth moving to a cheaper model overnight are exactly the ones that name a model directly in their code, and a policy that yielded to them would move nothing.

The override is bounded in one direction only. A policy chooses among the models the key can already reach, so an override moves a request sideways within that set and never widens it. Both the model the caller named and the destination the policy chose are kept on the decision record, and the response identifies the model that actually served the request.

Two kinds of request are outside routing altogether, whatever policy covers the key.

RequestWhy routing is skipped
One authenticated with the caller's own provider credentialThe credential reaches one provider, so the request cannot be moved to a model held anywhere else
One carrying no model in its bodyMCP, health, well-known, and OAuth paths are answered before a body is read, so there is nothing to rewrite

In both cases the model the caller named stands, no rule is evaluated, and the request carries no routing decision.

The default​

Every policy has a default, and it is required. A request that matches no rule is served by the default. A routing policy cannot reject a request, so a caller that sends a well-formed header and declares nothing, declares a value the attribute definition does not allow, or declares an attribute that does not exist is served the default rather than an error.

The default overrides the model the caller named, exactly as a matched rule does. An enforcing policy therefore gives every request it applies to a destination, whether a rule matched or not, subject to the two carve-outs under the model the caller named. Choosing a default is consequently a decision about all the traffic no rule describes: setting it to the model those callers are served today keeps them where they are, and setting it to anything else moves them the moment the policy is enforced.

Evaluation order​

The first matching rule wins. Rules are evaluated top to bottom, and evaluation stops at the first match.

Three things cause a rule not to match, and all three have the same consequence: evaluation continues at the next rule, and reaches the default only when every rule has been tried.

CauseExample
The attribute the condition names was not declared on the requestA rule on metadata.workload against a request that sent no attributes
The value declared is not the value the condition tests forA rule on metadata.workload is batch against workload=interactive
The rule's destination is not valid on the endpoint the request arrived onThe model named by the rule is not served in that endpoint's mode

Scope resolution decides which rule list applies to a key. The rule list is evaluated per request. A policy resolves to one rule list for a key, and which rule inside that list wins varies request by request.

Scope and resolution​

The scope set​

A routing policy targets one of two scopes.

ScopeSelects
ProjectEvery key issued in that project
KeyOne API key

Organization and tag selector scopes are not offered in this release. A routing policy cannot be created against either, and an attempt to do so is refused. An organization-scoped policy is intended to bind every project beneath it, which is the opposite of how every other policy family resolves, and the scope is withheld until that precedence rule is implemented rather than shipped with behavior a reader would have to guess at.

There is no team scope and no user scope. A policy selects keys by where they are issued, never by key ownership. Routing traffic per team is covered under things a routing rule cannot do.

Precedence​

A key resolves to exactly one routing policy, because two rule lists cannot be merged. Where both a key-scoped and a project-scoped policy could apply, the narrower target wins.

OrderScopeNote
1KeyWins for the one key it targets
2ProjectApplies to every key in the project that no key-scoped policy targets

The narrower policy replaces the wider one rather than combining with it. A key-scoped policy substitutes its whole rule list for the project's, and the project's rules are not consulted for that key, including the project's default. Resolution ignores mode: whichever policy wins is the one whose mode then decides whether traffic moves.

Coverage​

The coverage figure shown against a policy counts the API keys that resolve to a rule list. It counts keys covered, not routing decisions taken, so it rises the moment a policy is saved and says nothing about whether any request has matched a rule. What each rule would have matched is the monitor readout, described under Modes.

The inheritance view and the effective policy view that apply to every policy type apply to routing policies too, and answer what a given key resolves to.

Modes​

A routing policy is held in one of two modes.

ModeWhat happens to trafficWhat is recorded
MonitorNothing changes. Every request goes where it would have gone with no policy at allThe rules are evaluated against real traffic, and how many requests each rule would have matched over a recent window, with the model each would have moved to
EnforceThe rules decide the destinationThe full decision record

Every policy is created in monitor. Enforce is reached only by changing the mode on a policy that already exists, never as part of creating or editing one, so the audit log carries the moment traffic began to move as an event of its own. Reviewing a policy in monitor before it is enforced is covered in A policy has to be reviewed before it takes effect.

Mode is not status. A policy separately carries an enabled or disabled status, kept apart from mode so that a policy sitting in monitor can be paused and resumed without losing which of the two modes it held.

Monitor proves matching, not quality. The destination a rule names is never actually called in monitor, so nothing recorded there says whether its answers would have been suitable.

The decision record​

Every request records why it went where it went.

RecordedDetail
The routing that appliedWhich policy, which scope it was targeted at, and which mode it was in
The rule that matchedWhere in the rule order the match happened. A request that matched no rule records which of the three reasons the default applied for: no attributes were carried, none of them matched a rule, or every rule whose conditions did match named a destination not valid on the endpoint the request arrived on
The attributes evaluatedThe attributes the winning policy names, with the value declared for each, and any whose value fell outside the allowed set
The model the caller namedRecorded whenever the policy replaced it
The destination the policy choseThe model the rule or the default named
The model that servedWhat answered after any fallback or traffic split moved the request again
How it turned outWhether it succeeded, how long it took, and what it cost

Three model identities are kept apart, because they can all differ on one request: what the application asked for, what routing chose, and what finally served it after an availability fallback, a traffic split, or a budget action moved it again. Reading the destination alone hides that.

The response itself identifies the model that actually served the request, so a caller can tell what served a call without opening anything.

The values sit on the per-request record beside what that request cost and how long it took, so a routing decision can be reviewed against its own outcome. They are not aggregated. Attribute values never enter the usage rollups, and routing metadata is a separate namespace from the API key tags spend is grouped by, so reading spend along an attribute is not something this record supports. What the record holds, and who can read it, is covered under attribute values are visible to administrators.

Where each object is managed​

ObjectSurfaceOperations
Request attributeDefined centrally in the Admin ConsoleCreate, edit, remove, and see which policies reference one before changing it. The definition view shows how many policies hold a condition on the attribute and which rules compare against it, so an attribute no policy references is visible as removable
Routing policyThe Policies areaCreate, target, monitor, enforce, list

An attribute definition is organization-wide and shared by every project's rules, while a routing policy is targeted at one project or one key. Renaming an attribute or withdrawing one of its allowed values therefore reaches every project at once, which is why the two are held at different permission levels: defining attributes is an administrator operation, and writing a policy against them belongs to the project.

An attribute that a live rule condition still compares against cannot be deleted. The attempt is refused and names the policies blocking it, so the rules have to be changed first.

Everything available in the Admin Console is available in the admin API, so a routing policy can be managed as code.

Policy changes reach the data planes in under a minute rather than instantly.

Roles​

RoleOn routing
AdministratorDefines attributes centrally in the Admin Console, and creates, targets, and enforces routing policies
DeveloperDeclares attributes on the requests their application sends. Developers hold no Admin Console role in routing: they do not create or see routing policies

A rule matching on an attribute is inert until the application team that owns the calling code sends it.

What a routing policy is not​

Routing decides which of a project's models serves a request. Four nearby things are decided elsewhere, and each is a separate mechanism with its own guide.

Not a routing policyWhat decides itWhere it is covered
Which models a key may use at allThe project's model setCreate a project
Where a request goes when the destination failsA fallback policySet up fallback
A fixed percentage across two backendsA traffic splitReduce cost with traffic splitting
What happens when a spend limit is reachedA budget enforcement actionEnforce a budget cap

The cases a routing rule cannot cover, with what to do instead, are listed under things a routing rule cannot do.