Two bills. One of them ours.
Your provider bills the inference, on your account, at your rates. Relays charges for the edge around it: the routing decision, the policy gate, the ordered fallback, and the receipt trail.
Existing SDK, unchanged contract.
Route, gate, fallback, receipt.
Inference on your own account.
Relays edge priced per workload, agreed before you run it
Provider inference passed straight through to your provider bill
Provider inference
Pass-through
You bring the provider keys you already hold. Token spend lands on the account you already have, at the rates you already agreed. Relays does not resell it and takes no margin on it.
- Your keys, your provider contracts
- No Relays margin on tokens
- Native contracts and capabilities intact
The Relays edge
Agreed per workload
Routing, policy, fallback, and the receipt trail. We price it against the workload you bring and agree it before you run it, so you never have to guess which tier you belong in.
- Ordered, capability-gated routes
- Fallback without an application rewrite
- A receipt per call, tied to the request
- Early access: we set pricing with the first teams
See the cost before you commit.
Spend control is a product feature, not a billing tier. You set the rules that order and gate the routes, and every call leaves a receipt tied to the request that caused it.
- Decision
- which route ran, and why it cleared your rules
- Gate
- the capability and policy checks the route had to pass
- Attempts
- every attempt in order, with the outcome of each
- Tokens
- input and output, recorded per attempt
- Est. cost
- computed from your own provider's rates
- Request
- the call that caused all of the above
Ordered routes
Narrower or cheaper routes can run first. A route becomes eligible only once it clears the capability and policy rules you set.
Visible fallback
When a provider fails or rate-limits, Relays records each attempt on its own line instead of folding them into one opaque charge.
Attributable cost
Estimated cost stays attached to the request, the route, and the attempt that produced it. You can read a workload before you commit to it.
How the conversation works.
There is no quote form. A number written before the routes are mapped would be a guess. This is the sequence instead.
Bring a workload
Start with real work, not a seat count. An agent loop, a batch job, or a product surface that already calls a model.
We map the routes
Together we write the route order and the capability and policy boundary the workload runs inside.
Run it against receipts
The workload runs through the edge and leaves a receipt per call: decision, attempts, tokens, estimated cost.
Agree the price
We price against what actually ran, on evidence you can read. Published rates land on this page once they settle.
Straight answers
What engineering and finance ask first.

Bring us a workload.
We will map the routes and the boundary, run it against a receipt trail, and agree pricing on what actually ran.
