Pricing · Early access

Priced against the work, not the tokens.

Relays has not published prices yet. The shape of the deal is already fixed: your provider bills the inference on your own keys, and the Relays edge is agreed against your workload before you run it.

Two bills. One of them ours.

Your provider bills the inference, on your account, at your rates. Relays charges for the edge around it: the routing decision, the policy gate, the ordered fallback, and the receipt trail.

Your applicationOne request

Existing SDK, unchanged contract.

RelaysThe edge

Route, gate, fallback, receipt.

Your providerYour key

Inference on your own account.

Relays edge priced per workload, agreed before you run it

Provider inference passed straight through to your provider bill

Relays never becomes the account of record for your inference.

See the cost before you commit.

Spend control is a product feature, not a billing tier. You set the rules that order and gate the routes, and every call leaves a receipt tied to the request that caused it.

Route receipt · field guideIllustrative
Decision
which route ran, and why it cleared your rules
Gate
the capability and policy checks the route had to pass
Attempts
every attempt in order, with the outcome of each
Tokens
input and output, recorded per attempt
Est. cost
computed from your own provider's rates
Request
the call that caused all of the above
Illustrative structure, not a sample run. Relays fills these fields from your own calls, so nothing here is invented.
  • Ordered routes

    Narrower or cheaper routes can run first. A route becomes eligible only once it clears the capability and policy rules you set.

  • Visible fallback

    When a provider fails or rate-limits, Relays records each attempt on its own line instead of folding them into one opaque charge.

  • Attributable cost

    Estimated cost stays attached to the request, the route, and the attempt that produced it. You can read a workload before you commit to it.

How the conversation works.

There is no quote form. A number written before the routes are mapped would be a guess. This is the sequence instead.

  1. 01

    Bring a workload

    Start with real work, not a seat count. An agent loop, a batch job, or a product surface that already calls a model.

    You send it
  2. 02

    We map the routes

    Together we write the route order and the capability and policy boundary the workload runs inside.

    We map it
  3. 03

    Run it against receipts

    The workload runs through the edge and leaves a receipt per call: decision, attempts, tokens, estimated cost.

    You run it
  4. 04

    Agree the price

    We price against what actually ran, on evidence you can read. Published rates land on this page once they settle.

    We agree it

Straight answers

What engineering and finance ask first.

Bring us a workload.

We will map the routes and the boundary, run it against a receipt trail, and agree pricing on what actually ran.

Start a conversation