One alias in front of Groq.
Your application calls a Relays alias. Relays holds the Groq credential, applies the route policy, and hands back the provider response in the shape your SDK expects.
- Your application
Keeps the OpenAI-compatible request it already sends.
- Relays
Holds the key, picks the candidate, writes the receipt.
Groq
Receives the call on its own base URL.
- OpenAI-compatible
- Chat Completions
- Responses
- Bring your own Groq API key
GROQ_API_KEY https://api.groq.com/openai/v1- Maintained Relays preset
Relays’ Groq preset declares Chat Completions and Responses support, so native Responses payloads can remain on that edge when the selected model supports them.
What is measured, and what is not.
Availability comes from Groq. Latency, error rate, and cost per successful request come from your own traffic, so Relays leaves them empty until requests run.
Relays measurements
Models at the edge.
Model IDs sit behind a Relays alias, so a catalogue change at Groq never reaches your code.
Every route carries a policy.
A route is a list of candidates, a budget, and a record of what happened. You set all three before any traffic runs.
Candidates stay inside the contract.
Groq can be ordered with other configured candidates that support the requested Chat Completions or Responses endpoint and declared capabilities.
See how a route is chosenA ceiling on this route.
Set an alert threshold and a stop threshold for the route rather than for the whole application. Relays checks both before the call goes upstream.
Each attempt is written down.
Every call records the alias, the candidates tried, the status each returned, tokens in and out, and the estimated cost of the attempt that served it.
Read a route receipt
One endpoint is enough to start.
Keep the client you already wrote. Bring us the workload and we will map the first Groq route with you.
