Explainer

What is an LLM API gateway?

An LLM API gateway is a service that sits between your application and several model providers. Your code sends OpenAI-shaped requests to one base URL with one API key, and the gateway routes each request to whichever provider serves the model you named, then reports usage in one place.

Last updated: October 2026

What does an LLM API gateway actually do?

It does four jobs: it authenticates your requests, it routes each one to the provider that serves the requested model, it normalises the request and response shapes, and it records token usage so the cost lands on one balance.

  • Authentication and key management - one credential instead of one per provider.
  • Routing - the model name in the request decides which upstream endpoint is called.
  • Protocol translation - OpenAI-shaped requests are translated to providers that use a different wire format, such as Anthropic.
  • Metering and billing - token counts, per-request detail and a single balance rather than several invoices.
  • Operational controls - rate limits, retries and failover when an upstream provider degrades.

Why not call each provider directly?

Calling providers directly is fine until you need a second vendor. At that point you take on a second SDK, a second key store and a second invoice, and every experiment becomes an integration project instead of a configuration change.

  • Switching models becomes a one-line change rather than a new SDK and a new billing relationship.
  • Comparison work stops being biased by whichever provider was easiest to wire up first.
  • Credential handling stays in one place, which matters when several tools share the same budget.
  • Spend is attributable per key and per request, so a runaway job is visible instead of absorbed.

What does "OpenAI-compatible base URL" mean?

It means the endpoint accepts the same request and response shape as OpenAI’s chat completions API. Existing SDKs work after you change the base URL, the key and the model name - no new client library, and no rewrite of the code around it.

Same request shape, different endpoint
from openai import OpenAI

client = OpenAI(
    base_url="https://api.quickrouter.homes/v1",
    api_key="sk-qr-your-key",
)

response = client.chat.completions.create(
    model="claude-sonnet-4-5",
    messages=[{"role": "user", "content": "Summarise this changelog."}],
)

When is a gateway the wrong choice?

If you have committed to a single provider, negotiated enterprise terms with it and have no plan to test another model, a gateway adds a hop without adding a decision. The value comes from optionality, and optionality you never exercise is overhead.

  • A single-model product with no evaluation roadmap gains little from a routing layer.
  • Latency-sensitive paths should measure the extra hop rather than assume it is free.
  • Regulated workloads need the data-handling terms of every upstream provider reviewed, not just the gateway’s.

What should you check before choosing a gateway?

Check how billing works, whether keys are scoped per tool, which models are actually available and whether the pricing is the upstream list price or a marked-up rate. Those four answers decide whether the gateway saves you money or just adds a layer.

QuestionWhy it matters
Is usage billed at list price or at a markup?A markup quietly changes the cost of every experiment you run.
Can you create separate keys per tool?Per-key spend is what makes a runaway agent visible before the invoice arrives.
How many models are actually reachable?A long catalogue that excludes the model you need is not optionality.
What protocol does each tool need?Anthropic-style clients take the host root; OpenAI-style clients take the /v1 path.
What happens when an upstream provider degrades?Routing, retries and failover behaviour is the operational half of the product.

How does QuickRouter fit this definition?

QuickRouter is an LLM API gateway with two differences that matter for cost: usage is billed from your balance at the upstream list price of the model you call, and the same key works across every model in the catalogue.

  • One OpenAI-compatible base URL: https://api.quickrouter.homes/v1.
  • One key across the whole catalogue, including OpenAI, Anthropic, Google, DeepSeek and xAI models.
  • Gateway plans start at $19 per month, with model usage billed separately at upstream list prices.
  • Per-request detail in the console, so you can reconcile spend against your own logs.

Frequently asked questions

Is an LLM API gateway the same as a proxy?+

A proxy only forwards traffic. A gateway authenticates, routes by model, normalises protocols and meters usage, which is why it can produce a single bill and per-request cost detail.

Does a gateway add latency?+

It adds one network hop. For most workloads that is small relative to model inference time, but measure it on your own path if you are optimising a latency-sensitive surface.

Is an OpenAI-compatible endpoint the same as OpenAI?+

No. Compatibility describes the request shape, not the model behind it. A compatible endpoint can serve Anthropic, Google, DeepSeek or open-weight models while your code stays unchanged.

What does "billed at upstream list price" mean?+

You pay the published per-token rate of the model you called, rather than a rounded or marked-up rate, and the console shows per-request detail so the number is checkable.

What is the difference between a gateway and a model router?+

They overlap. A router chooses between models, often automatically; a gateway is the single authenticated endpoint, key and billing surface that the routing happens behind. Many products do both.

One key covers every model in this guide

Create an account, pick a plan and copy an API key. Gateway plans start at $19 per month and model usage is billed at upstream list prices.

What is an LLM API gateway? Definition, use cases and pricing