Sketio

Cloudflare AI Gateway Architecture for LLM Apps

Updated

AI Gateway is a proxy that sits between your application and the AI providers it calls. You send requests through a gateway, and the gateway adds what you would otherwise build yourself: a cache for repeated requests, rate limits, logs and cost analytics, and a fallback to another provider when a call fails. Workers AI, OpenAI, Anthropic, Google Gemini and other providers work with it, so one place can see and control all of an app's model traffic.

This template shows a Worker that calls a model through a gateway. Workers AI is the first choice, and an external provider is the fallback.

requestvia gatewayClientWorkerAI GatewayRate limitingCacheLogs and analyticsWorkers AIExternal providerrecordsfirst choicefallback

Scroll sideways to see the whole diagram

Cloudflare AI Gateway Architecture for LLM Apps. Open it in Sketio to change it.

Start from this diagram and edit it on your own board.

By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.

What each part does

Client
The browser or app of the person using your AI feature. It talks to your Worker and never to the gateway or a provider.
Worker
Your application code. For a Workers AI model it calls env.AI.run(model, input, { gateway: { id } }) with the gateway's ID. For other providers it sends the provider's own request format to the gateway URL, which starts with https://gateway.ai.cloudflare.com/v1/ and continues with your account ID, the gateway ID and the provider name.
AI Gateway
The proxy. A gateway can be created in the dashboard, or for Workers AI requests it is created on the first authenticated request under the name default. Settings such as caching, rate limiting and logging are set per gateway, and some can be overridden per request with cf-aig-* headers.
Cache
Answers identical requests without calling the provider. It is off by default and is turned on in the gateway's settings. By default the key is a hash of the provider, endpoint, model, authorization header and the whole request body, so any change to the prompt or parameters is a new entry. The cf-aig-cache-status header says HIT or MISS. The cache is volatile, and identical requests that arrive at the same moment may both go to the provider.
Rate limiting
Limits how many requests a gateway accepts in a time window, using a fixed window or a sliding window. A request over the limit gets a 429 Too Many Requests response and is not processed. The limit applies to all requests through that gateway.
Logs and analytics
Logs are on by default for each gateway and can include the prompt, the response, the provider, status, token usage, cost and duration. The cf-aig-collect-log: false header skips a log entry, and cf-aig-collect-log-payload: false keeps the metadata but not the bodies. Analytics shows requests, tokens, cost and errors.
Workers AI
Models that run on Cloudflare's network. In this diagram it is the first choice for each request.
External provider
A model from a provider such as OpenAI or Anthropic. In this diagram the gateway uses it when the first choice returns an error.

How a request flows

  1. The client asks your Worker for something, such as an answer or a summary.
  2. The Worker sends the model call through the gateway: with the gateway option on env.AI.run for Workers AI models, or by calling the gateway URL for an external provider.
  3. The gateway applies its settings to the request. One over the rate limit gets a 429 and goes no further. With caching on, an identical earlier request is answered from the cache, with no call to the provider.
  4. Otherwise the request goes to the first provider. To get a fallback, send the Universal endpoint an array of provider objects: Cloudflare tries them in order, and moves to the next when a request returns an error. The cf-aig-step response header tells which step answered, with 0 for the first.
  5. The response goes back to the Worker and the client, and the gateway records the log entry and the analytics: tokens, cost and status.

When to use it

Common variations

Add Guardrails

Guardrails check prompts and responses for harmful content and can flag or block them, across providers. They add latency to each request, and the documentation says streaming (stream: true) requests are not supported: on the gateway endpoints, Guardrails buffers the whole response, evaluates it and returns it as a single non-streamed payload.

Cap spend and tag requests

Spend limits set dollar budgets by model, provider or custom metadata and block requests once a budget is used up. Custom metadata, such as a user or team ID sent with cf-aig-metadata, lets you filter logs and analytics by it.

Keep provider keys out of the app

Bring Your Own Keys stores provider API keys in Cloudflare instead of in your code, and gateway authentication restricts who may call the gateway at all.

Route by rules instead of a fixed order

Dynamic routing builds request flows with if/else conditions on the request body, headers or metadata (for example sending paid and free users to different models), percentage splits for A/B tests and gradual rollouts, and rate or budget limits that switch to a fallback, instead of one fixed first choice and fallback.

Retry before falling back

Per-request headers set the retry policy: cf-aig-max-attempts (up to 5), cf-aig-retry-delay in milliseconds and cf-aig-backoff, which is constant, linear or exponential.

Make it yours

Rename the gateway, pick the first-choice and fallback models, and set the cache lifetime and rate limit to your traffic. Caching is off by default, so turn it on only for prompts whose answers can be reused.

Opens this diagram as a board you can edit.

By continuing, you agree to the Terms of Service and Privacy Policy, including sending images of your strokes, diagram labels and similar data to providers in the United States (Cloudflare, Inc. and TypeSafe AI, Inc.) for AI conversion.

All templates