docs
// Platform

Queueing & Load Shedding

A per-workspace switch. When the only compatible agent is saturated, a request either waits for capacity or is refused immediately with 429. Queueing is on by default.

Overview

A saturated agent has no free slot to admit another request. The router can respond in one of two ways: park the request until a slot frees, or refuse it and let the caller decide what to do. Which one it does is a property of the workspace the request attributes to, held on the inference_queueing_enabled flag.

The default is true: queue. A workspace that has never touched the setting queues, and so does a request whose workspace cannot be read at all -- the safe default for load shedding is to keep completing requests that succeed today rather than to start refusing them.

This governs saturation only. Per-key rate limits, quota exhaustion and billing refusals are separate and are unaffected; see Error Handling.

The Two Behaviours

Setting A request that finds its agent saturated
inference_queueing_enabled: true (default) Waits for capacity. A burst usually still completes, at the cost of latency. If the wait times out or the queue is full, the request fails with 503 and a Retry-After header.
inference_queueing_enabled: false Refused immediately with 429 and a Retry-After header. Nothing waits. The caller retries or routes elsewhere.

How a Request Waits

A waiting request is dispatched as soon as a slot frees on an agent that can serve it, and holds no agent's slot while it waits. The wait is bounded: the queue timeout, 30 seconds by default, counts from the moment the request first starts waiting, however many times it is woken. When the timeout passes, the request fails with 503 and a Retry-After header. When the queue for the agent is already full, it fails the same way at once, without waiting.

A streamed request that waits sends nothing until it is dispatched -- not even its status line -- so the bounded wait sits well inside common proxy idle timeouts. If it times out or the queue is full, it is refused with 503, a Retry-After header and a JSON error body, exactly as a request with stream: false is, and no event stream: an SDK retries it as it retries any 503. Once dispatched, the stream starts as usual.

A turn that dispatches more than once -- one that runs tools, for example -- can wait again after its stream has started. That wait is kept alive with SSE comments, which conforming clients ignore, and a timeout then arrives as an in-stream error event, like any failure after the stream starts. See Streaming Errors.

The Refusal

HTTP 429
HTTP/1.1 429 Too Many Requests Retry-After: 10 Content-Type: application/json { "error": { "message": "...", "type": "rate_limit_error", "code": "capacity_exceeded" } }

Branch on the status and code. message describes the capacity condition in prose and is not part of the contract.

429 rather than 503 is deliberate. A workspace with queueing off is shedding load on purpose, not failing: the request was understood, capacity was checked, and the policy chose to decline. A provider monitor counts 5xx against provider uptime and excludes 429, so answering 503 to a request you deliberately shed would score your own choice as an outage.

capacity_exceeded is the same code a transiently overloaded backend returns, so a client needs no special handling: wait for Retry-After and retry, with exponential backoff and jitter across successive refusals.

Which Requests It Governs

Every caller-facing dispatch path reads the setting: chat completions, responses, embeddings, rerank, score, and audio transcription, translation and speech. It applies on every base URL those endpoints are served on except direct model dispatch, which never waits: a request there that finds its agent saturated is refused immediately with 429 and a Retry-After header, whatever the setting.

The realtime transcription WebSocket does not read the setting and never waits: a turn that finds no agent free to take it fails at once with an error event whose code is capacity_exceeded, the code the HTTP endpoints answer the same refusal with, and the session stays open for the next utterance.

The setting is read from the workspace the request attributes to. On a workspace-scoped base URL that is the workspace named in the path; on a project-scoped one it is the project's default workspace. Requests that attribute to a different workspace are unaffected by this workspace's setting.

Inference the router performs for itself -- the internal steps behind a routed turn -- keeps queueing regardless. The switch governs what the router owes an external caller, not how it schedules its own work. The one exception is direct model dispatch, where nothing waits: an internal step of a direct request that finds its agent saturated fails at once instead.

Choosing a Setting

Leave it on when completion matters more than latency: batch-shaped work, background jobs, an internal tool where a slower answer beats an error the caller has to handle.

Turn it off when:

  • The caller has somewhere else to go. A client with a fallback provider acts on a fast 429; it cannot act on a request that is silently parked.
  • You are enforcing a latency budget. A refusal inside the budget is more useful than an answer outside it.

Marketplace traffic needs neither. A marketplace listing is served through direct model dispatch, which never waits, so a monitor measures the fleet rather than a queue, and a saturated agent answers 429, which a provider monitor excludes from uptime, rather than a timed-out 503, which it counts. The listing has to be readable as well: the manifest answers 401 to a fetch carrying no API key until the workspace its URL binds publishes it.

Reading and Setting It

In the chat workspace settings panel, the Queue when the agent is busy card holds the switch. Only a workspace owner or manager sees a live switch; everyone else sees the current state read-only.

Over the API, the flag is on the workspace document:

GET /proj_ABC123/v1/workspaces/{workspace_id}
{ "id": "ws_abc123", "name": "Marketplace", "inference_queueing_enabled": false, "zero_data_retention": false }
cURL
curl -X PATCH "https://api.erebine.ai/proj_ABC123/v1/workspaces/ws_abc123" \ -H "Authorization: Bearer $EREBINE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"inference_queueing_enabled": false}'

zero_data_retention on the same document is the unrelated retention mode. The two settings sit together and do nothing to each other.