Queueing & Load Shedding
A per-workspace switch. When the only compatible agent is saturated, a request either waits for capacity or is refused immediately with 429. Queueing is on by default.
Overview
A saturated agent has no free slot to admit another request. The router can
respond in one of two ways: park the request until a slot frees, or refuse
it and let the caller decide what to do. Which one it does is a property of
the workspace the request attributes to, held on the
inference_queueing_enabled flag.
The default is true: queue. A workspace that has never touched
the setting queues, and so does a request whose workspace cannot be read at
all -- the safe default for load shedding is to keep completing requests
that succeed today rather than to start refusing them.
This governs saturation only. Per-key rate limits, quota exhaustion and billing refusals are separate and are unaffected; see Error Handling.
The Two Behaviours
| Setting | A request that finds its agent saturated |
|---|---|
inference_queueing_enabled: true (default) |
Waits for capacity. A burst usually still completes, at the cost of latency. If the wait times out or the queue is full, the request fails with 503 and a Retry-After header. |
inference_queueing_enabled: false |
Refused immediately with 429 and a Retry-After header. Nothing waits. The caller retries or routes elsewhere. |
How a Request Waits
A waiting request is dispatched as soon as a slot frees on an agent that can
serve it, and holds no agent's slot while it waits. The wait is bounded: the
queue timeout, 30 seconds by default, counts from the moment the request
first starts waiting, however many times it is woken. When the timeout
passes, the request fails with
503 and a Retry-After header. When the queue for
the agent is already full, it fails the same way at once, without waiting.
A streamed request that waits sends nothing until it is dispatched -- not
even its status line -- so the bounded wait sits well inside common proxy
idle timeouts. If it times out or the queue is full, it is refused with
503, a Retry-After header and a JSON error body,
exactly as a request with stream: false is, and no event
stream: an SDK retries it as it retries any 503. Once
dispatched, the stream starts as usual.
A turn that dispatches more than once -- one that runs tools, for example -- can wait again after its stream has started. That wait is kept alive with SSE comments, which conforming clients ignore, and a timeout then arrives as an in-stream error event, like any failure after the stream starts. See Streaming Errors.
The Refusal
HTTP/1.1 429 Too Many Requests
Retry-After: 10
Content-Type: application/json
{
"error": {
"message": "...",
"type": "rate_limit_error",
"code": "capacity_exceeded"
}
}
Branch on the status and code. message describes
the capacity condition in prose and is not part of the contract.
429 rather than 503 is deliberate. A workspace with
queueing off is shedding load on purpose, not failing: the request was
understood, capacity was checked, and the policy chose to decline. A
provider monitor counts 5xx against provider uptime and excludes
429, so answering 503 to a request you deliberately
shed would score your own choice as an outage.
capacity_exceeded is the same code a transiently overloaded
backend returns, so a client needs no special handling: wait for
Retry-After and retry, with exponential backoff and jitter
across successive refusals.
Which Requests It Governs
Every caller-facing dispatch path reads the setting: chat completions,
responses, embeddings, rerank, score, and audio transcription,
translation and speech. It applies on every base URL those endpoints are
served on except direct model dispatch, which
never waits: a request there that finds its agent saturated is refused
immediately with 429 and a Retry-After header,
whatever the setting.
The realtime transcription WebSocket
does not read the setting and never waits: a turn that finds no agent
free to take it fails at once with an error event whose
code is capacity_exceeded, the code the HTTP
endpoints answer the same refusal with, and the session stays open for
the next utterance.
The setting is read from the workspace the request attributes to. On a workspace-scoped base URL that is the workspace named in the path; on a project-scoped one it is the project's default workspace. Requests that attribute to a different workspace are unaffected by this workspace's setting.
Inference the router performs for itself -- the internal steps behind a routed turn -- keeps queueing regardless. The switch governs what the router owes an external caller, not how it schedules its own work. The one exception is direct model dispatch, where nothing waits: an internal step of a direct request that finds its agent saturated fails at once instead.
Choosing a Setting
Leave it on when completion matters more than latency: batch-shaped work, background jobs, an internal tool where a slower answer beats an error the caller has to handle.
Turn it off when:
-
The caller has somewhere else to go. A client with a fallback provider
acts on a fast
429; it cannot act on a request that is silently parked. - You are enforcing a latency budget. A refusal inside the budget is more useful than an answer outside it.
Marketplace traffic needs neither. A
marketplace listing is served through
direct model dispatch, which never waits, so a monitor measures the fleet
rather than a queue, and a saturated agent answers 429, which a
provider monitor excludes from uptime, rather than a timed-out
503, which it counts. The listing has to be readable as well:
the manifest answers 401 to a fetch carrying no API key until
the workspace its URL binds
publishes it.
Reading and Setting It
In the chat workspace settings panel, the Queue when the agent is busy card holds the switch. Only a workspace owner or manager sees a live switch; everyone else sees the current state read-only.
Over the API, the flag is on the workspace document:
{
"id": "ws_abc123",
"name": "Marketplace",
"inference_queueing_enabled": false,
"zero_data_retention": false
}
curl -X PATCH "https://api.erebine.ai/proj_ABC123/v1/workspaces/ws_abc123" \
-H "Authorization: Bearer $EREBINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"inference_queueing_enabled": false}'
zero_data_retention on the same document is the unrelated
retention mode. The two settings sit
together and do nothing to each other.