Direct Model Dispatch
One base URL for every model in the project. The model travels in the request payload instead of the path, so an SDK client is instantiated once and changes the model argument per call.
Overview
The endpoint-scoped base URL binds one client to one deployment:
/proj_ABC123/ENDPOINT_SLUG/v1 names the deployment in the
path. Serving several models from one process means several clients, or a
base URL rebuilt per request.
_direct removes the slug. The request body's model
field carries a namespaced anchor id, the router resolves it to one of the
project's deployments, and dispatch proceeds through the ordinary
per-endpoint pipeline from there. Request and response bodies are the
endpoint-scoped route's, unchanged.
_direct is not the
semantic router. The
semantic router picks a model for you and reports its choice; here the
caller names the model and no routing decision is made.
Base URLs
Two forms are registered. Every inference path is served on both.
https://api.erebine.ai/proj_ABC123/WORKSPACE_ID/_direct/v1
https://api.erebine.ai/proj_ABC123/_direct/v1
WORKSPACE_ID is the ws_-prefixed workspace id.
Only the prefixed form attributes usage to that workspace; the bare form
binds the project's default workspace. Usage attribution and the retention
mode follow the workspace the request attributes to, so the prefixed form is
the one to configure in a client.
from openai import OpenAI
client = OpenAI(
base_url="https://api.erebine.ai/proj_ABC123/ws_abc123/_direct/v1",
api_key="ere_yourproject_yourkey",
)
for model_id in ("erebine/fast-chat", "erebine/long-context"):
reply = client.chat.completions.create(
model=model_id,
messages=[{"role": "user", "content": "Say ok"}],
)
print(model_id, reply.choices[0].message.content)
Custom Domains
A project can serve this surface, and every endpoint URL, on a hostname it controls. The domain belongs to the project, so one DNS setup covers every endpoint in it, and a certificate for the hostname is issued once the domain is verified. Only the project owner can add one, in project Settings > Custom Domains. Adding the domain shows the two records to publish. Once it is verified, each workspace's Connect card lists its base URLs on the new hostname.
| Record | Name | Value | Purpose |
|---|---|---|---|
TXT |
_erebine-challenge.api.acme.com |
erebine-domain-verification=TOKEN |
Required. Proves you control the zone; the domain verifies on this record alone. |
CNAME |
api.acme.com |
api.erebine.ai |
Routes traffic. Without it the domain verifies but nothing reaches it. |
The TXT record is checked every few minutes. A domain still unverified 72 hours after it was added is marked failed; Recheck starts it again. Once verified, the paths this page documents are served with the project id dropped, because the host supplies it:
https://api.acme.com/v1
https://api.acme.com/WORKSPACE_ID/_direct/v1
https://api.acme.com/ENDPOINT_SLUG/v1
https://api.acme.com/_direct/v1 is the first form spelled out.
The project-level paths your users call drop the project id the same way:
/v1/files on the custom domain is the project's files, and so are
/v1/uploads, /v1/batches,
/v1/conversations, /v1/endpoints,
/v1/slos and /v1/webhooks.
/v1/models lists the models this surface accepts.
The API is authoritative about that list: a path a custom domain does
not serve answers 404 and names the host to call instead.
-
An apex name (
acme.com) cannot carry aCNAME. Use a subdomain, anALIASorANAMErecord if your DNS provider offers one, orAandAAAArecords pointing at the addressesapi.erebine.airesolves to. -
The realtime WebSocket API is not served on a custom domain: a WebSocket
handshake there is refused. Connect to it on
api.erebine.ai. A plain HTTP request to a realtime path on the custom domain answers404realtime_unavailable_on_custom_domain. -
Only the API is served on the custom domain. The dashboard stays on its
own address, and so do the account and project management APIs: keys,
usage, settings, members, exports and the agent surfaces are called on
api.erebine.aiwith the project id spelled out. On the custom domain they answer404surface_unavailable_on_custom_domain, naming the surface and the address to call it on.
Listing the Models
GET /proj_ABC123/WORKSPACE_ID/_direct/v1/models returns the
ids this surface accepts, in the OpenAI list shape. Send one verbatim as the
model field.
{
"object": "list",
"data": [
{
"id": "erebine/fast-chat",
"object": "model",
"created": 1784518031,
"owned_by": "yourproject",
"name": "fast-chat"
}
]
}
Ids here are namespaced anchor slugs (erebine/<slug>), not
the model UUIDs the other listings emit. An id taken from another listing
does not resolve on _direct.
name is an Erebine extension. OpenAI's model object is
id, object, created,
owned_by. The official Python and Node clients both pass the
extra key through rather than rejecting the response, but do not write code
that requires it against other providers.
GET .../_direct/v1/models/{id} retrieves one entry, so
models.retrieve() works against this base URL alongside
models.list(). Send an id this listing emitted. The raw
erebine/<slug> spelling, the percent-encoded
erebine%2F<slug> an OpenAI SDK puts on the wire, and the
bare slug all resolve to the same entry, and the response carries the
namespaced id in every case. An id that resolves to nothing
answers 404 model_not_found, which is the status
the OpenAI specification defines for a miss.
Four base URLs emit the OpenAI model list. They differ in what id
means and in which other SDK calls they answer:
| Base URL | id is |
models.list() |
models.retrieve() |
Inference |
|---|---|---|---|---|
/proj_ABC123/ENDPOINT_SLUG/v1 |
Model UUID | Yes | Yes | Yes |
/proj_ABC123/_direct/v1 |
erebine/<slug> |
Yes | Yes | Yes |
/proj_ABC123/_semantic/router/v1 |
Model UUID | Yes | Yes | Yes |
/proj_ABC123/v1 |
Model UUID | Yes | Yes | No, 404 |
Point an OpenAI SDK at the endpoint-scoped or _direct base and
one client answers models.list(),
models.retrieve() and the inference calls that base serves. The
semantic-routed base serves the listing, retrieve-one and inference, so
models.retrieve() resolves there against the model UUIDs its
listing emits. The fourth is the project management listing, not an
inference base URL.
What the Surface Serves
Paths are shown relative to the base URL. Every one of them is registered on both forms. On the provider manifest the form chosen also decides which workspace governs the document; see Marketplace Listing.
/chat/completionsChat completions, streaming included
POST/responsesResponses API
GET/responsesList stored responses
GET/responses/{id}Retrieve a stored response
GET/responses/{id}/input_itemsInput items of a stored response
POST/responses/{id}/cancelCancel a run
DELETE/responses/{id}Delete a stored response
POST/embeddingsEmbeddings
POST/rerankRerank a candidate list
POST/scoreScore query-document pairs
POST/audio/speechText to speech, raw audio bytes
POST/audio/transcriptionsSpeech to text, multipart
POST/audio/translationsSpeech to English text, multipart
GET/modelsIds this surface accepts
GET/provider-manifestMarketplace listing document
The two audio upload paths read model out of the multipart body
alongside the clip, which is how an OpenAI audio client already sends it.
/audio/speech answers raw encoded audio with the matching
Content-Type, not a JSON envelope.
Multimodal input works here as it does on the endpoint-scoped route: image,
video and audio content parts on
chat completions, subject
to what the addressed model declares. Realtime voice over
/v1/realtime is not on this list: that WebSocket is mounted on
the endpoint-scoped base URL, /proj_ABC123/ENDPOINT_SLUG/v1/realtime.
See Realtime.
Addressing Rules
| You send | Result |
|---|---|
"model": "erebine/fast-chat" |
Resolves. The response model field echoes the string you sent, character for character. |
"model": "fast-chat" |
Resolves. The namespace is optional on input; the response echoes the bare spelling back, not a rewritten one. |
"model": "openai/fast-chat" |
404 model_not_found. A foreign namespace is refused rather than ignored. |
| An id from any other listing | 404 model_not_found. The other listings emit model UUIDs; this one accepts anchor slugs. |
An embeddings model on /chat/completions |
404 model_not_found. Resolution is gated on the task the addressed path performs, so a mismatch is refused at the door rather than failing inside the worker. |
| A model this project deploys that is still loading, or is draining or suspended | 429 capacity_exceeded with a Retry-After header. The model exists; it has no capacity serving it right now. |
A model appears in the listing only once its deployment is actually serving.
A deployment that exists but is still pulling weights or loading is not
advertised, because an advertised id that dispatch then refuses is worse
than a shorter list. Dispatch answers it with 429 rather than
404, so a client that already holds the id retries it.
Errors
A model you cannot use on the path you called answers 404
model_not_found with the body the OpenAI API sends, so an OpenAI
client raises NotFoundError for it on dispatch and on
GET /models/{id} alike. That is the status the OpenAI
specification defines for a missing model, and it is served in preference to
scoring well on someone else's uptime metric. A path this surface does not
serve answers 404 too.
{
"error": {
"message": "The model `erebine/no-such-model` does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": null,
"code": "model_not_found"
}
}
Every other refusal answers 400 or 429 -- never
402, never 5xx. Provider monitors compute uptime as
successes over total and count 401, 402,
404 and 5xx against the provider while excluding
400, 403, 413 and 429, so
a malformed request or a load condition must not read as an outage.
| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request |
Missing or empty model, or a malformed body. The error names model in param when that is the missing field. |
| 400 | retention_not_permitted |
The bound workspace runs in zero data retention and the request asked the router to store something. |
| 404 | route_not_found |
The path is not one this surface serves. |
| 404 | model_not_found |
The one addressing refusal, on dispatch and on GET /models/{id}. Unknown anchor, foreign namespace, an id from another listing, a model not deployed in this project, or a model addressed on a path its task mode does not serve. param is null and the message names the id you sent. |
| 429 | capacity_exceeded |
No capacity. Returned immediately: this surface never waits for a saturated agent, whatever the workspace's queueing setting. Retry after the delay in the Retry-After header. A model this project deploys that is still loading, or is draining or suspended, answers this too, and so does one whose agents are all unhealthy, offline or drained. |
| 429 | listing_concurrency_exceeded |
The model already has its published concurrency of requests in flight, counted across every caller. A Responses request with background: true counts until it completes, fails or is cancelled, not only until its 202 returns. Returned immediately, whatever the queueing setting; retry after the delay in the Retry-After header. |
401 is not on this table because it is not a dispatch answer. An
anonymous fetch of the provider manifest returns it unless the workspace the
URL binds has published the document -- the workspace named in the path, or
the project's default workspace when the path names none. See
Serving It Without a Key.
An unknown project id does not leak its non-existence: the listing and the
manifest answer an empty 200 document, and dispatch and
GET /models/{id} answer the same 404 a served project
gives for an unknown model.
Response Headers
| Header | Present | Carries |
|---|---|---|
X-Request-ID |
Always | Correlation id for the request. Quote it in support requests. |
X-Erebine-Worker-ID |
Always | The agent that served the request. |
X-Erebine-Routed-Model |
Never | Reports a routing decision. On _direct the caller named the model, so there is no decision to report. |
X-Erebine-Routed-Endpoint |
Never | As above. |
Marketplace Listing
The provider manifest publishes the same inventory the models listing advertises, as the catalogue document a model marketplace polls: identity, pricing, declared capacity, the datacenters the serving agents sit in, and the workspace's retention claim. Only ids that are actually serving appear, so the manifest never names an id the listing does not. An entry the operator has neither priced nor listed as free is left out of the manifest; the listing still serves it to your own callers.
GET /proj_ABC123/WORKSPACE_ID/_direct/v1/provider-manifest
GET /proj_ABC123/_direct/v1/provider-manifest
Both forms are served and both return the same inventory. They differ in
which workspace governs the document: the workspace-scoped form binds the
workspace named in the path, and the project-scoped form binds the project's
default workspace. That binding fixes compliance.zdr on every
entry and decides whether an anonymous fetch is served at all.
It is a separate route from /models on purpose. Folding a
marketplace document into the OpenAI list shape would break every SDK client
pointed at the same base URL, so the two documents stay apart and
/models keeps the OpenAI shape.
The document declares its own schema_version, at the root and
again on every entry. Both read 2.4 today. Read the field rather
than assuming a version: the format tracks the marketplace schema it is
published against, and it will change when that schema does.
Document Shape
The root carries the version and a data array. Each element is
one model document, and pricing and capacity live inside the modality they
belong to rather than at the root.
{
"schema_version": "2.4",
"data": [
{
"schema_version": "2.4",
"id": "erebine/fast-chat",
"name": "fast-chat",
"created": 1784518031,
"quantization": "bf16",
"hugging_face_id": "meta-llama/Llama-3.1-8B-Instruct",
"openrouter": { "slug": "erebine/fast-chat" },
"input_modalities": [
{
"type": "text",
"supported_inputs": { "max_context_length": { "value": 128000 } },
"pricing": [
{ "type": "prompt", "unit": "token", "cost_usd": "0.000001000000" },
{ "type": "cached_prompt", "unit": "token", "cost_usd": "0.000000250000" }
]
}
],
"output_modalities": [
{
"type": "text",
"streaming": true,
"max_length": { "value": 4096 },
"supported_parameters": {
"tools": { "type": "boolean" },
"reasoning": { "type": "boolean" },
"structured_outputs": { "type": "boolean" }
},
"pricing": [
{ "type": "completion", "unit": "token", "cost_usd": "0.000001000000" }
]
}
],
"capacity": [
{ "type": "concurrency", "unit": "request", "value": 48 }
],
"datacenters": [ { "country_code": "US" } ],
"compliance": { "zdr": false },
"is_ready": true,
"is_free": false,
"discount_to_user": 0
}
]
}
| Field | Type | Carries |
|---|---|---|
id |
string | The same namespaced anchor id /models advertises. Send it verbatim as model. |
name |
string | The bare slug, for display. |
created |
integer | Unix seconds. Dates the model record, and matches the model's created on the project management listing. |
quantization |
string or null | Weight precision the deployment is serving, e.g. bf16. null when the deployment does not declare one. |
hugging_face_id |
string | The upstream Hugging Face repository id, when the model carries one. Detected from config.json's _name_or_path on import and settable by an operator; absent when neither happened. Never guessed from a file or directory name. |
openrouter.slug |
string | The slug a marketplace lists the entry under: the marketplace's own catalog slug when an operator has recorded one, otherwise the entry's id. |
input_modalities |
array | At least one. Each names a type (text, image, audio, video, file) and owns its pricing. The text entry carries supported_inputs.max_context_length. |
output_modalities |
array | At least one. Each names a type (text, speech, embeddings, rerank among them), owns its pricing, and declares streaming and the supported_parameters that modality accepts. |
max_length |
object | The largest output this modality will produce, wrapped as { "value": <n> } like every other length limit in the schema. Present only on the output branches whose schema declares it. |
capacity |
array | Declared throughput. A concurrency entry is the number of requests the model admits at once, counted across every caller, and carries no window. It is enforced: one request over it is refused with 429 listing_concurrency_exceeded. A background: true Responses request counts toward it until the response completes, fails or is cancelled. It follows the agents serving the model and the size of the model's recent requests, so it changes as agents come and go and as request sizes shift, and it is omitted while no limit is known for any serving agent. |
datacenters |
array | One entry per location the serving agents sit in, each an ISO 3166-1 alpha-2 country_code and an optional operator region label. |
deployment_region |
string | The region the serving deployment runs in, as an operator records it. Display only, and absent when none is recorded. |
compliance.zdr |
boolean | Derived from the bound workspace's retention mode, not set independently. |
is_ready |
boolean | Whether the entry is servable. An id is published only once its deployment is serving, so this is true on every entry the document carries. |
is_free |
boolean | Whether the operator lists the entry as free. Free is always an explicit choice; a free entry publishes zero prices. |
discount_to_user |
number | Fraction taken off the published price, below 1. 0 is no discount. |
supported_parameters maps each request parameter the modality
accepts to a descriptor of the values it takes; a parameter that is absent is
not supported. A model that calls tools declares tools,
tool_choice and parallel_tool_calls. Where the model
does not support a forced tool choice, tool_choice is
{ "type": "enum", "values": ["auto", "none"] }, and a request that
sends "required" or a named function is refused with
400 unsupported_parameter.
Pricing is per modality. An input modality prices prompt and
cached_prompt; an output modality prices
completion. cost_usd is USD per unit
and is a string, not a number, so a client parses it into a
decimal type and loses no precision to a float. unit is
token on every text entry.
The document describes one workspace's compliance posture: the workspace the
URL binds. Each entry's compliance.zdr is derived from that
workspace's retention mode rather
than being a flag an operator sets, so the published claim and the enforced
behaviour cannot drift apart. A workspace-scoped fetch publishes that
workspace's retention posture; a project-scoped fetch publishes the default
workspace's.
Give a marketplace the workspace-scoped form. It states in the URL which workspace it returns, and it does not change meaning if the project's default workspace is later reassigned. Both forms stay supported; this is a recommendation, not a requirement.
Two fetches of an unchanged inventory are byte-identical, so a monitor that diffs the document does not read a reordering as a live change.
Serving It Without a Key
A marketplace cannot present a credential. Its application form takes a URL and nothing else, and the poller that keeps the listing current runs on the same terms. A document behind an API key is a document that never gets read, and a listing that never gets updated.
marketplace_manifest_public is the per-workspace switch that
settles it. It defaults to false. Turned on, the manifest is
served at the URL it already has to a request carrying no key. Turned off, an
anonymous fetch of it is refused. A request that does carry a key is
unaffected either way.
The switch is read from the workspace the URL binds: the workspace named in
the path, or the project's default workspace when the path names none. The
router resolves the workspace by path prefix first, then the
X-Workspace-Id header, then the project default; a marketplace
poller sends a URL and no headers, so here the URL is what decides. The two
rows below that name marketplace_manifest_public read that
workspace's value.
| Request | Answer |
|---|---|
| With an API key, either setting | The manifest. An authenticated caller is not gated on this switch. |
No key, marketplace_manifest_public: true |
The manifest. The same document an authenticated fetch gets. |
No key, marketplace_manifest_public: false (default) |
401. |
| No key, unknown project id | An empty 200 document, exactly as an authenticated fetch gets. The router does not confirm or deny that a project exists. |
The switch and the published compliance.zdr read the same
binding, so turning the switch on in a workspace the URL does not bind
publishes nothing.
Give a marketplace the workspace-scoped URL. It states which workspace it returns, and it does not change meaning if the project's default workspace is later reassigned. The project-scoped URL is served on the same terms, governed by the default workspace's switch. That is the whole call:
curl "https://api.erebine.ai/proj_ABC123/ws_abc123/_direct/v1/provider-manifest"
# governed by the project's default workspace instead
curl "https://api.erebine.ai/proj_ABC123/_direct/v1/provider-manifest"
What Becomes Public
Exactly the document described above and
nothing more: model ids and names, context lengths, prices, declared
capacity, datacenter country and region, quantization, and the
compliance.zdr flag. It is a price list.
No chat content, no keys, no workspace data, no user data. The switch governs
this one document and grants nothing beyond it.
It is off by default because publishing is a decision, not a consequence. A project id is not a secret -- it sits in every base URL and every client config -- so reading anything addressed by one without a key has to be asked for rather than assumed.
Workspace Policy
Two per-workspace settings change what this surface does. Each follows the workspace the URL names, and the project's default workspace when the URL names none.
-
Zero data retention refuses any
request that asks the router to store something -- an explicit
store: true, orbackground: trueon Responses. - Manifest publication decides whether the provider manifest is readable without an API key. It is off by default.
Queueing is not one of them. This surface never
waits for capacity: a request that finds its agent saturated is refused at
once with 429 capacity_exceeded and a
Retry-After header, whatever the workspace's setting, so a
monitor measures the fleet rather than a queue. Neither does any step the
router runs on the request's behalf, such as a research run or workspace
analysis a tool call starts: a saturated agent fails that step at once. The
setting still governs the workspace's other base URLs.