Model Management
List, update, share, revalidate, and delete every custom model in your project from one REST surface. Mounted on the inference route group, gated by the same project API key you already use for chat completions.
Note: Model Management API endpoints use the path /proj_ABC123/v1/models. Mutations (PATCH, DELETE, POST share/unshare/revalidate/versions) run on the same inference route group as /v1/chat/completions, gated by standard project API-key authentication. There is no separate "Model Management" scope; any valid project API key is accepted. See the Endpoint scope table below.
The Model Management API allows you to:
- List all models in your project
- Retrieve details for a specific model
- Update model metadata (description, architecture, quantization, and more)
- Set a model's speculative decoding: method, token count, and a draft model from the catalog
- Delete models you no longer need
- Trigger model revalidation to re-extract metadata
- Share a model to the public catalog or unshare it
- List, create, promote, and roll back model versions
- View catalog information for shared models
For uploading models, see the Model Upload guide. For versioning details, see Model Versioning.
Base URL
Every endpoint below is mounted under one project-scoped base URL. Substitute your project external id and your API key once; the remaining sections show only the path and the body.
https://api.erebine.ai/proj_ABC123/v1
Authorization: Bearer ere_myproject_your_api_key
The external id has the shape ws_<hex> for workspaces and is visible in the project header. The API key format is ere_<project>_<secret> and is issued from the API Keys page.
List Project Models
GET/proj_ABC123/v1/models
Lists all models uploaded to your project. The default response is the OpenAI-compatible list envelope. Pass ?extended=true to receive Erebine-specific metadata fields (format, size, status, architecture, version, sharing, and more) along with pagination counters.
limit and offset are Erebine extensions; OpenAI's
models listing defines neither. Omit limit and the standard
envelope returns the whole inventory unpaginated, which is what an SDK
enumerating models expects. An explicit limit is bounded to
100; a larger value is served as 100. The ?extended=true
envelope defaults to 100 rows and reports the truncation on
has_more and total. On THIS listing a negative
limit or offset answers 400
invalid_request with param naming the one at
fault, rather than being clamped silently -- see
Error Responses. The inference-facing listings
(an endpoint-scoped base, _direct and
_semantic/router) do not refuse a negative bound; they
ignore it and serve the full visible inventory.
curl https://api.erebine.ai/proj_ABC123/v1/models \
-H "Authorization: Bearer ere_myproject_your_api_key"
import requests
headers = {"Authorization": "Bearer ere_myproject_your_api_key"}
response = requests.get(
"https://api.erebine.ai/proj_ABC123/v1/models",
headers=headers
)
for model in response.json()["data"]:
print(f"{model['id']} ({model.get('name', '')}) owned_by={model['owned_by']}")
const response = await fetch(
"https://api.erebine.ai/proj_ABC123/v1/models",
{
headers: {
"Authorization": "Bearer ere_myproject_your_api_key"
}
}
);
const payload = await response.json();
for (const model of payload.data) {
console.log(`${model.id} (${model.name ?? ""}) owned_by=${model.owned_by}`);
}
Response (standard)
{
"object": "list",
"data": [
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"created": 1706123456,
"owned_by": "my-project",
"name": "my-custom-model"
}
]
}
Response (extended)
Request with ?extended=true. Optional fields are omitted from the JSON body when unset (the encoder does not emit explicit null).
{
"object": "list",
"data": [
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"created": 1706123456,
"owned_by": "my-project",
"name": "my-custom-model",
"format": "safetensors",
"size_bytes": 4294967296,
"size_formatted": "4.00 GB",
"status": "ready",
"version": "1.0.0",
"is_latest": true,
"architecture": "LlamaForCausalLM",
"quantization": "awq",
"context_length": 32768,
"parameter_count": 7000000000,
"workload_type": "chat",
"is_shared": false
}
],
"has_more": false,
"total": 1
}
Get Model Details
GET/proj_ABC123/v1/models/{modelId}
Retrieves detailed information about a specific model. Pass ?extended=true to include additional metadata such as format, size, architecture, and catalog status.
curl https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000?extended=true \
-H "Authorization: Bearer ere_myproject_your_api_key"
import requests
MODEL_ID = "00000000-1111-0000-1111-000000000000"
headers = {"Authorization": "Bearer ere_myproject_your_api_key"}
response = requests.get(
f"https://api.erebine.ai/proj_ABC123/v1/models/{MODEL_ID}",
params={"extended": "true"},
headers=headers
)
model = response.json()
print(f"{model['name']} - {model['status']}")
const MODEL_ID = "00000000-1111-0000-1111-000000000000";
const response = await fetch(
`https://api.erebine.ai/proj_ABC123/v1/models/${MODEL_ID}?extended=true`,
{
headers: {
"Authorization": "Bearer ere_myproject_your_api_key"
}
}
);
const model = await response.json();
console.log(`${model.name} - ${model.status}`);
Response (standard)
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"created": 1706123456,
"owned_by": "my-project",
"name": "my-custom-model"
}
Response (extended)
Optional fields are omitted from the JSON body when unset; do not rely on a key being present with an explicit null value.
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"created": 1706123456,
"owned_by": "my-project",
"name": "my-custom-model",
"format": "safetensors",
"size_bytes": 4294967296,
"size_formatted": "4.00 GB",
"status": "ready",
"version": "1.0.0",
"is_latest": true,
"description": "A fine-tuned LLM for code generation",
"architecture": "LlamaForCausalLM",
"quantization": "awq",
"context_length": 32768,
"parameter_count": 7000000000,
"workload_type": "chat",
"is_shared": false
}
Update Model Metadata
PATCH/proj_ABC123/v1/models/{modelId}
Updates metadata fields on an existing model. All fields in the request body are optional, only provided fields will be updated.
Request Body
| Parameter | Type | Description |
|---|---|---|
| descriptionoptional | string | Description of the model |
| licenseoptional | string | License identifier (e.g., "Apache-2.0") |
| architectureoptional | string | Model architecture (e.g., "LlamaForCausalLM") |
| hugging_face_idoptional | string or null | Upstream Hugging Face repository id (e.g. meta-llama/Llama-3.1-8B-Instruct, or a single-segment id such as gpt2). Omit the key to leave it unchanged; send null or an empty string to clear it. A value that cannot be a repository id returns HTTP 400 with "Invalid Hugging Face repository id: <value>". Detected from config.json's _name_or_path at import when the checkpoint carries one; never guessed from a file or directory name. |
| quantizationoptional | string | Quantization method. Must be one of the runtime-supported values: native, awq, bitsandbytes, bitblas, gguf, gptq, ipex, int4, int8, fp8, modelopt, quark, torchao, or compressed-tensors. Any other non-empty value returns HTTP 400 with "Invalid quantization method: <value>". Passing the empty string skips validation and is treated as "do not change". |
| context_lengthoptional | integer | Maximum context length in tokens |
| parameter_countoptional | integer | Total number of model parameters |
| hidden_sizeoptional | integer | Hidden dimension size extracted from model config |
| num_layersoptional | integer | Number of transformer layers |
| vocab_sizeoptional | integer | Vocabulary size extracted from model config |
| workload_typeoptional | string | Workload type: chat, code, reasoning, embedding, or multilingual. Defaults to chat. |
| reasoning_effortsoptional | array of strings or null | The reasoning_effort values the model accepts. Omit the key to leave it unchanged; send null for the model family's behaviour, [] for a model that does not reason, or a list drawn from none, minimal, low, medium, high, xhigh, max. Any other value returns HTTP 400. See Reasoning Effort. |
| speculative_decodingoptional | object or null | The model's speculative decoding setting. The keys you send are merged onto the stored setting and the result is checked as a whole; each key takes a value, or null to clear it. Send null for the whole object to reset the setting to off with nothing chosen. A setting that breaks a rule returns HTTP 400 naming the rule. See Speculative Decoding for the keys. |
curl -X PATCH https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000 \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"description": "Fine-tuned for code generation tasks",
"architecture": "LlamaForCausalLM",
"context_length": 32768
}'
import requests
MODEL_ID = "00000000-1111-0000-1111-000000000000"
headers = {
"Authorization": "Bearer ere_myproject_your_api_key",
"Content-Type": "application/json"
}
response = requests.patch(
f"https://api.erebine.ai/proj_ABC123/v1/models/{MODEL_ID}",
headers=headers,
json={
"description": "Fine-tuned for code generation tasks",
"architecture": "LlamaForCausalLM",
"context_length": 32768
}
)
model = response.json()
print(f"Updated: {model['name']}")
const MODEL_ID = "00000000-1111-0000-1111-000000000000";
const response = await fetch(
`https://api.erebine.ai/proj_ABC123/v1/models/${MODEL_ID}`,
{
method: "PATCH",
headers: {
"Authorization": "Bearer ere_myproject_your_api_key",
"Content-Type": "application/json"
},
body: JSON.stringify({
description: "Fine-tuned for code generation tasks",
architecture: "LlamaForCausalLM",
context_length: 32768
})
}
);
const model = await response.json();
console.log(`Updated: ${model.name}`);
Response
Optional fields that have no value are omitted entirely rather than emitted as null. For example, if parameter_count has not been set, the key will be absent from the body.
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"created": 1706123456,
"owned_by": "my-project",
"name": "my-custom-model",
"format": "safetensors",
"size_bytes": 4294967296,
"size_formatted": "4.00 GB",
"status": "ready",
"version": "1.0.0",
"is_latest": true,
"description": "Fine-tuned for code generation tasks",
"architecture": "LlamaForCausalLM",
"quantization": "native",
"context_length": 32768,
"workload_type": "chat",
"is_shared": false
}
Reasoning Effort
A model states which reasoning_effort values it accepts. The setting governs reasoning_effort on chat completions, reasoning.effort on the Responses API, and the reasoning object's effort, enabled and max_tokens. It is read on every path that takes them: endpoint-scoped routes, _direct, and the semantic routes.
| Setting | reasoning_efforts |
Behaviour |
|---|---|---|
| Default (family behaviour) | unset | Every model starts here. Each value maps onto the levels the model's chat template or family declares; a model whose family declares none receives the value unchanged. |
| Does not reason | [] |
Every reasoning control is refused with HTTP 400 unsupported_parameter, as OpenAI refuses reasoning_effort on a non-reasoning model. reasoning.exclude and include_reasoning are accepted: they only shape the response, and there is no reasoning to leave out. The model's provider listing stops advertising reasoning. |
| Only these values | a list | Only the listed values are accepted; any other returns HTTP 400 unsupported_value naming the list. An accepted value still maps onto the model's own levels as it does by default. reasoning.enabled: false counts as none, and reasoning.enabled: true alone counts as medium. |
On a semantic route the router picks the model, so a request cannot know which one answers. There a model that does not reason has the reasoning controls dropped, and a listed model takes the nearest listed value, instead of the request being refused. A request that names the model or endpoint gets the refusal.
Set it
In the dashboard, open the model, pick a setting in its Reasoning effort section, and save the section; the help text shows the levels the model's family declares. Through the API, send reasoning_efforts to Update Model Metadata. From the CLI:
erectl models 00000000-1111-0000-1111-000000000000 --reasoning-efforts unsupported
erectl models 00000000-1111-0000-1111-000000000000 --reasoning-efforts low,medium,high
erectl models 00000000-1111-0000-1111-000000000000 --reasoning-efforts default
Refusal on a model that does not reason
{
"error": {
"message": "Unsupported parameter: 'reasoning_effort' is not supported with this model.",
"type": "invalid_request_error",
"param": "reasoning_effort",
"code": "unsupported_parameter"
}
}
Speculative Decoding
Speculative decoding drafts several tokens ahead and has the model verify them in one forward pass, so a response streams faster without changing what the model samples. It is a setting on the model. Every endpoint that serves the model uses it, and every agent that runs one of those endpoints applies it when it starts the engine: the agent downloads the draft model, verifies it like any model, and passes it to the engine. Nothing is set on the agent and no files are placed by hand.
The setting
| Key | Type | Description |
|---|---|---|
| enabled | boolean | The model speculates by default on every endpoint. Requires a method. Defaults to false. |
| method | string or null | One of the methods below. |
| num_speculative_tokens | integer or null | Tokens drafted per step, 1 to 32. Unset uses the method's default. More tokens pay off only while most of them are accepted. |
| draft_model_id | string or null | The id of the model that drafts. Required by the methods that run a draft and refused by the others. |
| ngram_fallback | boolean | mtp only: speculate with ngram when the model's checkpoint ships no multi-token-prediction layers, instead of not speculating. Defaults to false. |
| temperature, top_p, top_k | number or null | Sampling defaults while the engine speculates: temperature 0 to 2, top_p above 0 and at most 1, top_k 1 or more. See Sampling defaults. |
Methods
| Method | Draft model | What drafts |
|---|---|---|
mtp |
Refused | The model's own multi-token-prediction layers. The checkpoint must ship them; ngram_fallback covers one that does not. |
ngram |
Refused | Prompt lookup: continues an n-gram already in the context. Pays off on output that repeats its input, such as code edits and extraction. |
eagle, eagle3 |
Required | An EAGLE or EAGLE-3 head trained for this model. |
dflash |
Required | A DFlash drafter trained for this model. |
draft_model |
Required | A smaller model that shares this model's tokenizer. |
medusa, mlp_speculator |
Required | Medusa or MLP speculator heads trained for this model. |
The draft model
A draft is a model in your catalog: one of your project's models, or a model in the shared catalog. Upload it like any other model; its name does not matter. It cannot be the model itself or a preseeded model, it cannot have a draft of its own, and a model that is another model's draft cannot take one. A draft still validating can be chosen ahead of time; it is used once it is ready.
Saving the setting, on the model page or through the API, restarts the engines on endpoints that serve the model with the new setting. When the method needs a draft, the engine waits up to five minutes for a draft still downloading, then starts without speculative decoding. A draft no longer named is no longer held for the model and can be evicted to free space. A model that is another model's draft cannot be deleted: Delete Model returns 409 naming the models in your project that use it and counting the rest. A shared model that a model in another project uses as its draft cannot leave the catalog: Unshare Model returns 409 counting those models. Deleting the endpoints that serve it does not take it out of the catalog either, and cancelling the upload that created it returns the delete's 409.
With draft_model, the draft must have the same vocabulary size as the model when both are known; the engine refuses the pair otherwise, so the setting is refused with HTTP 400 instead. An engine older than the release that introduced the chosen method starts without speculative decoding rather than failing to start.
When an endpoint speculates
- An endpoint's own Speculative decoding setting wins:
offnever speculates;onspeculates with this setting's method, or with the model's own multi-token-prediction layers when no method is chosen. - On
auto, the endpoint follows the model:enabledspeculates. A model not enabled speculates only on an endpoint whose workload profile islatency. - The operator of a self-hosted agent can override the setting for every model on that agent; see EIM Advanced Configuration.
Sampling defaults
temperature, top_p and top_k fill in only the parameters a request leaves unset, and only while the engine speculates from this setting. A request's own value always wins.
They exist for thinking budgets. With speculative decoding on, the current engine release corrupts a request that reaches its thinking budget when the request samples with top_p below 1 and no top_k (vllm-project/vllm#58231). A budget is therefore accepted only from a request that is greedy (temperature 0), sends top_p 1.0, or carries a top_k; any other request that sends reasoning.max_tokens gets HTTP 400 naming these ways out, and a budget derived from reasoning_effort is not sent. A top_k default makes every request that sends no top_k safe; top_p 1.0 or temperature 0 covers only requests that leave those unset.
HF config overrides
Speculative decoding and HF config overrides do not combine: the engine applies overrides to the model and not to its draft (vllm-project/vllm#37435), so a RoPE or YaRN override would have the model verify at one context geometry while the draft proposes at another. Enabling speculative decoding on a model with overrides is refused, and so is saving overrides on a model that speculates. When an endpoint turns speculation on for a model that has overrides, the engine launches without the overrides.
Set it
In the dashboard, open the model and use its Speculative decoding section: turn it on, pick the method, the draft and the token count, and save the section. Through the API, send speculative_decoding to Update Model Metadata; an extended read of a model you own returns the setting under the same key. From the CLI:
erectl models 00000000-1111-0000-1111-000000000000 --speculative on --speculative-method dflash \
--speculative-tokens 7 --draft-model 00000000-2222-0000-2222-000000000000
erectl models 00000000-1111-0000-1111-000000000000 --speculative on --speculative-method mtp --ngram-fallback on
erectl models 00000000-1111-0000-1111-000000000000 --speculative-top-k 20
erectl models 00000000-1111-0000-1111-000000000000 --speculative off
erectl models 00000000-1111-0000-1111-000000000000 --speculative-reset
curl -X PATCH https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000 \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"speculative_decoding": {
"enabled": true,
"method": "dflash",
"num_speculative_tokens": 7,
"draft_model_id": "00000000-2222-0000-2222-000000000000",
"top_k": 20
}
}'
In the extended read
"speculative_decoding": {
"enabled": true,
"method": "dflash",
"num_speculative_tokens": 7,
"draft_model_id": "00000000-2222-0000-2222-000000000000",
"ngram_fallback": false,
"top_k": 20
}
Refusal of a draft method with no draft
{
"error": {
"message": "Method dflash runs a draft model: choose one from the catalog.",
"type": "invalid_request_error",
"code": "invalid_request"
}
}
Delete Model
DELETE/proj_ABC123/v1/models/{modelId}
Permanently deletes a model from your project. This action cannot be undone.
Cascade Deletion: Deleting a model removes associated rows in the project owned by the model. The exact set of dependent records (for example endpoints, endpoint workers, routing rules, model-shared entries, version rows, and any pending tasks) is owned by the model service and database constraints; treat this list as informational. This action cannot be undone.
Requirements
- Valid project API key (the standard inference API-key authentication)
- The model must belong to the calling project
- No other model may use it as its speculative decoding draft. The delete returns
409naming those models; change or clear their draft first. See Speculative Decoding.
curl -X DELETE https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000 \
-H "Authorization: Bearer ere_myproject_your_api_key"
import requests
API_KEY = "ere_myproject_your_api_key"
BASE_URL = "https://api.erebine.ai/proj_ABC123/v1"
MODEL_ID = "00000000-1111-0000-1111-000000000000"
response = requests.delete(
f"{BASE_URL}/models/{MODEL_ID}",
headers={"Authorization": f"Bearer {API_KEY}"}
)
if response.status_code == 200:
print(f"Model {MODEL_ID} deleted successfully")
else:
print(f"Error: {response.json()}")
const MODEL_ID = "00000000-1111-0000-1111-000000000000";
const response = await fetch(
`https://api.erebine.ai/proj_ABC123/v1/models/${MODEL_ID}`,
{
method: "DELETE",
headers: {
"Authorization": "Bearer ere_myproject_your_api_key"
}
}
);
if (response.ok) {
console.log(`Model ${MODEL_ID} deleted successfully`);
} else {
const error = await response.json();
console.log(`Error: ${JSON.stringify(error)}`);
}
Response
{
"id": "00000000-1111-0000-1111-000000000000",
"object": "model",
"deleted": true
}
Revalidate Model
POST/proj_ABC123/v1/models/{modelId}/revalidate
Triggers re-validation of a model. The endpoint queues a revalidation job via the database service and returns status: "validating" immediately on success; the model row's status column is updated by the worker that picks up the job. Whether the request is accepted depends on the model's current status as enforced by the database service. Errors are surfaced as 404 Model not found.
curl -X POST https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000/revalidate \
-H "Authorization: Bearer ere_myproject_your_api_key"
import requests
MODEL_ID = "00000000-1111-0000-1111-000000000000"
headers = {"Authorization": "Bearer ere_myproject_your_api_key"}
response = requests.post(
f"https://api.erebine.ai/proj_ABC123/v1/models/{MODEL_ID}/revalidate",
headers=headers
)
result = response.json()
print(f"Status: {result['status']} - {result['message']}")
const MODEL_ID = "00000000-1111-0000-1111-000000000000";
const response = await fetch(
`https://api.erebine.ai/proj_ABC123/v1/models/${MODEL_ID}/revalidate`,
{
method: "POST",
headers: {
"Authorization": "Bearer ere_myproject_your_api_key"
}
}
);
const result = await response.json();
console.log(`Status: ${result.status} - ${result.message}`);
Response
{
"id": "00000000-1111-0000-1111-000000000000",
"status": "validating",
"message": "Model validation queued"
}
Get Catalog Info
GET/proj_ABC123/v1/models/{modelId}/catalog-info
Returns catalog information for a model that has been shared to the public catalog. Returns 404 if the model is not in the catalog.
Response Fields
| Field | Type | Description |
|---|---|---|
| id | string | Model id. |
| catalog_role | string | One of deployable or shared. See the Share Model section for the meaning of each. |
| deployment_count | integer | Number of projects that have deployed this model from the catalog. |
| is_on_shared_agents | boolean | Derived. true when catalog_role == "shared"; false otherwise. |
| is_public | boolean | Whether the catalog entry is discoverable by other projects. |
| featured | boolean | Editorial flag set by platform operators; not user-controlled. |
| shared_at | string | ISO-8601 timestamp of the first share. |
curl https://api.erebine.ai/proj_ABC123/v1/models/00000000-1111-0000-1111-000000000000/catalog-info \
-H "Authorization: Bearer ere_myproject_your_api_key"
import requests
MODEL_ID = "00000000-1111-0000-1111-000000000000"
headers = {"Authorization": "Bearer ere_myproject_your_api_key"}
response = requests.get(
f"https://api.erebine.ai/proj_ABC123/v1/models/{MODEL_ID}/catalog-info",
headers=headers
)
if response.status_code == 200:
info = response.json()
print(f"Role: {info['catalog_role']}, Deployments: {info['deployment_count']}")
else:
print("Model is not in the catalog")
const MODEL_ID = "00000000-1111-0000-1111-000000000000";
const response = await fetch(
`https://api.erebine.ai/proj_ABC123/v1/models/${MODEL_ID}/catalog-info`,
{
headers: {
"Authorization": "Bearer ere_myproject_your_api_key"
}
}
);
if (response.ok) {
const info = await response.json();
console.log(`Role: ${info.catalog_role}, Deployments: ${info.deployment_count}`);
} else {
console.log("Model is not in the catalog");
}
Response
{
"id": "00000000-1111-0000-1111-000000000000",
"catalog_role": "deployable",
"deployment_count": 5,
"is_on_shared_agents": false,
"is_public": true,
"featured": false,
"shared_at": "2026-01-15T10:30:00Z"
}
Model Versions
Four endpoints manage model versions. Detailed semantics and request bodies are covered in the Model Versioning guide; the endpoint surface is summarized here.
| Method | Path | Purpose |
|---|---|---|
| GET | /proj_ABC123/v1/models/{modelId}/versions |
List all versions of a model. |
| POST | /proj_ABC123/v1/models/{modelId}/versions |
Create a new version row for a model. |
| POST | /proj_ABC123/v1/models/{modelId}/versions/{version}/promote |
Promote the named version to be the latest. |
| POST | /proj_ABC123/v1/models/{modelId}/versions/{version}/rollback |
Roll back to the named version. |
Preseeded Models
A model badged Preseeded in the dashboard was imported from a private agent's local disk instead of uploaded. Its weights live only on the agents that imported them; the platform holds no copy and moves no bytes. Extended responses report this on the origin field, which is "preseeded" for these models and "uploaded" for everything else. The Private Agents guide covers the import itself.
Metadata works the same as it does for an uploaded model. Update Model Metadata accepts the same fields, every section of the model's dashboard page is open, Hugging Face overrides included, and the values you save drive what the agent launches the engine with: architecture, context length, quantization, and parsers all apply. The platform never harvests metadata out of the files, so what you enter is what the model reports.
The verbs that need platform-side bytes return 409 with the reason in the message: revalidate, share, unshare, and creating, promoting, or rolling back a version. A preseeded model has no catalog entry, no versions, and counts nothing against your project's storage quota. Delete returns 409 while an endpoint or an anchor endpoint binds the model; once unbound, the delete removes the row and every holding agent purges its local copy. Importing the same directory again afterwards creates a new model with a new id.
Error Responses
Every endpoint in this surface returns OpenAI-shaped error envelopes. The error classes below cover the failure modes you are likely to hit when integrating.
401 Unauthorized
Returned when no Authorization header is present or the bearer token is malformed.
{
"error": {
"message": "Missing or invalid Authorization header",
"type": "invalid_request_error",
"code": "missing_authorization"
}
}
403 Forbidden
Returned when the API key is valid but belongs to a different project than the one in the URL path, or when the key has been revoked.
{
"error": {
"message": "API key does not authorize access to this project",
"type": "invalid_request_error",
"code": "project_mismatch"
}
}
404 Model not found
GET /models/{modelId} answers an id it cannot find -- not a model id, no such model, or a model the calling project cannot see -- with the body the OpenAI API sends for a model the caller cannot use, which names the id you sent. DELETE /models/{modelId}, the route an OpenAI client's models.delete() calls, answers a model the calling project does not own with the same body.
{
"error": {
"message": "The model `gpt-4o-mini` does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": null,
"code": "model_not_found"
}
}
The other /models/{modelId} endpoints are model management the OpenAI API does not have. PATCH /models/{modelId}, revalidate, GET versions and catalog-info answer a model id that does not exist in the calling project with the body below; catalog-info answers the same way, with the message Model is not in the catalog, when the model exists but is not in the public catalog, and unshare with the message Model not found or not shared.
{
"error": {
"message": "Model not found",
"type": "not_found_error",
"code": "not_found"
}
}
400 Invalid quantization method
Returned by PATCH /models/{modelId} when the quantization body field is a non-empty value that is not in the runtime-supported set. The offending value is echoed in the message for diagnosis.
{
"error": {
"message": "Invalid quantization method: fp4",
"type": "invalid_request_error",
"code": "invalid_quantization"
}
}
400 Invalid speculative decoding setting
Returned by PATCH /models/{modelId} when the merged speculative_decoding setting breaks a rule: enabled without a method, a method that runs a draft with no draft_model_id, a draft on mtp or ngram, ngram_fallback on a method other than mtp, a value out of range, a draft that is not usable (see The draft model), or speculation on a model with HF config overrides. The message names the rule and nothing is stored.
409 Model is a speculative decoding draft
Returned by DELETE /models/{modelId} when another model uses this one as its speculative decoding draft. The message names those models.
400 Invalid pagination bound
Returned by the management listing GET /proj_ABC123/v1/models when limit or offset carries a negative value; the bound is refused rather than clamped. The param field names the offending parameter; limit is checked first, so a request that makes both negative is told about limit. The inference-facing listings do not share this refusal -- a negative bound there is ignored and the full visible inventory is returned.
{
"error": {
"message": "Invalid value for 'limit': must not be negative",
"type": "invalid_request_error",
"code": "invalid_request",
"param": "limit"
}
}
Keyboard Shortcuts
This page registers a small set of chord shortcuts for jumping between endpoints. Shortcuts are suppressed while typing in inputs, textareas, or contenteditable surfaces. Press ? to toggle the help dialog.
| Chord | Action |
|---|---|
| g o | Jump to Overview |
| g b | Jump to Base URL |
| g l | Jump to List Models |
| g d | Jump to Get Model Details |
| g u | Jump to Update Model |
| g x | Jump to Delete Model |
| g r | Jump to Revalidate Model |
| g s | Jump to Share Model |
| g n | Jump to Unshare Model |
| g c | Jump to Catalog Info |
| g v | Jump to Model Versions |
| g p | Jump to Preseeded Models |
| g e | Jump to Error Responses |
| g t | Jump to Scope Table |
| ? | Toggle this shortcuts dialog |
| Esc | Close the shortcuts dialog |
Endpoint Scope Table
All Model Management endpoints below are mounted on the project inference route group and gated by the same project API-key middleware as /v1/chat/completions. There is no separate models or Model Management scope enforced at the route layer at this time; any valid project API key is accepted. A scope gate of this kind is tracked as a cross-cutting follow-up and may change in a future release.
| Method | Path | Route group | Auth |
|---|---|---|---|
| GET | /v1/models | inference | project API key |
| GET | /v1/models/{modelId} | inference | project API key |
| PATCH | /v1/models/{modelId} | inference | project API key |
| DELETE | /v1/models/{modelId} | inference | project API key |
| POST | /v1/models/{modelId}/revalidate | inference | project API key |
| POST | /v1/models/{modelId}/share | inference | project API key |
| POST | /v1/models/{modelId}/unshare | inference | project API key |
| GET | /v1/models/{modelId}/versions | inference | project API key |
| POST | /v1/models/{modelId}/versions | inference | project API key |
| POST | /v1/models/{modelId}/versions/{version}/promote | inference | project API key |
| POST | /v1/models/{modelId}/versions/{version}/rollback | inference | project API key |
| GET | /v1/models/{modelId}/catalog-info | inference | project API key |