Responses API
Stateful generation without rebuilding the conversation each turn. Chain responses by id, queue long jobs in the background, stream incremental reasoning, and let the server keep the message history. Same wire shape as the OpenAI Responses API.
Overview
When to Use Responses API
- Multi-turn conversations, Chain responses together with
previous_response_idinstead of resending the full message history. - Persistent storage, Responses are stored and retrievable by ID for later reference.
- Background processing, Queue long-running requests and poll for completion.
- Client SDK support, OpenAI Python/Node.js SDKs natively support the Responses API.
When to Use Chat Completions
- You need full control over the message history.
- You are using a client or tool that only supports the Chat Completions API.
- You do not need server-side storage of responses.
Translation layer. Internally, every Responses API request is converted to a Chat Completion, routed through the same inference pipeline, and converted back to the Response format. Model behavior is identical.
Quick Start
Create a response with a simple text input:
from openai import OpenAI
client = OpenAI(
base_url="https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
api_key="ere_myproject_your_api_key"
)
response = client.responses.create(
model="llama-3.1-8b",
input="What is the capital of France?"
)
print(response.output[0].content[0].text)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
apiKey: "ere_myproject_your_api_key"
});
const response = await client.responses.create({
model: "llama-3.1-8b",
input: "What is the capital of France?"
});
console.log(response.output[0].content[0].text);
curl -X POST https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What is the capital of France?"
}'
Response
{
"id": "resp_abc123def456ghi789jkl012",
"object": "response",
"model": "llama-3.1-8b",
"status": "completed",
"output": [
{
"type": "message",
"id": "msg_001",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "The capital of France is Paris."
}
],
"status": "completed"
}
],
"usage": {
"input_tokens": 12,
"output_tokens": 8,
"total_tokens": 20
},
"created_at": 1706123456,
"service_tier": "default",
"store": true,
"background": false,
"metadata": null
}
Authentication
All Responses API endpoints require a valid API key with the
inference scope. Pass it in the Authorization header:
Authorization: Bearer ere_myproject_your_api_key
See Authentication & Security for details on creating and managing API keys.
Endpoints
| Method | Path | Description |
|---|---|---|
| POST | /v1/responses | Create a response |
| GET | /v1/responses | List responses |
| GET | /v1/responses/{response_id} | Get a response by ID |
| DELETE | /v1/responses/{response_id} | Delete a response |
| POST | /v1/responses/{response_id}/cancel | Cancel an in-progress response |
| GET | /v1/responses/{response_id}/input_items | List input items for a response |
All paths are relative to your endpoint base URL:
https://api.erebine.ai/proj_ABC123/ENDPOINT_SLUG
Create Response
POST /v1/responses
Request Body
| Parameter | Type | Description |
|---|---|---|
| modelrequired | string | Model identifier. Used by the router for validation gates (e.g. model_not_found, invalid_model_id); the endpoint configuration ultimately determines which backend model serves the request. |
| inputrequired | string | array | Input content. Can be a plain text string, an array of messages ({role, content}), or an array of input items ({type, role, content, call_id, output}). |
| instructionsoptional | string | System/developer instructions. Prepended as a system message if not already present from the response chain. |
| streamoptional | boolean | If true, the response is streamed as Server-Sent Events. Default: false |
| storeoptional | boolean | Whether to persist the response for later retrieval. Default: true. In a zero-data-retention workspace an explicit true is refused with 400 retention_not_permitted and an absent value is treated as false; see Zero data retention. |
| backgroundoptional | boolean | If true, the request returns immediately with a queued status. Poll the response ID for completion. With store: false the response can be polled while it runs and for about 10 minutes after it finishes, then returns 404; see Storage & Retention. Default: false. Refused with 400 retention_not_permitted in a zero-data-retention workspace; see Zero data retention. |
| previous_response_idoptional | string | ID of a previous response to chain onto. The previous response's context is automatically prepended. Mutually exclusive with conversation. See Conversation Chaining. |
| conversationoptional | object | Link this response to a server-side conversation. Pass {"id": "conv_xxx"} to prepend the conversation's existing items as context and append the response output as new items. Mutually exclusive with previous_response_id; sending both returns 400 invalid_request_error (code invalid_request, param conversation). See Conversations. |
| max_output_tokensoptional | integer | Maximum number of output tokens to generate. Sent to the model verbatim. If it does not fit alongside the input inside the model's context window, the request returns 400 context_length_exceeded unless clamping is opted in. See Passthrough and Augmentation. |
| temperatureoptional | number | Sampling temperature (0.0-2.0). Higher values produce more random output. |
| top_poptional | number | Nucleus sampling parameter (0.0-1.0). |
| toolsoptional | array | Tool definitions the model may call, and the web_search tool. Some hosted tools return 400; see Hosted Tools. See Tool Calling and Web Search. |
| tool_choiceoptional | string | object | Controls tool selection: "auto", "none", "required", {"type":"function","name":"fn_name"}, or {"type":"allowed_tools","mode":"auto","tools":[...]}. The nested Chat Completions spelling {"type":"function","function":{"name":"fn_name"}} is also accepted. "none", and an allowed_tools list without a {"type":"web_search"} entry, run no web search. The choice is echoed as sent. |
| parallel_tool_callsoptional | boolean | Allow multiple tool calls in a single response. Default: true |
| textoptional | object | Text format configuration. Supports {"format":{"type":"text"}}, {"format":{"type":"json_object"}}, or {"format":{"type":"json_schema","json_schema":{...}}} |
| reasoningoptional | object | Reasoning configuration for reasoning models. {"effort":"none|minimal|low|medium|high|xhigh|max"}, mapped onto the model's levels as reasoning_effort is on Chat Completions. A value the model cannot honour returns 400 with code unsupported_value and param: "reasoning.effort". A model its operator marked as not reasoning refuses reasoning.effort with code unsupported_parameter (see Reasoning Effort). |
| metadataoptional | object | Up to 16 key-value pairs. Keys max 64 characters, values max 512 characters. |
| useroptional | string | End-user identifier for abuse monitoring and usage tracking. |
| truncationoptional | string | Truncation strategy when the prior context plus the new input exceeds the model context window. Applies the same way to a previous_response_id chain and to a server-side conversation. "disabled" (the default) returns 400 context_length_exceeded and trims nothing; "auto" drops middle items to fit the deployed model's context window, keeping the earliest anchor and the most recent tail. The "disabled" rejection fires only when the model's context length is resolvable -- for a model whose window cannot be determined, no token check applies. A server-side conversation is additionally bounded at 2000 items per request: under "disabled" a longer conversation returns 400 context_length_exceeded rather than answering from part of it, and under "auto" the most recent 2000 items are used. See Conversation Chaining. |
| service_tieroptional | string | Processing lane: "auto", "default", "flex", "scale", or "priority". "flex" runs at lower priority and is billed at 0.5x; "priority" runs at higher priority and is billed at 1.25x; every other value, and an omitted field, resolves to "default". The resolved lane is echoed as service_tier on the response object, its streaming events, and a retrieved response. See Service Tiers. |
| includeoptional | array | Additional output data to include. Values are validated against the OpenAI include vocabulary; an unrecognised value returns 400. web_search_call.action.sources adds the pages each search returned to its web_search_call item (see Web Search). Every other well-formed value is accepted and has no effect on the response body. |
| research_depthoptional, vendor extension | string | Erebine addition, not part of the OpenAI schema. Sets how much research the turn does: off, auto (default), or deep. See Research Depth. |
Input Formats
The input field accepts three formats:
"input": "What is the capital of France?"
"input": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there!"},
{"role": "user", "content": "What is 2+2?"}
]
"input": [
{"type": "message", "role": "user", "content": "Call get_weather for Paris"},
{"type": "function_call_output", "call_id": "call_abc", "output": "{\"temp\":18}"}
]
List Responses
GET /v1/responses
Returns a paginated list of the stored responses of the workspace the request resolves to (see Workspace scope), ordered by creation time (newest first).
Query Parameters
| Parameter | Type | Description |
|---|---|---|
| afteroptional | string | Cursor for forward pagination. Pass the id of the last response from the previous page. |
| limitoptional | integer | Number of responses to return. Default: 20. Maximum: 100. |
Response
{
"object": "list",
"data": [
{
"id": "resp_abc123",
"object": "response",
"model": "llama-3.1-8b",
"status": "completed",
"created_at": 1706123456,
"completed_at": 1706123458,
"input_tokens": 12,
"output_tokens": 8,
"store": true,
"background": false,
"metadata": null
}
],
"first_id": "resp_abc123",
"last_id": "resp_abc123",
"has_more": false
}
curl "https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses?limit=20" \
-H "Authorization: Bearer ere_myproject_your_api_key"
Get Response
GET /v1/responses/{response_id}
Retrieves the full response object for a stored response, including output content. A background response sent with "store": false can be retrieved while it is queued or in progress, and with its output for about 10 minutes after it finishes; see Storage & Retention.
curl "https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123" \
-H "Authorization: Bearer ere_myproject_your_api_key"
Response
{
"id": "resp_abc123",
"object": "response",
"model": "llama-3.1-8b",
"status": "completed",
"created_at": 1706123456,
"completed_at": 1706123458,
"input_tokens": 12,
"output_tokens": 24,
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{"type": "output_text", "text": "Paris is the capital of France."}
]
}
],
"store": true,
"background": false,
"metadata": null
}
See Response Object for the full field list. Returns 404 if the response does not exist, belongs to a different project, or belongs to another workspace than the one the request resolves to (see Workspace scope).
Delete Response
DELETE /v1/responses/{response_id}
Deletes a response. This action cannot be undone. Afterwards the response cannot be retrieved, listed, chained from with previous_response_id or read through input_items, and a second delete returns 404, exactly as an unknown ID does.
A background response can be deleted at any point. One still queued is cancelled first and never runs, so it is not charged. One already in_progress cannot be stopped: it finishes and is charged for the work it did, but its result is discarded. A background response sent with "store": false stops being retrievable at once instead of about 10 minutes after it finishes.
Returns 404 for any response Get Response would not return: one that does not exist, belongs to a different project or to another workspace than the one the request resolves to, was sent with "store": false (other than a background response inside its polling window), or was already deleted.
curl -X DELETE \
"https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123" \
-H "Authorization: Bearer ere_myproject_your_api_key"
Response
{
"id": "resp_abc123",
"object": "response.deleted",
"deleted": true
}
Cancel Response
POST /v1/responses/{response_id}/cancel
Cancels an in-progress response. Only responses with status in_progress
or queued can be cancelled. Completed, failed, or already-cancelled
responses return a 400 error.
curl -X POST \
"https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123/cancel" \
-H "Authorization: Bearer ere_myproject_your_api_key"
Returns the updated response object with status: "cancelled".
List Input Items
GET /v1/responses/{response_id}/input_items
Returns the input items that were submitted with the response request. The listing runs through the same loader the generation path uses, so it reports the pool the model's context was assembled from. When the response is linked to a conversation, the conversation's items are listed; the reconstructed items from a previous_response_id chain are used only when no conversation is attached.
The item pool the listing pages over is capped at 2000 conversation items -- the same cap the generation path applies. Under the default truncation: "disabled" a request against a longer conversation is rejected with 400 context_length_exceeded, so such a stored response can only have been generated with truncation: "auto" -- from the most recent 2000 items. Each response records the truncation strategy it resolved to and the context window it was budgeted against, and the listing replays that recorded selection, so it reports the same window end and the same budget generation used; items appended after generation, including the response's own output, are part of the pool the selection is replayed over.
Responses stored before the strategy was recorded (before v0.7.5) carry none, and are listed from the conversation's oldest items -- the end they were generated from.
Query Parameters
| Parameter | Type | Description |
|---|---|---|
| afteroptional | string | Cursor for forward pagination. |
| limitoptional | integer | Number of items to return. Default: 20. Maximum: 100. |
curl "https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123/input_items" \
-H "Authorization: Bearer ere_myproject_your_api_key"
Response Object
A completed response contains the model's output, usage information, and metadata.
{
"id": "resp_abc123def456ghi789jkl012",
"object": "response",
"model": "llama-3.1-8b",
"status": "completed",
"previous_response_id": null,
"output": [
{
"type": "message",
"id": "msg_001",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "The capital of France is Paris.",
"annotations": []
}
],
"status": "completed"
}
],
"usage": {
"input_tokens": 12,
"output_tokens": 8,
"total_tokens": 20,
"input_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0
},
"output_tokens_details": {
"reasoning_tokens": 0
}
},
"created_at": 1706123456,
"completed_at": 1706123457,
"service_tier": "default",
"store": true,
"background": false,
"metadata": null,
"tools": [],
"tool_choice": "auto",
"parallel_tool_calls": true
}
Every output_text part carries annotations. It is empty unless the answer cites a page its web search returned; see Web Search.
Every response object carries tools, tool_choice and parallel_tool_calls, echoing the request or, when the request left one out, the default: [], "auto" and true. usage always carries input_tokens_details (cached_tokens, cache_write_tokens) and output_tokens_details (reasoning_tokens); a count the model server did not report is 0, and cache_write_tokens is always 0.
background is true for a response created with "background": true and false for any other. Every response object carries it: the create response, each streamed snapshot (response.created, response.in_progress, response.completed, response.failed), Get Response, Cancel Response and List Responses.
Status Values
| Status | Description |
|---|---|
queued | Request received, waiting for processing. |
in_progress | Inference is actively running. |
completed | Response generated successfully. |
failed | An error occurred during generation. |
cancelled | Cancelled by the user before completion. |
incomplete | Generation stopped early (max tokens, content filter, etc.). |
Output Item Types
| Type | Description |
|---|---|
message |
Assistant text message. Contains role, content[] (array of content parts), and status. |
function_call |
Tool/function call. Contains call_id, name, arguments (JSON string), and status. |
web_search_call |
A web search or page fetch the research tools ran, ahead of the message. Contains id, status (in_progress, completed, failed), and action: {"type": "search", "query": ...} or {"type": "open_page", "url": ...}. A search action carries sources when the request's include asks for them. |
reasoning |
Reasoning summary from models that produce think-tag content. Contains id and summary text. |
Echo Fields
The response object includes echo fields that mirror request parameters, making it easy to see the exact configuration used.
| Field | Type | Description |
|---|---|---|
| temperature | number | null | The temperature value used for this response. |
| top_p | number | null | The top_p value used for this response. |
| max_output_tokens | integer | null | The maximum output tokens configured for this response. |
| tools | array | null | The tools that were available for this response. |
| tool_choice | string | object | null | The tool choice setting used for this response. |
| text | object | null | The text format configuration used for this response. |
| reasoning | object | null | The reasoning configuration used for this response. |
| truncation | string | null | The truncation strategy used for this response. |
| instructions | string | null | The system instructions used for this response. |
| parallel_tool_calls | boolean | null | Whether parallel tool calls were enabled. |
Streaming
Set "stream": true to receive the response as Server-Sent Events.
The Responses API uses named event types for structured streaming.
Lifecycle Events
| Event | Description |
|---|---|
response.created | Response record created (status: queued). Contains the initial response object. |
response.in_progress | Inference started (status: in_progress). |
response.completed | Response completed normally. Contains the final response object with full usage data. |
response.failed | An error occurred during generation. Contains the response object with error details, the output items produced before the failure (a message cut short has status incomplete), and the usage consumed. The error's code is one of the Response object's codes; see Error Handling. Mid-stream errors are surfaced via this event (the Responses stream does not emit a separate event: error named line). |
Output Item Events
| Event | Description |
|---|---|
response.output_item.added | New output item (message, function_call, or reasoning) started. Contains output_index and initial item object. |
response.output_item.done | Output item finished. Contains the completed item object. |
response.content_part.added | New content part added to an output item. Contains output_index, content_index, and initial part object. |
response.content_part.done | Content part finished. Contains the completed part object. |
Text Delta Events
| Event | Description |
|---|---|
response.output_text.delta | Incremental text content. Contains item_id, output_index, content_index, and delta string. |
response.output_text.annotation.added | A url_citation annotation on the text, sent before response.output_text.done. Contains item_id, output_index, content_index, annotation_index, and annotation. |
response.output_text.done | Text content part complete. Contains item_id, output_index, content_index, and full accumulated text. |
Tool Call Events
| Event | Description |
|---|---|
response.function_call_arguments.done | Function call arguments complete. Contains output_index, item_id, and full arguments JSON string. (Arguments are emitted as a single terminal event; no incremental .delta stream is produced.) |
Reasoning Summary Events
Emitted by models that produce think-tag reasoning content (e.g. Qwen3, DeepSeek-R1).
| Event | Description |
|---|---|
response.reasoning_summary_part.added | Reasoning summary part started within a reasoning output item. |
response.reasoning_summary_text.delta | Incremental reasoning summary text. Contains item_id, output_index, summary_index, and delta string. |
response.reasoning_summary_text.done | Reasoning summary text complete. Contains the full accumulated text. |
response.reasoning_summary_part.done | Reasoning summary part finished. |
Streaming Example
from openai import OpenAI
client = OpenAI(
base_url="https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
api_key="ere_myproject_your_api_key"
)
stream = client.responses.create(
model="llama-3.1-8b",
input="What is the capital of France?",
stream=True
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
apiKey: "ere_myproject_your_api_key"
});
const stream = await client.responses.create({
model: "llama-3.1-8b",
input: "What is the capital of France?",
stream: true
});
for await (const event of stream) {
if (event.type === "response.output_text.delta") {
process.stdout.write(event.delta);
}
}
curl --no-buffer -X POST \
https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What is the capital of France?",
"stream": true
}'
event: response.created
data: {"id":"resp_abc123","object":"response","status":"queued","model":"llama-3.1-8b"}
event: response.in_progress
data: {"id":"resp_abc123","status":"in_progress"}
event: response.output_item.added
data: {"output_index":0,"item":{"type":"message","id":"msg_001","role":"assistant","content":[],"status":"in_progress"}}
event: response.content_part.added
data: {"output_index":0,"content_index":0,"part":{"type":"output_text","text":""}}
event: response.output_text.delta
data: {"item_id":"msg_001","output_index":0,"content_index":0,"delta":"The capital of France is Paris."}
event: response.content_part.done
data: {"output_index":0,"content_index":0,"part":{"type":"output_text","text":"The capital of France is Paris."}}
event: response.output_item.done
data: {"output_index":0,"item":{"type":"message","id":"msg_001","role":"assistant","content":[{"type":"output_text","text":"The capital of France is Paris."}],"status":"completed"}}
event: response.completed
data: {"id":"resp_abc123","object":"response","status":"completed","usage":{"input_tokens":12,"output_tokens":8,"total_tokens":20}}
event: done
data: [DONE]
Vendor Events
In addition to standard response.* events, the platform emits
vendor-prefixed events (x_*) for Erebine-specific features such
as research mode, the research pipeline, and user interaction. Vendor events ride on the
Responses stream in two wire shapes:
- The dominant shape is a bare
data: {"type":"x_*",...}line with no precedingevent:name. Standard OpenAI SDK clients ignore unrecognized data lines, so this shape is wire-compatible. - The
x_artifact.*family on the Responses path uses a named SSE event:event: x_artifact.created\ndata: {json}\n\nwith camelCase payload keys. Strict OpenAI SDK clients will surface these as unknown event types.
Some vendor families listed below are emitted only by the Erebine dashboard
surface and are not produced on the public Responses stream;
those rows are marked dashboard-only. Public SDK consumers should
not depend on dashboard-only events arriving on /v1/responses.
Compatibility. Standard OpenAI SDK clients ignore unrecognized data lines. Vendor events are only relevant when building a custom stream consumer that wants to render research progress, deep think status, or artifact notifications.
Research Events (x_research.*)
Emitted during agentic research loops (research mode). Each event is a
data: line containing a JSON object with a type
field, name (tool name), arguments (JSON string),
and optional metadata.
| Event type | Description |
|---|---|
x_research.searching | Web search tool invoked. arguments contains the search query JSON. |
x_research.reading | URL fetch tool invoked. arguments contains {"url":"..."}. |
x_research.code_searching | Code search tool invoked (GitLab or local index). |
x_research.calculating | Calculator tool invoked. |
x_research.result | Tool returned a result. metadata contains a brief summary of the result. |
x_research.complete (dashboard-only) | Research loop finished. Emitted only by the Erebine dashboard surface; the public Responses stream does not produce this event. Dashboard payload contains elapsed_ms, input_tokens, output_tokens, iterations, and sources counts. |
data: {"type":"x_research.searching","name":"x_web_search","arguments":"{\"query\":\"Paris weather\"}"}
data: {"type":"x_research.reading","name":"x_fetch_url","arguments":"{\"url\":\"https://example.com/\"}"}
data: {"type":"x_research.result","name":"x_web_search","arguments":"{\"query\":\"Paris weather\"}","metadata":{"summary":"Paris is currently 18C and sunny."}}
data: {"type":"x_research.complete","elapsed_ms":4200,"input_tokens":3200,"output_tokens":480,"iterations":3,"sources":5}
Research Pipeline Events (x_research.*)
Emitted during a research pipeline run (a deep turn, or an
escalated auto turn): planning, sub-task fan-out, deepening,
and synthesis. Clients can use these to render a sub-task progress panel.
Every pipeline event carries a common envelope: run_id,
depth (off / auto / deep),
stage, round (1-based), and, where applicable,
subtask_index and attempt.
| Event type | Description |
|---|---|
x_research.escalation_proposed | The system proposes escalating an auto turn into a research run. |
x_research.plan_created | The planner produced the initial research plan. |
x_research.plan_updated | The plan changed (critique pass or per-round follow-ups). |
x_research.subtask_started | A research sub-task began executing. |
x_research.subtask_delta | Output-token delta from a sub-task. |
x_research.subtask_completed | A research sub-task finished. |
x_research.subtask_retried | A weak sub-task result triggered a retry attempt. |
x_research.cross_reference | A cross-reference between sub-task findings. |
x_research.gaps | Gap-analysis result for a deepening round. |
x_research.round_completed | A deepening round finished. |
x_research.converged | Gap analysis found nothing actionable; deepening stops early. |
x_research.synthesis_started | Final synthesis began streaming. |
x_research.synthesis_delta | Synthesis output-token delta. |
x_research.synthesis_completed | Final synthesis finished. |
x_research.claim | A structured claim extracted during synthesis. |
x_research.intelligence_persisted | The intelligence bundle (memories, artifacts, project-intelligence) was persisted. |
x_research.run_detached | The live run detached from the turn and continues in the background. |
x_research.run_completed | Terminal: the run completed successfully. |
x_research.run_failed | Terminal: the run failed. |
x_research.error | A non-terminal pipeline error surfaced to the consumer. |
All pipeline lifecycle events on the Responses API use the
x_research.* names and the common envelope shown above; match on the
x_research. prefix and dispatch on the event type.
Ask User Events (x_ask_user.*)
Emitted when the model needs clarification before it can continue. The stream pauses and the application should prompt the user, then resume with the answer.
| Event type | Description |
|---|---|
x_ask_user.question | The model requires user input before continuing. Contains ask_user_id (correlation ID) and question text. Submit the answer via the chat answer endpoint. |
x_ask_user.pending_state | Captures the assistant content and tool calls accumulated before the pause, enabling conversation resumption after the user answers. |
Artifact Events (x_artifact.*)
Emitted when code artifacts (code blocks, documents) are created or updated
during a generation. On the Responses path these are written as named
SSE events (event: x_artifact.created\ndata: {json}\n\n)
with camelCase payload keys, this differs from the
chat-completions surface, which emits the same family as bare
data: lines with snake_case keys.
| Event type | Description |
|---|---|
x_artifact.created | New artifact created. Responses-path payload contains artifactId, identifier, title, language, contentType, and contentBase64 (base64-encoded artifact bytes). |
x_artifact.updated | Existing artifact updated. Same payload shape as x_artifact.created with the new version of the content. |
Context Fork Event (x_context_fork) (dashboard-only)
Emitted by the Erebine dashboard surface when the user's message triggered
creation of a new conversation branch. The public Responses stream does not
emit this event. Unlike other vendor events the name has no
.<suffix> segment.
Chat Metadata Event (x_chat.metadata) (dashboard-only)
Emitted by the Erebine dashboard surface as the final data event before
[DONE]. The public Responses stream (/v1/responses)
does not emit this event; public SDK consumers should not wait for it.
Payload uses camelCase keys (dashboard convention) rather than the snake_case
used by the rest of the vendor surface.
| Field | Description |
|---|---|
type | "x_chat.metadata" |
messageId | Server-assigned external ID for the persisted assistant message. |
userMessageId | External ID for the persisted user message. |
sequence | Monotonically increasing sequence number for the assistant message within the conversation. |
context | Context budget breakdown: systemTokens, summaryTokens, retrievedTokens, recentTokens, fileTokens, currentMessageTokens, totalTokens, inputBudget, retrievedCount, recentCount, usedSemanticRetrieval, semanticRetrievalActive, chunkSelectionMethod. |
usage | Combined token usage including model inference plus any research-pipeline overhead: input_tokens, output_tokens, total_tokens. |
Analyst Events (x_analyst.*)
Emitted when the analyst mode builds or refreshes the workspace context brief before generating a response.
| Event type | Description |
|---|---|
x_analyst.context_gathering | Workspace context gathering has begun. |
x_analyst.context_completed | Context gathering finished. Contains counts of gathered items. |
x_analyst.context_brief_created | The LLM-generated context brief is ready. The brief summarizes the workspace for the response. |
x_analyst.context_refreshed | A previously cached context brief was refreshed due to workspace changes. |
Conversation Chaining
Use previous_response_id to build multi-turn conversations without
resending the full message history. The server automatically retrieves the
previous response's context and prepends it to your new input.
curl -X POST https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What is the capital of France?"
}'
# Returns: {"id": "resp_abc123", ...}
curl -X POST https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What about Germany?",
"previous_response_id": "resp_abc123"
}'
Chain Requirements
- The previous response must exist in the same project and belong to the workspace the request resolves to (see Workspace scope); otherwise the request returns
404, as for a response that does not exist. - The previous response must have succeeded:
completedorincomplete. Chaining from afailedorcancelledresponse returns400 invalid_request_errorwith codeinvalid_state-- a request error to fix, not a server fault to retry. Aqueuedorin_progressresponse returns the same error. - The previous response must have been stored (
storedefaults to true). Chaining from astore: falseresponse returns400 invalid_state. - Maximum chain depth is 50 responses.
- Circular references are detected and rejected.
previous_response_idandconversationare mutually exclusive. A request carrying both returns400 invalid_request_error(codeinvalid_request, paramconversation); neither source is applied.- The replayed chain is subject to
truncation. Under the default"disabled"an over-budget chain returns400 context_length_exceeded; settruncation: "auto"to drop middle items instead.
Tool Calling
Define function tools in the tools parameter. The model may generate
function_call output items that your application executes.
A Responses API function tool is flat: name,
description, parameters and strict
sit at the top level of the tool, with no nested function
object. The nested Chat Completions spelling is also accepted here.
{
"model": "llama-3.1-8b",
"input": "What is the weather in Paris?",
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
]
}
{
"id": "resp_tool123",
"status": "completed",
"output": [
{
"type": "function_call",
"id": "fc_001",
"call_id": "call_abc123",
"name": "get_weather",
"arguments": "{\"location\":\"Paris\"}",
"status": "completed"
}
]
}
{
"model": "llama-3.1-8b",
"previous_response_id": "resp_tool123",
"input": [
{
"type": "function_call_output",
"call_id": "call_abc123",
"output": "{\"temperature\": 18, \"condition\": \"sunny\"}"
}
]
}
Hosted Tools
Of the OpenAI hosted tools, only web search runs here. A hosted tool that
would change the answer if it were ignored returns 400
invalid_request_error with code unsupported_value
and param naming its type, for example
tools[1].type. Remove it, or declare a function tool your
application executes instead.
| Tool type | Behavior |
|---|---|
web_search, web_search_preview and their dated types |
Runs. See Web Search. |
file_search, code_interpreter, image_generation, computer, computer_use_preview, mcp, tool_search, programmatic_tool_calling |
Returns 400 unsupported_value. |
namespace |
Its functions are offered to the model beside your other functions, under their own names, and a call to one comes back as a function_call item carrying the namespace. Its custom tools behave as below. The tool is echoed as sent. Function names must be unique across namespaces and top-level functions; a repeated name returns 400 unsupported_value with param such as tools[1].tools[0].name. defer_loading is ignored: every function is offered up front. |
custom, local_shell, shell, apply_patch |
Accepted and echoed in tools exactly as sent, but not offered to the model, which cannot call them. |
Web Search
Add {"type": "web_search"} to tools and the
router searches the web and reads pages before it answers. Each search
and page fetch comes back as a web_search_call output item
ahead of the message; your application executes nothing. The legacy
web_search_preview and the dated
web_search_2025_08_26 and
web_search_preview_2025_03_11 run the same search.
{
"model": "llama-3.1-8b",
"input": "What was a positive news story from today?",
"tools": [{"type": "web_search"}]
}
| Field | Behavior |
|---|---|
search_context_size |
low, medium (default) or high. Sets how many results each search returns when the model does not ask for a number: 3, 5 or 8. Echoed. |
user_location |
Accepted and echoed. Results are not localized. |
search_content_types |
Accepted and echoed on web_search_preview. Results are text. |
filters.allowed_domains |
Not supported: the search cannot be restricted to domains. Returns 400 invalid_request_error with code unsupported_parameter and param naming the field, for example tools[0].filters.allowed_domains. Remove it to search the whole web. |
external_web_access |
Search always reads the live web, so true is the only value served. false returns 400 with code unsupported_value and param tools[0].external_web_access. |
A web search tool cannot be forced: a tool_choice naming it
returns 400. The model decides when to search.
"tool_choice": "none", and an allowed_tools
choice without a {"type": "web_search"} entry (the tool's own
type), run no search. See
Server-Side Tools for the research
tools, the x_tools selection and the iteration cap.
Sources. With
"include": ["web_search_call.action.sources"], each completed
search action lists the public pages it returned as
sources: [{"type": "url", "url": ...}].
Citations. The answer cites a source as a Markdown link.
Each link to a page a search returned or the research opened carries a
url_citation annotation on the output_text part:
start_index and end_index span the link, counted
in Unicode code points, with the page's url and
title. A link to a page the research never returned is not
annotated. Retrieval returns the same annotations and sources.
Storage & Retention
Responses are stored by default ("store": true) using the
platform's standard two-tier storage architecture. Content is encrypted at
rest and retained based on the endpoint's service tier. For details on
storage tiers, encryption, retention, and billing, see
Storage.
Set "store": false to skip storage. The response is still
returned, but it is not retrievable by ID afterward, and its output,
instructions and tools are not kept with it; a
conversation the request names still receives the turn's
items. A record of the run remains: its status and any error, the model,
sampling settings, token counts and timings, and the metadata
and user values sent with the request.
A background response ("background": true) sent with
"store": false can be polled with
GET /v1/responses/{response_id}, as on OpenAI: it returns
the response's status while it is queued or in progress, and the finished
response, with its output, for about 10 minutes after it finishes. After
that it returns 404, exactly as an unknown ID does. Its output
is never written to storage, and it cannot be listed, chained from with
previous_response_id or read through input_items.
It can be cancelled while it is queued or in progress, and
deleting it discards its output at once. A zero-data-retention
workspace refuses background responses; see
Zero data retention.
Workspace scope
A stored response belongs to the workspace its create request resolved to: the
ws_ path prefix, else the X-Workspace-Id header, else
the only workspace of an API key confined to one, else the project's default
workspace. Listing, retrieving, deleting, cancelling and reading input items
resolve the workspace the same way and reach only that workspace's responses,
and previous_response_id must name one of them. A response in
another workspace answers 404, exactly as a missing one does.
- An API key confined to some of the project's workspaces reaches only those.
Naming another returns
403 Forbidden(workspace_not_allowed), and a key confined to several workspaces must name one. See Authentication & Security. - The project's default workspace is never an
operational workspace:
making one the default is refused with
400 Bad Request(operational_default_forbidden). In a project with no default workspace, a key that belongs to a user must name a workspace; a request that names none returns403 Forbidden(workspace_required). A key with no user behind it reaches every response of the project there. - Responses stored before responses recorded their workspace belong to none. Only a key with no user behind it and no workspace list reaches them, whichever workspace the request names. A confined key never does, and neither does a key that belongs to a user: nothing shows such a response lies in a workspace that user can read. Input items read with such a key leave out an older response or conversation that a response continued.
- A
conversationfollows the same rule; see Conversations.
Zero data retention
A workspace can be put into zero-data-retention mode. On this endpoint the mode changes what the API does rather than only what is kept on disk, so the rules are exact. What the mode suppresses everywhere else, what it keeps, and the other calls it refuses are on Zero Data Retention.
| Request | Result |
|---|---|
"store": true |
400, code retention_not_permitted, param of store. Storage cannot be bought back per request. |
store absent |
Accepted, and treated as false. The response is returned inline and is not retrievable by ID afterward. |
"store": false |
Accepted. Identical to the line above. |
"background": true |
400, code retention_not_permitted, param of background. A background run has no inline delivery: it requires server-side state to poll for, and that state is what the mode does not keep. |
An absent store is downgraded rather than refused on purpose.
The field defaults to true, so refusing the default would reject every
ordinary request and make the endpoint unusable in such a workspace. An
explicit true is a different thing: it asks for storage, and
the honest answer is no.
The same rule applies to store on
Chat Completions, which has no
background parameter to refuse.
Turning the mode on is forward-looking. It stops new content being
written; it does not delete responses stored before it was enabled. Use
DELETE /v1/responses/{response_id} or the data-retention
controls in your account settings for those. See
Forward-Looking.
Error Handling
Common Error Codes
| HTTP Status | Error Code | Description |
|---|---|---|
| 400 | invalid_request |
Missing or invalid parameters. |
| 400 | invalid_state |
Previous response is not in a terminal state, or response cannot be cancelled. |
| 400 | chain_depth_exceeded |
Response chain exceeds the maximum depth of 50. |
| 400 | retention_not_permitted |
The workspace runs in zero-data-retention mode and the request asked for store: true or background: true. See Zero data retention. |
| 401 | authentication_error |
Invalid or missing API key. |
| 403 | workspace_not_allowed |
The request names a workspace outside the API key's workspace list, or a key confined to several workspaces names none. See Workspace scope. |
| 403 | workspace_not_in_project |
The ws_ path prefix or X-Workspace-Id header does not name a workspace of the project. |
| 403 | workspace_required |
The project has no default workspace and the request, made with a key that belongs to a user, names none. See Workspace scope. |
| 404 | not_found |
Response or previous response not found, or it belongs to another workspace than the one the request resolves to. |
| 429 | rate_limit_exceeded |
Too many requests. Check Retry-After header. |
| 429 / 503 | capacity_exceeded |
No available workers. 429 when the request is refused at once, because the workspace sheds load or the request came through direct model dispatch; 503 when a queued request's wait times out or the queue is full. See Queueing. Check Retry-After header. |
| 503 | backend_unavailable |
No agent the endpoint is bound to can take traffic: it is offline, unhealthy or drained. Retry after a short delay. |
A response object that failed carries an error whose
code is one of the Response object's codes: a rate or
capacity refusal reads rate_limit_exceeded, and any other
failure, such as a timeout, reads server_error, with the
cause in message. HTTP error bodies keep the codes in the
table above.
A request that runs web search and fails after the research did work
answers with the header x-should-retry: false, so the OpenAI
SDKs do not retry it and run the research again. A refusal before any work,
such as capacity_exceeded with Retry-After, carries
no such header and the SDKs retry it as usual.
{
"error": {
"message": "Previous response is not complete: resp_abc123",
"type": "invalid_request_error",
"code": "invalid_state"
}
}
Client Integrations
opencode
OpenCode
supports the Responses API via the @ai-sdk/openai-compatible
adapter. Configure it in ~/.config/opencode/opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"erebine": {
"npm": "@ai-sdk/openai-compatible",
"name": "erebine",
"options": {
"baseURL": "https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
"headers": {
"Authorization": "Bearer ere_myproject_your_api_key"
}
},
"models": {
"my-model": {
"name": "llama-3.1-8b",
"reasoning": true,
"tool_call": true,
"tools": true
}
}
}
},
"model": "erebine/my-model"
}
See OpenCode Integration for full configuration details and troubleshooting.
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
api_key="ere_myproject_your_api_key"
)
# Non-streaming
response = client.responses.create(
model="llama-3.1-8b",
input="What is the capital of France?"
)
print(response.output[0].content[0].text)
# Streaming
stream = client.responses.create(
model="llama-3.1-8b",
input="Explain quantum computing in simple terms.",
stream=True
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
# Conversation chaining
first = client.responses.create(
model="llama-3.1-8b",
input="What is the capital of France?"
)
second = client.responses.create(
model="llama-3.1-8b",
input="What about Germany?",
previous_response_id=first.id
)
Node.js (OpenAI SDK)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.erebine.ai/proj_ABC123/my-endpoint/v1",
apiKey: "ere_myproject_your_api_key",
});
// Non-streaming
const response = await client.responses.create({
model: "llama-3.1-8b",
input: "What is the capital of France?",
});
console.log(response.output[0].content[0].text);
// Streaming
const stream = await client.responses.create({
model: "llama-3.1-8b",
input: "Explain quantum computing in simple terms.",
stream: true,
});
for await (const event of stream) {
if (event.type === "response.output_text.delta") {
process.stdout.write(event.delta);
}
}
curl (Non-Streaming)
curl -X POST https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What is the capital of France?"
}'
curl (Streaming)
curl --no-buffer -X POST \
https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses \
-H "Authorization: Bearer ere_myproject_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b",
"input": "What is the capital of France?",
"stream": true
}'
List and Retrieve
# List responses
curl https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses?limit=10 \
-H "Authorization: Bearer ere_myproject_your_api_key"
# Get a specific response
curl https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123 \
-H "Authorization: Bearer ere_myproject_your_api_key"
# Get input items
curl https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123/input_items \
-H "Authorization: Bearer ere_myproject_your_api_key"
# Cancel an in-progress response
curl -X POST https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123/cancel \
-H "Authorization: Bearer ere_myproject_your_api_key"
# Delete a response
curl -X DELETE https://api.erebine.ai/proj_ABC123/my-endpoint/v1/responses/resp_abc123 \
-H "Authorization: Bearer ere_myproject_your_api_key"