Zero Data Retention
A per-workspace mode. Nothing the model derives from your turns is written. It changes what lands on disk, and in a small, exact set of places it changes what the API answers.
Overview
Zero data retention is a property of a workspace, not of a request or an API key. Turn it on for a workspace and every request that attributes to that workspace runs under it: the completions and responses it serves, the chats held in it, and every background service that would otherwise read a turn and write something derived from it.
The mode is not inherited. A workspace is in it or it is not; turning it on for one workspace does not change any other, and a new workspace does not pick it up from its neighbours.
A request that cannot be attributed to any workspace is treated as zero-retention. Storage needs a workspace to attribute to, and an unresolvable one resolves to the stricter answer rather than the convenient one.
What Stops
Nothing the model derives from your turns is stored. Concretely:
- Messages, their embeddings, rolling summaries and the archived activity blob are not written.
- Completions and responses are not persisted, so they are not retrievable by id afterwards.
These workspace services stop running for the workspace:
- Automatic project intelligence extraction and the workspace graph.
- Evolution and MCP guidance.
- Voice learning.
- Chat memory and execution memory.
- Document indexing.
- Research payloads.
- Mockup bundles.
- Link unfurling.
Each of these is stopped before the work runs, not after. A service that generated its output and then discarded it would still have paid for the inference and still have put the content in flight.
What Is Kept
"Zero data retention" is a claim about turn-derived content. It is not a claim that the workspace holds nothing. A reader who finds an uploaded file still there deserves to have been told, so this list is exhaustive at the store level.
| Kept | Why |
|---|---|
| The chat and its name | The container, not the content. Turns inside it are not written. |
| Your uploads and the documents you add | You put them there deliberately. Delete them the way you added them. |
| Token counts and timings | Billing. Content-free. |
| A content-scanning hash and verdict | Scanning still runs. It records a hash and a pass or fail, never the text. |
| The execution audit trail, reduced | Tool name, outcome and timing only. Arguments and results are dropped. |
| A query embedding, for up to 60 seconds | The router caches an embedding of the turn to route it. The write hard-caps its own lifetime at 60 seconds. |
A Responses run's metadata and user |
The run record keeps the two fields you sent until it expires, 90 days by default. Instructions, tools and output are not written. |
Forward-Looking
DELETE /v1/responses/{response_id},
DELETE
/v1/chat/completions/{id}, the workspace file controls, or the retention
controls in your account settings to remove what already exists.
API Refusals
Most of what the mode does is invisible from outside: rows that are not
written, workers that do not run. The exception is the set below, where a
request asks the router to store something and gets an honest refusal
instead of a silent no-op. A caller that believed its write succeeded would
be worse off than one that got a 400.
Every one of them carries the same error code,
retention_not_permitted, so one client-side branch covers all
of them.
| Request | Answer |
|---|---|
Chat completions with "store": true |
400 retention_not_permitted |
Responses with "store": true |
400 retention_not_permitted, param of store |
Responses with "background": true |
400 retention_not_permitted, param of background. A background run has no inline delivery: it hands back a handle to poll, and the state behind that handle is what the mode does not keep. |
| Creating a batch job | 400 retention_not_permitted. Batch is store-and-forward: the request bodies and the result file must both be persisted for the job to be fetchable. Send the requests to /v1/chat/completions, /v1/responses or /v1/embeddings instead. |
| Creating or editing a chat memory or a workspace memory | 400 retention_not_permitted. A memory is turn-derived content. |
| Persisting a generated chat artifact | 400 retention_not_permitted. A generated artifact is model output. |
| Filing an execution learning | 400 retention_not_permitted. A learning is distilled from what a tool run produced. |
| Rolling evolution guidance back to a loop-authored version | 422 retention_not_permitted. The version exists in the changelog, so 404 would be a lie; restoring it would re-write loop-generated content into a workspace that does not keep it. |
None of these name a parameter you could change to make them succeed, apart
from store and background. There is no variant of
the request that a zero-retention workspace accepts.
The store Rule
store is the one field whose treatment differs between an
explicit value and an absent one, so it is worth stating exactly.
| You send | Result |
|---|---|
"store": true |
Refused. Storage cannot be bought back per request. |
"store": false |
Served. The response is returned inline and is not retrievable by id afterwards. |
store absent, on Responses |
Served, and treated as false. Identical to the line above. |
An absent store is downgraded rather than refused because the
field documents a default of true on the Responses API. Refusing the default
would reject every ordinary request and make the endpoint unusable in such a
workspace. An explicit true is a different thing: it asks for
storage, and the honest answer is no.
Refusing rather than ignoring matters for the same reason. An ignored
store: true would leave an OpenAI-compatible client
404ing on its follow-up GET, which is a worse
failure than a 400 at the point of the mistake.
{
"error": {
"message": "This workspace runs in zero-data-retention mode and cannot store responses. Remove `store` or send store=false.",
"type": "invalid_request_error",
"param": "store",
"code": "retention_not_permitted"
}
}
Turning It On
In the chat workspace settings panel, the Zero data retention card holds the switch. Only a workspace owner or manager sees a live switch; everyone else sees the current state read-only.
Over the API, PATCH the workspace:
curl -X PATCH "https://api.erebine.ai/proj_ABC123/v1/workspaces/ws_abc123" \
-H "Authorization: Bearer $EREBINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"zero_data_retention": true}'
Reading the Mode
The workspace document carries the flag, so a client can branch on it before it sends a request it knows will be refused.
{
"id": "ws_abc123",
"name": "Regulated",
"zero_data_retention": true,
"inference_queueing_enabled": false
}
zero_data_retention is always present and always a boolean.
inference_queueing_enabled is the separate
queueing toggle; the two are unrelated
settings that happen to sit on the same document.
Publishing the Claim
The provider manifest publishes a
compliance.zdr boolean per model entry. It is derived from the
retention mode of the workspace the manifest request binds, not set as an
independent flag. Bound to a zero-retention workspace every entry publishes
true; bound to any other workspace, every entry publishes
false.
The URL decides which workspace that is. The workspace-scoped form,
/proj_ABC123/WORKSPACE_ID/_direct/v1/provider-manifest,
publishes the retention posture of the workspace named in the path; the
project-scoped form publishes the project's default workspace's. Publish a
listing from the workspace-scoped URL and a marketplace reads the posture of
the workspace you named.
That derivation is the point. The claim a marketplace republishes on your behalf and the behaviour the router enforces are one fact, so they cannot drift apart.
Setting the mode does not on its own make the claim readable. The manifest
answers 401 to a fetch carrying no API key until the workspace
that URL binds turns on marketplace_manifest_public, and a
marketplace polls with no credential. See
Serving It Without a Key.