docs
// Platform

Zero Data Retention

A per-workspace mode. Nothing the model derives from your turns is written. It changes what lands on disk, and in a small, exact set of places it changes what the API answers.

Overview

Zero data retention is a property of a workspace, not of a request or an API key. Turn it on for a workspace and every request that attributes to that workspace runs under it: the completions and responses it serves, the chats held in it, and every background service that would otherwise read a turn and write something derived from it.

The mode is not inherited. A workspace is in it or it is not; turning it on for one workspace does not change any other, and a new workspace does not pick it up from its neighbours.

A request that cannot be attributed to any workspace is treated as zero-retention. Storage needs a workspace to attribute to, and an unresolvable one resolves to the stricter answer rather than the convenient one.

What Stops

Nothing the model derives from your turns is stored. Concretely:

  • Messages, their embeddings, rolling summaries and the archived activity blob are not written.
  • Completions and responses are not persisted, so they are not retrievable by id afterwards.

These workspace services stop running for the workspace:

Each of these is stopped before the work runs, not after. A service that generated its output and then discarded it would still have paid for the inference and still have put the content in flight.

What Is Kept

"Zero data retention" is a claim about turn-derived content. It is not a claim that the workspace holds nothing. A reader who finds an uploaded file still there deserves to have been told, so this list is exhaustive at the store level.

Kept Why
The chat and its name The container, not the content. Turns inside it are not written.
Your uploads and the documents you add You put them there deliberately. Delete them the way you added them.
Token counts and timings Billing. Content-free.
A content-scanning hash and verdict Scanning still runs. It records a hash and a pass or fail, never the text.
The execution audit trail, reduced Tool name, outcome and timing only. Arguments and results are dropped.
A query embedding, for up to 60 seconds The router caches an embedding of the turn to route it. The write hard-caps its own lifetime at 60 seconds.
A Responses run's metadata and user The run record keeps the two fields you sent until it expires, 90 days by default. Instructions, tools and output are not written.

Forward-Looking

Turning the mode on deletes nothing. It stops new content being written from that moment. Anything stored before you turned it on is still stored, and is still retrievable, until you delete it. Use DELETE /v1/responses/{response_id}, DELETE /v1/chat/completions/{id}, the workspace file controls, or the retention controls in your account settings to remove what already exists.

API Refusals

Most of what the mode does is invisible from outside: rows that are not written, workers that do not run. The exception is the set below, where a request asks the router to store something and gets an honest refusal instead of a silent no-op. A caller that believed its write succeeded would be worse off than one that got a 400.

Every one of them carries the same error code, retention_not_permitted, so one client-side branch covers all of them.

Request Answer
Chat completions with "store": true 400 retention_not_permitted
Responses with "store": true 400 retention_not_permitted, param of store
Responses with "background": true 400 retention_not_permitted, param of background. A background run has no inline delivery: it hands back a handle to poll, and the state behind that handle is what the mode does not keep.
Creating a batch job 400 retention_not_permitted. Batch is store-and-forward: the request bodies and the result file must both be persisted for the job to be fetchable. Send the requests to /v1/chat/completions, /v1/responses or /v1/embeddings instead.
Creating or editing a chat memory or a workspace memory 400 retention_not_permitted. A memory is turn-derived content.
Persisting a generated chat artifact 400 retention_not_permitted. A generated artifact is model output.
Filing an execution learning 400 retention_not_permitted. A learning is distilled from what a tool run produced.
Rolling evolution guidance back to a loop-authored version 422 retention_not_permitted. The version exists in the changelog, so 404 would be a lie; restoring it would re-write loop-generated content into a workspace that does not keep it.

None of these name a parameter you could change to make them succeed, apart from store and background. There is no variant of the request that a zero-retention workspace accepts.

The store Rule

store is the one field whose treatment differs between an explicit value and an absent one, so it is worth stating exactly.

You send Result
"store": true Refused. Storage cannot be bought back per request.
"store": false Served. The response is returned inline and is not retrievable by id afterwards.
store absent, on Responses Served, and treated as false. Identical to the line above.

An absent store is downgraded rather than refused because the field documents a default of true on the Responses API. Refusing the default would reject every ordinary request and make the endpoint unusable in such a workspace. An explicit true is a different thing: it asks for storage, and the honest answer is no.

Refusing rather than ignoring matters for the same reason. An ignored store: true would leave an OpenAI-compatible client 404ing on its follow-up GET, which is a worse failure than a 400 at the point of the mistake.

Refusal body
{ "error": { "message": "This workspace runs in zero-data-retention mode and cannot store responses. Remove `store` or send store=false.", "type": "invalid_request_error", "param": "store", "code": "retention_not_permitted" } }

Turning It On

In the chat workspace settings panel, the Zero data retention card holds the switch. Only a workspace owner or manager sees a live switch; everyone else sees the current state read-only.

Over the API, PATCH the workspace:

cURL
curl -X PATCH "https://api.erebine.ai/proj_ABC123/v1/workspaces/ws_abc123" \ -H "Authorization: Bearer $EREBINE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"zero_data_retention": true}'

Reading the Mode

The workspace document carries the flag, so a client can branch on it before it sends a request it knows will be refused.

GET /proj_ABC123/v1/workspaces/{workspace_id}
{ "id": "ws_abc123", "name": "Regulated", "zero_data_retention": true, "inference_queueing_enabled": false }

zero_data_retention is always present and always a boolean. inference_queueing_enabled is the separate queueing toggle; the two are unrelated settings that happen to sit on the same document.

Publishing the Claim

The provider manifest publishes a compliance.zdr boolean per model entry. It is derived from the retention mode of the workspace the manifest request binds, not set as an independent flag. Bound to a zero-retention workspace every entry publishes true; bound to any other workspace, every entry publishes false.

The URL decides which workspace that is. The workspace-scoped form, /proj_ABC123/WORKSPACE_ID/_direct/v1/provider-manifest, publishes the retention posture of the workspace named in the path; the project-scoped form publishes the project's default workspace's. Publish a listing from the workspace-scoped URL and a marketplace reads the posture of the workspace you named.

That derivation is the point. The claim a marketplace republishes on your behalf and the behaviour the router enforces are one fact, so they cannot drift apart.

Setting the mode does not on its own make the claim readable. The manifest answers 401 to a fetch carrying no API key until the workspace that URL binds turns on marketplace_manifest_public, and a marketplace polls with no credential. See Serving It Without a Key.