Private Agents
Bring your own hardware under the router. Mint a time-bounded join key, run the EIM agent, watch it transition from pending to active, and the router treats it as a first-class worker for the project that owns it. This page is the practical CRUD: join keys, enrollment, lifecycle, monitoring.
Overview
For a side-by-side of private agents vs. shared agents, see the Agent Types overview.
This page documents the project-scoped Frontend API (authenticated
with a project API key or dashboard session). The router also exposes
a parallel API-key-authed management surface at
/:project_id/v1/management/join-keys and
/:project_id/v1/management/agents; see the API reference
for details. Defaults may differ between the two surfaces.
Agent Fields
The following fields describe an enrolled agent. They are returned by the agent list and detail API endpoints.
| Field | Type | Description |
|---|---|---|
| id | string (UUID) | Unique agent identifier. |
| name | string | Human-readable name for the agent. Settable at enrollment and updatable via PATCH. |
| worker_id | string | Unique worker identifier used in the mesh protocol. |
| region | string | Region where the agent is located (1-24 ASCII characters, inherited from the join key used at enrollment). |
| max_concurrent_requests | integer | Admit slots: the concurrent inference requests this agent accepts, reported by the agent (safe slots times its over-provisioning factor). On create and update it is the ceiling the agent's admit slots are capped at; 0 means auto-configured by the platform. |
| safe_concurrent_requests | integer | Safe slots: the concurrent requests this agent's KV cache holds without preemption, as last reported; 0 until the agent reports. Read-only. Agents serving an anchor are routed by safe slots; other agents by admit slots. |
| description | string or null | Optional description or notes about this agent. Updatable via PATCH. |
| supported_tiers | array of strings or null | Service tier IDs this agent supports. When null, the platform uses defaults for the agent's accelerator type. Example: ["gpu_nvidia_shared", "self_hosted"]. |
| agent_public_key | string | Public key the agent authenticates the mesh connection with. Set at enrollment and immutable afterwards. |
| country_code | string | ISO 3166-1 alpha-2 country the agent serves from. An operator assertion: nothing on the wire proves where a host sits, so you set it and the response echoes back what was stored. Updatable via PATCH. |
| allow_model_preseed | boolean | Whether this agent may register models imported from its local disk. Defaults to false. Inference agents only. See Preseeding a Local Model. |
| status | string | Current lifecycle status. See Agent Lifecycle for all values. |
| last_heartbeat_at | string (ISO 8601) or null | Timestamp of the most recent heartbeat received from the agent. |
| enrolled_at | string (ISO 8601) or null | Timestamp when the agent successfully completed enrollment. |
| created_at | string (ISO 8601) or null | Timestamp when the agent record was created. |
| updated_at | string (ISO 8601) or null | Timestamp of the most recent change to the agent record. |
| agent_role | string or null | Role the agent enrolled as: inference or execution. Returned by the list and detail reads. Null on the create and PATCH responses, and null whenever the role cannot be read back - treat a missing role as "not proven", never as inference. |
| capability_manifest_hash | string or null | SHA-256 of the capability manifest an execution agent published at enrollment. Returned by the list and detail reads for execution agents only; null for inference agents, and null on the create and PATCH responses, where no manifest exists yet. Use it to confirm the agent you started published the manifest you expect. |
| events | array of objects or null | Recent agent events, newest first, embedded by the single-agent read so a dashboard renders the timeline in one round-trip. Null on the list response; page the dedicated events endpoint for the full history. |
Join Key Management
Join keys are secure tokens that authorize a backend agent to enroll in your project. Each key is tied to a specific project, carries a region assignment, and expires after a configurable TTL (maximum 1 hour). Keys can be single-use or multi-use.
API Endpoints
| Method | Endpoint | Description |
|---|---|---|
| POST | /:projectId/v1/management/join-keys | Create a new join key. |
| GET | /:projectId/v1/management/join-keys | List all join keys for the project (paginated). |
| DELETE | /:projectId/v1/management/join-keys/:keyId | Revoke a join key. Agents already enrolled are not affected. Returns 204 No Content on success. |
| GET | /:projectId/v1/management/tiers | List the platform's service tiers. The slug values are what supported_tiers accepts. |
Join Key Fields
The following fields are accepted when creating a join key (POST).
| Field | Type | Required | Description |
|---|---|---|---|
| name | string | Yes | Human-readable label for the join key (e.g. "datacenter-rack-3"). |
| region | string | Yes | Region to assign to agents enrolled with this key. Must be 1-24 ASCII characters (e.g. "us-east-1"). |
| expires_in_hours | integer or null | No | Time-to-live in hours. Defaults to the maximum (1 hour). The hard limit is 1 hour regardless of the value supplied. Values less than or equal to 0 are accepted but produce an already-expired key; supply a positive integer. |
| max_enrollments | integer or null | No | Maximum number of agents that may enroll with this key. Defaults to 1 (single-use). Note: the router-side management endpoint (/v1/management/join-keys) uses a different default (10), so always set max_enrollments explicitly when calling either surface. |
| router_addresses | array of strings or null | No | Router addresses for agents to connect to. When omitted, platform defaults are used. |
The response for a create request includes the full key value exactly once. It cannot be retrieved later because only the SHA-256 hash is stored.
The list and detail responses include the following join key fields:
| Field | Type | Description |
|---|---|---|
| id | string (UUID) | Unique join key identifier. |
| name | string | Human-readable label. |
| key_prefix | string | Visible prefix of the key for identification. The first 8 characters include the ejk_ tag and the leading characters of the project slug (e.g. for project slug myproj, the prefix renders as ejk_mypr****wxyz). The full key is never returned after creation. |
| region | string | Region assigned to agents enrolled with this key. |
| max_enrollments | integer | Maximum enrollments allowed. |
| current_enrollments | integer | Number of agents that have enrolled using this key so far. |
| expires_at | string (ISO 8601) | Timestamp when the key expires. |
| status | string | Current status: active, used, expired, or revoked. |
| can_enroll | boolean | Whether the key can still be used for enrollment (active, not expired, enrollments remaining). |
| created_at | string (ISO 8601) or null | Timestamp when the key was created. |
| storage_limit_enabled | boolean or null | Whether a storage limit was set on this key. The limit travels to the agent in its enrollment payload; it is a property of the key, not of the enrolled agent, so the agent object does not carry it. |
| storage_limit_bytes | integer (int64) or null | Storage ceiling in bytes applied to agents enrolled with this key. Relevant only when storage_limit_enabled is true. |
Join Key API Examples
In the examples below, replace proj_ABC123 with your
project external id (e.g. proj_a1b2c3d4). The project
external id is the leading path segment, and the API key must carry
the management scope.
Create a Join Key
curl -X POST https://api.erebine.ai/proj_ABC123/v1/management/join-keys \
-H "Authorization: Bearer ere_my-project_abc123" \
-H "Content-Type: application/json" \
-d '{
"name": "datacenter-rack-3",
"region": "us-east-1",
"expires_in_hours": 1,
"max_enrollments": 1
}'
import requests
headers = {
"Authorization": "Bearer ere_my-project_abc123",
"Content-Type": "application/json"
}
response = requests.post(
"https://api.erebine.ai/proj_ABC123/v1/management/join-keys",
headers=headers,
json={
"name": "datacenter-rack-3",
"region": "us-east-1",
"expires_in_hours": 1,
"max_enrollments": 1
}
)
result = response.json()
# Save the key value immediately, it cannot be retrieved later
print(f"Join Key: {result['key']}")
const response = await fetch(
"https://api.erebine.ai/proj_ABC123/v1/management/join-keys",
{
method: "POST",
headers: {
"Authorization": "Bearer ere_my-project_abc123",
"Content-Type": "application/json"
},
body: JSON.stringify({
name: "datacenter-rack-3",
region: "us-east-1",
expires_in_hours: 1,
max_enrollments: 1
})
}
);
const result = await response.json();
// Save the key value immediately, it cannot be retrieved later
console.log(`Join Key: ${result.key}`);
List Join Keys
curl https://api.erebine.ai/proj_ABC123/v1/management/join-keys \
-H "Authorization: Bearer ere_my-project_abc123"
Revoke a Join Key
curl -X DELETE https://api.erebine.ai/proj_ABC123/v1/management/join-keys/KEY_ID \
-H "Authorization: Bearer ere_my-project_abc123"
Agent API
Agents are enrolled automatically using a join key (see Agent Enrollment). Once enrolled, you can list, view, update, and remove agents via the following endpoints.
| Method | Endpoint | Description |
|---|---|---|
| GET | /:projectId/v1/management/agents | List all enrolled agents for the project (paginated). Supports ?status= filter. |
| GET | /:projectId/v1/management/agents/:agentId | Get full details of a specific agent. |
| PATCH | /:projectId/v1/management/agents/:agentId | Update agent name, description, or status (suspend/resume). |
| DELETE | /:projectId/v1/management/agents/:agentId | Remove an agent from the project. The stored status is set to disconnected and a disconnected event is emitted with reason=removed_by_user. The agent remains visible in the list under the disconnected status. Returns 204 No Content on success. |
| GET | /:projectId/v1/management/agents/:agentId/events | Retrieve the event log for a specific agent. |
PATCH Request Body
All fields are optional. Only provided fields are updated.
| Field | Type | Description |
|---|---|---|
| name | string or null | New human-readable name for the agent. |
| description | string or null | New description or notes for the agent. |
| allow_model_preseed | boolean or null | Grant (true) or revoke (false) local model preseeding for this agent. Granting it to an execution agent is refused. See Preseeding a Local Model. |
| status | string or null |
Status change request. Valid values:
|
Agent API Examples
List Agents
curl https://api.erebine.ai/proj_ABC123/v1/management/agents \
-H "Authorization: Bearer ere_my-project_abc123"
Get Agent Details
curl https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID \
-H "Authorization: Bearer ere_my-project_abc123"
Suspend an Agent
curl -X PATCH https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID \
-H "Authorization: Bearer ere_my-project_abc123" \
-H "Content-Type: application/json" \
-d '{"status": "suspended"}'
Resume an Agent
curl -X PATCH https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID \
-H "Authorization: Bearer ere_my-project_abc123" \
-H "Content-Type: application/json" \
-d '{"status": "active"}'
Get Agent Events
curl https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID/events \
-H "Authorization: Bearer ere_my-project_abc123"
Remove an Agent
curl -X DELETE https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID \
-H "Authorization: Bearer ere_my-project_abc123"
Agent Enrollment
To enroll a private agent:
- Create a join key in the dashboard or via the API (see Join Key API Examples).
- Save the full key value immediately, it is shown only once.
- Run the EIM agent with the
enrollsubcommand, passing the join key. - The agent authenticates with the router and establishes a secure encrypted connection.
- The agent transitions to
activestatus and begins sending heartbeats. It is then ready to receive inference requests.
# Enroll an EIM node. The key goes in the environment:
# a command line is visible to every account on the host.
EREBINE_AGENT_JOIN_KEY="ejk_my-project_..." erebine-eim-agent enroll
After enrollment, start the agent in serve mode to begin handling inference requests. See the EIM Deployment guide for detailed setup instructions.
Preseeding a Local Model
A preseeded model is one that is already on the agent's disk. Point the agent at the directory and it registers the model with the router: the model joins your project's model list, endpoints can serve it, and the weights never cross the network. The platform keeps no copy of the files.
Private inference agents only. Execution agents cannot hold models, so the permission below is refused for them.
1. Allow the Agent to Preseed
New agents are created with Allow model preseed already ticked; clear it on the create form if you would rather this agent never registered local weights. Agents created before this default shipped have it off. To change it on an enrolled agent, open its Actions menu, choose Edit, and set Allow model preseed there -- or use the API:
curl -X PATCH https://api.erebine.ai/proj_ABC123/v1/management/agents/AGENT_ID \
-H "Authorization: Bearer ere_my-project_abc123" \
-H "Content-Type: application/json" \
-d '{"allow_model_preseed": true}'
The CLI carries the same switch:
erectl agents AGENT_ID --update --allow-preseed,
and --no-allow-preseed to take it back.
The import is a local filesystem operation, so nothing checks the permission while it runs. The check happens when the agent registers the model with the router. Grant it first, or the registration is rejected and the model never appears.
2. Import the Model
The source must be a Hugging Face model directory: a
config.json plus at least one
.safetensors or .bin weight file. A
source that is a symlink into a Hugging Face cache is resolved
to the real files first.
Stop the agent before importing. The import takes exclusive access to the model cache and refuses to start while a running agent holds it.
sudo systemctl stop erebine-eim-agent
erebine-eim-agent import-model /srv/models/my-model \
--model-name my-model \
--link-mode auto
sudo systemctl start erebine-eim-agent
The command prints the model id, the link mode it resolved, the file count, and the size. That id is the id the model carries in your project once it is registered.
| Argument | Description |
|---|---|
<path> |
Required. Path to the model directory. |
--model-name |
Name shown in your model list. Defaults to the source directory name. |
--link-mode |
auto, hardlink, copy, or symlink. Defaults to auto. |
--model-cache-path |
Cache path to import into. Defaults to EREBINE_AGENT_MODEL_CACHE_PATH, then to the agent's built-in path. |
--log-level |
trace, debug, info, warning, or error. Defaults to info. |
Link Modes
The link mode decides how the files get into the cache, and what it costs you in disk.
| Mode | Behavior |
|---|---|
auto |
Hardlink, falling back to a copy when the source and the cache are on different filesystems. This is the default. |
hardlink |
Hardlink only. Fails when the source and the cache are on different filesystems. |
copy |
Always copy. Consumes cache space; a copy larger than the space left is refused before any bytes move. |
symlink |
Symlink into the source. Costs nothing, and the model breaks if you move or delete the source. |
Importing the same directory twice is a no-op. A source whose files already match a cache entry returns that entry and writes nothing.
3. Start the Agent
The agent registers its imported models when it starts and again on every reconnect. Once the router accepts one, it appears in the project's model list with a Preseeded badge; see Model Management for what you can do with it from there.
A rejection is logged with its reason: the agent is not allowed to preseed, the name is already taken by another model in the project, or the model was deleted from the project. In the delete case the agent purges its own copy; in the others, fix the cause and the next registration goes through.
Importing from the macOS App
The macOS application has an Import Model pane: pick the directory, pick the link mode, watch the checksum progress. The app owns its agent, so you do not stop anything first. With the agent running the model is registered immediately; with it stopped, at the next start. See EIM on macOS.
Importing in a Container
The import must run inside the agent's mount namespace, or it writes to a cache the agent never reads. It also cannot run while the agent holds that cache. So: stop the agent container, then run the import as a one-shot container carrying the same cache volume plus the source directory. The image entrypoint passes the arguments straight through.
docker run --rm \
-v MODEL_CACHE_VOLUME:/var/lib/inference/.cache/erebine/models \
-v /srv/models/my-model:/models/my-model \
ghcr.io/erebine/container-agents/eim-vllm-cuda:latest \
import-model /models/my-model
Two constraints follow the mounts. Hardlinking needs the
source and the cache on one filesystem, so a source bind-mounted
from elsewhere lands on copy instead. And the files
have to be readable by the UID the agent runs as.
Listing and Removing Imported Models
The agent can tell you what it holds and where each model stands with the router. Safe to run while the agent is running:
erebine-eim-agent list-models
Each model reports registered (the router
accepted it and it is in your model list),
refused with the reason, or
not announced when no verdict has arrived yet.
Removal has two halves, and they are not interchangeable. Deleting the model on the platform is the complete removal: the router tombstones it and asks every agent holding it to purge, and an agent that misses that command self-purges on its next announce. Do that from your model list, the desktop app's Import Model pane, or the API.
Removing the local copy deletes the files on one machine and nothing else -- the model stays in your project's list, pointing at bytes that are gone, until you delete it on the platform too. An agent holds no management credential, so this is all the CLI can do:
erebine-eim-agent remove-model MODEL_ID
The agent must be stopped for remove-model (it
holds the cache lock while running). Source files outside
the model cache are never touched, so a hardlink or symlink
import leaves your originals alone. Models the platform
downloaded are refused: removing one only forces the next
provision to download it again.
What a Preseeded Model Does Not Do
- It is not shared. The weights stay on the agents that imported them. The platform makes no copy, publishes nothing to the catalog, and refuses share and unshare.
- Nothing reads the files in the background. Revalidation and metadata refresh are refused, and the background sweeps skip preseeded models. The only reader is the engine serving it.
- It has no versions. Creating, promoting, and rolling back a version are refused.
- It runs only where its bytes are. Endpoints can only bind agents that hold the files. The endpoint wizard blocks the other pairings and names the holders; creating or attaching past it is refused.
- It uses no storage quota. Your quota counts bytes the platform stores, and it stores none of these.
Metadata is the exception: it stays yours to edit and it still drives what the agent launches the engine with. See Model Management.
Deleting a Preseeded Model
- Unbind it first. Delete is refused while an endpoint or an anchor endpoint binds the model.
- Deleting removes it everywhere. The model leaves your project and every agent holding the files purges its cached copy. An agent that was offline at the time purges on its next registration attempt.
- The purge stays inside the agent's model cache. Your source directory is untouched.
- Importing the same directory again creates a new model with a new id. Deleting bans the entry, not the files.
Agent Monitoring
The agent dashboard at /agents provides real-time visibility
into your enrolled agents:
- Heartbeat status: whether each agent is actively communicating with the router.
- Current load: active request count and capacity.
- Model information: which model is loaded and its serving state.
- Event history: per-agent log of significant lifecycle events.
Programmatic Access
For programmatic snapshot access to your fleet, use the documented
REST endpoint
GET /:projectId/v1/management/agents
(see Agent API) with a project API key.
Heartbeat Tracking
The router tracks the last_heartbeat_at timestamp for each
agent. An agent with a heartbeat older than the staleness threshold
(default 30 seconds; the deployment sets it) is considered
disconnected and excluded from routing decisions until it reconnects and
resumes sending heartbeats. The effective status shown in the dashboard
and API is derived from both the stored status and heartbeat freshness.
Agent Controls
Each row's Actions menu carries three verbs beyond Edit and Delete. All three act on agents you own; none of them re-enrol the agent or touch its downloaded models.
| Action | Endpoint | Effect |
|---|---|---|
| Reset agent | POST /agents/AGENT_ID/reset |
The agent drops its endpoint assignment and its pinned model, restarts, and re-attaches from the router. Cached model files and enrolment survive, so recovery costs no re-download. In-flight requests are lost. Use this when the agent stays connected but will not serve. |
| Disable agent | POST /agents/AGENT_ID/disable |
Parks the agent: it stops serving and frees its accelerator, but stays enrolled and keeps renewing its lease. A maintenance window opens for the duration, so the interval is excluded from the agent's uptime. Refused while the agent is suspended. |
| Enable agent | POST /agents/AGENT_ID/enable |
Reverses a disable and closes the maintenance window. The agent's endpoints return to service on the router's next health scan. Replaces the Disable entry in the menu while the agent is disabled. |
Reset and Disable both ask for confirmation in the dashboard. How an agent comes back from a Reset depends on how it runs: a container with a restart policy or a systemd unit starts a fresh process; the macOS desktop app hosts the agent itself and cycles it in place, without closing the app; an agent run in the foreground stays stopped until you start it again.
Agent Lifecycle
Agents transition through the following status states. The status
field in the database reflects the stored lifecycle state. The effective_status
is computed at runtime from the stored status and heartbeat freshness.
| Status | Description |
|---|---|
pending |
Agent record has been created (join key issued) but the agent has not yet connected and authenticated. |
active |
Agent is connected and sending regular heartbeats. If the most recent heartbeat is older than the heartbeat staleness threshold (default 30 seconds; the deployment sets it), the effective status is disconnected even while the stored status remains active. |
disconnected |
Stored after the heartbeat-timeout worker observes a stale agent or after a DELETE remove call; also computed as the effective_status when the stored status is active but the most recent heartbeat is stale. The router excludes the agent from routing until it reconnects and heartbeats resume. The agent may reconnect and return to active. |
suspended |
Agent has been suspended by the project owner. Suspended agents do not receive new requests. Resume with PATCH status=active. |
dead |
Terminal state. The agent failed to come online within the join key lifespan or was declared dead by the platform. Dead agents do not receive requests and cannot be revived. |
Allowed Transitions
| From | To | Trigger |
|---|---|---|
pending |
active |
Agent connects and completes enrollment handshake. |
pending |
dead |
Join key expired before the agent connected. |
active |
suspended |
Owner calls PATCH with status=suspended. |
active |
disconnected |
Heartbeat timeout detected by the router. |
active |
dead |
Platform declares agent dead after extended absence. |
disconnected |
active |
Agent reconnects and heartbeats resume. |
disconnected |
suspended |
Owner suspends a disconnected agent. |
disconnected |
dead |
Platform declares agent dead after extended absence. |
suspended |
active |
Owner calls PATCH with status=active. |
suspended |
dead |
Platform declares agent dead after extended absence. |
State Diagram
stateDiagram-v2
[*] --> pending: join key issued
pending --> active: enrollment handshake
pending --> dead: join key expired
active --> suspended: PATCH status=suspended
active --> disconnected: heartbeat timeout
active --> dead: extended absence
disconnected --> active: heartbeat resumes
disconnected --> suspended: owner suspends
disconnected --> dead: extended absence
suspended --> active: PATCH status=active
suspended --> dead: extended absence
dead --> [*]
Single-Agent Queuing
When only one agent is available for an endpoint's tier, the router queues incoming requests instead of immediately returning a 503 error. This matters for EIM deployments with a single backend agent.
The queue timeout is 30 seconds by default; the deployment sets it. If the agent does not become available within the timeout, the request fails with a timeout error.