docs
// Tools

erectl deploy

One command from a bare Linux host to enrolled agents. erectl deploy agent eim and erectl deploy agent eem detect the container runtime (and, for inference, the hardware), mint a join key, generate the compose stack, optionally install systemd units, start the agents, and verify enrollment end to end. With --runtime baremetal they install the agent binary and run it under systemd instead.

Overview

The manual recipe in the container-agents compose files -- directories, sysctl, join key, the right file for your hardware -- is fully automated:

bash
sudo erectl deploy agent eim \
  --auto-hardware-detect --count 2 --service \
  --region us-east

The command plans first, shows you every host mutation it will make, and asks before proceeding. --dry-run prints the plan and every command it would run without touching the host, and is the one mode that does not need root. Re-running with the same flags converges: directories are kept, the compose stack reconciles, and a still-valid join key is reused rather than re-minted.

Kinds

The kind is required. An inference stack and an execution agent are different host mutations, and a root-privileged command does not guess between them: the deploy agent group with no kind prints the two kinds, and the same flags with no kind in front of them are a usage error rather than a default.

KindDeploys
eimN inference agents running vLLM on GPU or CPU. Hardware detection, GPU strategies, MIG orchestration, kernel tuning.
eemOne execution agent running tools. No GPU, no kernel tuning; it serves workspaces rather than models.

Both kinds coexist on one host. Each keeps its own deploy directory, systemd unit, data directory and active-deploy pointer, so deploying one never touches the other's stack, and each has its own --down and --update. Pointing --deploy-dir at the other kind's directory is refused rather than crossed.

A join key carries the role it was minted for. A key minted for an inference agent cannot deploy an execution agent or the reverse: erectl refuses locally before touching the host, and the router refuses the enrollment too.

Prerequisites

  • Linux with systemd, and root (via sudo).
  • podman (with podman-compose) or docker (with the compose v2 plugin). podman is preferred when both are present; override with --runtime. With neither, the command stops and points at --runtime baremetal; it never switches runtime on its own.
  • For eim on NVIDIA GPUs: the NVIDIA Container Toolkit. The command verifies it before changing anything. eem needs no GPU and no toolkit.
  • An erectl configuration with a management-scope API key (see erectl setup). Under sudo, root's configuration is the one that is read; pass --base-url and --api-key explicitly if root has none.
  • No agent of the same kind already running under another unit. Every deploy, container or bare metal, refuses while any enabled or running systemd unit other than its own runs that kind's agent binary (erebine-eim-agent or erebine-eem-agent), whatever the unit is called -- the deb and rpm packages ship such units -- because two agents on one host would enroll twice. systemctl disable --now the unit the error names and re-run.

Quick Start

Minting a join key needs one thing: the region. For eim the tier defaults to self_hosted, which makes the agents user-owned -- serving your project, not shared platform traffic; pass --tier to serve other tiers (list the valid slugs with erectl tiers). The router address defaults to tcp://<host>:5555 derived from your configuration's base_url; pass --router-addr when the agent plane lives elsewhere. Alternatively, bring your own key with --join-key-file (or - for stdin) and skip minting entirely.

bash
# Detect hardware, deploy one user-owned inference agent, no systemd unit:
sudo erectl deploy agent eim --auto-hardware-detect --region us-east

# Preview only (no root needed):
erectl deploy agent eim --auto-hardware-detect --dry-run --region us-east

# Serve a specific tier instead of the self_hosted default:
sudo erectl deploy agent eim --auto-hardware-detect \
  --region us-east --tier priority

# Agent plane on a different host than the API:
sudo erectl deploy agent eim --auto-hardware-detect \
  --region us-east --router-addr tcp://agents.example.com:5555

One execution agent, with a systemd unit so it comes back after a reboot:

bash
# Deploy, enroll and verify one execution agent:
sudo erectl deploy agent eem --service --region us-east

# Preview only (no root needed):
erectl deploy agent eem --dry-run --region us-east

# Name the registration and the project explicitly:
sudo erectl deploy agent eem --service --region us-east \
  --registration-name eem-fleet-01 --project proj_ABC123

An omitted --project defaults to the project your base_url is scoped to; an omitted --registration-name defaults to eem-<hostname>, which must be unique within the project.

Join keys expire after one hour, so the command pulls images first and mints as late as possible. After a verified deploy it prints the exact command to revoke the key -- for both kinds, enrollment state persists in the agent's data directory, so the key is only needed once.

Bare Metal

No container runtime on the host, or no reason for one: --runtime baremetal installs the agent binary and runs it under systemd, behind the same unit names, lifecycle flags and verification gates as a container deploy. It is never chosen automatically.

bash
# Inference agents on the host, running your vLLM:
sudo erectl deploy agent eim --runtime baremetal --service \
  --auto-hardware-detect --region us-east --vllm-path /opt/vllm

# One execution agent:
sudo erectl deploy agent eem --runtime baremetal --service --region us-east

--service is required: systemd is the only supervisor a bare-metal agent has. Linux on x86_64 only, the one architecture every release channel publishes.

Install method

The release defaults to latest, the same default the container path uses for its image tag; pin one with --image-version 2.4.0. Bare metal needs agent release 2.4.0 or newer: an older pinned release is refused, and so is a brew tap that still pins one. The binary comes from the first of these that applies, or from --install-method:

  1. Already installed at the wanted version: reused, nothing downloaded.
  2. deb on hosts with apt, rpm on hosts with dnf, yum or zypper.
  3. brew, the erebine/homebrew-tap formula, run as the account that owns Homebrew. latest only: the tap pins one version.
  4. binary, the release asset, installed to /usr/local/bin.

Every download lands in a fresh directory only root can reach, under /var/cache/erebine (a /var/cache/erebine that is not a root-owned directory only root can write is refused, not repaired), and is checked against the sha256 digest the release publishes before anything installs it; a mismatch stops the deploy, and the directory is removed either way. After the install the binary has to start, so a missing shared library fails the deploy with the packages to install instead of failing the first boot.

On RHEL, Rocky Linux and AlmaLinux the rpm's dependencies (zeromq, libsodium) come from EPEL: enable it (dnf install epel-release) before the deploy, or the install stops on unresolved dependencies. The plan says so on those hosts.

The deb and rpm packages ship units of their own that run the same binary. The deploy never enables them, and refuses while any unit running the agent binary is enabled or running, as every deploy does (see Prerequisites); it checks again after installing a package.

vLLM

The inference agent runs your vLLM; erectl never installs Python. It uses vllm on PATH, or --vllm-path: the executable, or the Python environment that holds bin/vllm. Under sudo PATH is root's, so pass the flag. The version is read from the environment without starting Python. A version other than the one this release is built and tested with (vLLM 0.30.0) is a plan warning; no vLLM refuses the deploy.

The deploy writes the agent images' vLLM and LMCache launchers into <deploy-dir>/bin, bound to your environment's interpreter, so a bare-metal agent keeps the image's offline defaults and usage accounting. Nothing is written into your environment. Engine settings that belong to a particular build -- a tcmalloc LD_PRELOAD on CPU, ROCm tuning flags, CUDA_HOME for a host CUDA toolkit -- are yours to set in agents.local.env.

The agents' account must be able to read the environment. A venv under a home directory closed to other users is refused by name; /opt is the usual place.

Units and agents

erebine-agents.service is still the unit you operate. On bare metal it starts one erebine-agents@<n>.service per agent, and each instance stops and restarts with it: systemctl stop erebine-agents stops every agent, and journalctl -u 'erebine-agents@*' -f follows them all. The execution agent runs as erebine-eem.service.

A deploy clears any failed instance, enables the stack unit and restarts it when it is not running or when a generated file or the binary changed; otherwise it starts only what is stopped, such as an instance that had failed, and leaves the running agents alone. It then fails unless the unit is active. Over a running stack, --restart restarts it regardless and --no-restart leaves the running agents as they are. The header comment erectl writes at the top of each file does not count as a change, so re-running with the flags in another order restarts nothing.

What a container gives each agent, each instance gets explicitly:

ResourceBare metal
Accounterebine-eim (uid 5152) or erebine-eem (uid 5153), created when absent. The uids are the container images', so the data directories need no re-chown.
Home and state<data-dir>/agent<n>, the only directory the agent can write. Models and the LMCache disk tier stay in agent<n>/models and agent<n>/lmcache, where a container deploy keeps them.
PortsMetrics on 127.0.0.1:9100+n; LMCache, when enabled, on 6555+n and 8090+n. The execution agent keeps 127.0.0.1:9095 for metrics and health, so both kinds share a host without a clash.
DevicesThe GPU strategy's plan through CUDA_VISIBLE_DEVICES or HIP_VISIBLE_DEVICES; a MIG agent's slice is looked up at every start.
/tmpPrivate to each instance.
Memory limitsLocked memory and address space unlimited (LimitMEMLOCK=infinity, LimitAS=infinity), where a container gets an unlimited memlock ulimit.
CredentialsThe execution agent's --credentials-dir is owned by root, readable by the agent's group, and read-only in its unit, as the container mounts it.

Units run hardened: read-only system, private /tmp, no new privileges. On AMD GPUs the agents join the host's video and render groups -- only the ones the host has. Device selection is by those variables, not a device cgroup, so it is not the isolation a pinned container gives.

Everything else is the container path's: sysctl and transparent hugepages, MIG carving and its boot unit, the device-readiness gate, the verification gates, --no-start, --update, --down and --revoke-key. The in-container device check is skipped: there is no container to look into. --update installs a newer release through the method the deploy used, the way a container update pulls a newer image, and --down keeps the installed binary and the account, as it keeps images. To change runtime, tear the running deployment down first; a deploy over it with the other runtime is refused.

Hardware Types

eim only; an execution agent runs no model and needs no accelerator.

TypeImageDetected by
nvidia-gpueim-vllm-cudanvidia-smi lists a GPU
amd-gpueim-vllm-rocm/dev/kfd with a compute node
amd-cpueim-vllm-zendnnAMD CPU, no GPU
cpueim-vllm-cpueverything else

Pass --hardware-type to skip detection. The plan output always states what was detected and why.

GPU Strategies

eim only. With --count above one on a GPU host, --gpu-strategy picks the shape:

  • pinned (default): one GPU per agent, tensor parallelism 1. No GPU-to-GPU interconnect needed -- the right shape for VFIO passthrough and for models that fit a single GPU. Higher aggregate throughput than one tensor-parallel group.
  • shared (NVIDIA only): every agent sees every GPU with memory split evenly. For many small models per board; agents contend for the same devices.
  • mig (NVIDIA only): hard-partitioned MIG slices, one per agent. Deterministic startup and real isolation; see below.

A single agent (--count 1) always uses the canonical single-agent shape with tensor parallelism derived from every visible GPU, whatever the strategy -- except mig, which always pins its slice.

When --count exceeds the GPUs on the host and no strategy was passed, the default falls back from pinned to shared automatically: all agents run against the available GPU(s) with vLLM memory limits split evenly (0.90/count each) -- multiple agents per GPU with no MIG required. The plan labels the fallback; an explicit --gpu-strategy pinned refuses instead.

MIG Orchestration

eim only. Compose files cannot create MIG instances; erectl can. --mig (shorthand for --hardware-type nvidia-gpu --gpu-strategy mig) with --mig-profile handles the whole lifecycle:

bash
sudo erectl deploy agent eim --mig --mig-profile 2g.12gb --count 2 --service \
  --region us-east
  • Verifies the boards support MIG and the requested profile fits, per GPU, before changing anything.
  • Adopts an existing matching layout as-is; refuses to destroy a different one without --mig-repartition.
  • Enables MIG mode where needed (never resetting a busy GPU), carves the instances, and regenerates the CDI spec.
  • Installs a boot service that re-creates the layout: MIG instances never survive a reboot, and on Hopper and newer boards MIG mode itself does not either. This is why mig requires --service.

Profile names are per board -- run nvidia-smi mig -lgip to see yours. Use the name (like 2g.12gb), never the numeric ID.

Kernel Settings

eim only; an execution agent leaves the kernel alone. Every inference deploy ensures three settings in /etc/sysctl.d/99-erebine.conf and applies the file:

SettingValue
net.core.rmem_max, net.core.wmem_maxAt least 4194304
vm.max_map_countAt least 262144

Each is only ever raised. A value already higher in the file is kept, and vm.max_map_count is written only while the host runs below 262144: a host that runs higher, such as an Elasticsearch host at 1048576, is left alone, and the file never carries a value that would lower it. The plan shows the current value and the target, or that the host is already at or above it. The command warns when a file that sorts after 99-erebine.conf sets a lower value, because that file wins at boot. --no-sysctl skips the step on the single-agent shapes that raise the buffers themselves; --down keeps the file.

Transparent Hugepages

eim only; an execution agent leaves the kernel alone. Every inference deploy sets the host's transparent hugepage mode to madvise for both enabled and defrag, so hugepages and their compaction stalls are confined to memory that asks for them. The runtime value takes effect immediately, and the command makes it persistent:

  • Debian and Ubuntu with GRUB: adds transparent_hugepage=madvise to the kernel command line through a drop-in, /etc/default/grub.d/99-erebine-thp.cfg, then runs update-grub. /etc/default/grub is never edited.
  • RHEL, Rocky Linux, AlmaLinux, CentOS, and Oracle Linux: adds the same argument to every installed kernel with grubby.
  • Both: /etc/tmpfiles.d/99-erebine-thp.conf restores defrag at boot, because no kernel argument sets it.

A new kernel argument applies at the next reboot. The command says when one is needed and never reboots the host. Anything it will not change it reports with the exact commands to run, and the deploy continues: inside a container, on image-mode hosts (rpm-ostree kargs), on other distributions, or when a different transparent_hugepage= value is already configured. When the active tuned profile sets its own transparent hugepage value, the command warns and names the profile change that stops the override. Re-running changes nothing that is already in place; until the reboot, the Debian path runs update-grub again so an interrupted run is never mistaken for a finished one.

Pass --no-thp to skip the step on hosts that manage kernel arguments elsewhere.

EEM Options

The execution agent has no hardware to detect and no kernel to tune. What it needs instead is an identity in the project and the directory it reads credentials from; its tools are compiled into the agent:

FlagMeaning
--registration-nameThe agent's name, unique per project (default eem-<hostname>)
--projectProject the agent reports to (default: the project the base_url is scoped to). Refused under --update, which re-renders the recorded project
--credentials-dirCredential bundles, mounted read-only (default <data-dir>/credentials)
--data-dirData root holding credentials, state and config (default /data/erebine-eem)
--domainsDomain the agent may serve (repeatable; default every domain its tools declare)
--disable-domainsDomain the agent must not serve (repeatable)
--workspaceWorkspace the agent declares, by its name in the project (repeatable). Refused under --update and --down

The three data subdirectories are created and handed to the agent's own uid, so nothing in them ends up owned by root. On bare metal the credentials directory stays root's instead, readable by the agent's group: without a container to mount it read-only, ownership is what keeps the agent from writing it. The agent writes its signing key, its CURVE key and its enrollment record into <data-dir>/config, which is why a restart costs no join-key slot.

--workspace seeds the workspaces file the agent enforces, into that same config directory. The names are read from your project and resolved before anything on the host changes: a name the project does not have, or one that names a chat workspace, refuses the whole run rather than deploying an agent that binds nothing. A declaration binds on the name, so that check is what stops a typo from silently serving nothing.

bash
sudo erectl deploy agent eem --service --region us-east \
  --workspace prod-k8s --workspace erebine-repo

Without the flag no file is written and the agent keeps its single local workspace; the command prints the erectl workspaces declare line that writes one later. A code workspace arrives without repo_roots -- those name directories on this host, and the project has nowhere to put them -- so the deploy says which ones to fill in before the agent serves them. The file is yours after the first write: an edit takes effect without a restart, and --update leaves it exactly as it found it.

Everything else -- the region, the key source, the endpoint, the unit, the timeouts -- comes from the common options both kinds share.

Systemd and Reboots

--service installs a unit per kind -- erebine-agents.service for eim, erebine-eem.service for eem -- so the stack starts at boot on every runtime (on bare metal the same names run the agents directly). Without it: docker restarts the containers itself; podman only does so on 5.8.2 or newer with podman-restart.service enabled -- the plan output tells you exactly what your host will do.

Every unit the command installs, for both kinds and on every runtime, lifts the locked-memory and address-space limits (LimitMEMLOCK=infinity, LimitAS=infinity): vLLM and CUDA map very large virtual address ranges, and NCCL registers memory against the locked-memory limit. docker starts containers from its own daemon rather than from the unit, so every inference service in the generated compose also carries ulimits: memlock: -1 (soft and hard), which docker and podman both apply.

The unit is a oneshot that runs up -d and stays active; it does not hold the containers attached. Container output reaches the journal through the journald log driver both kinds' generated compose pins, so journalctl -u erebine-agents -f (or -u erebine-eem) follows it on docker as well as podman. systemctl restart erebine-agents bounces the containers; enrollment state persists, so agents relaunch without re-enrolling.

Stopping the service runs compose down, which removes the containers on purpose: left in place, the next up -d adopted the old container and kept serving the old image. Nothing in the data directory is touched -- keys, enrollment records and model state all survive -- so the next start creates fresh containers that come back as the same agents. A reboot brings everything back, including the MIG layout when the persistence unit is installed. After an NVIDIA driver update on a MIG host, run systemctl restart erebine-mig-setup.

Verification

A deploy is not finished when the containers are up. Both kinds hold the run open until the router agrees the agents are serving, and fail the deploy rather than reporting success over a stack that enrolled and then went quiet.

erectl deploy agent eim runs four gates: the join key's enrollment counter rises by the number of agents that needed enrolling; every container is running on two samples; the accelerator is visible inside each service (pinned multi-GPU and MIG shapes); and the router reports the new agents connected and heartbeating.

erectl deploy agent eem runs three:

  1. Enrollment. The join key's enrollment counter rises by one. Skipped when the agent already holds an enrollment record, because it restores its identity and redeems no slot.
  2. Liveness. The container is running, sampled twice so a crash loop cannot pass between samples.
  3. Router health and capability manifest. The router reports this agent -- matched by the worker id the agent persisted, never a sibling on another host -- connected, heartbeating, and having published a capability manifest. An agent that enrolls but publishes no manifest advertises no tools, so this gate treats it as a failed deploy and says so.

A verified deploy prints the command that revokes the join key. Run it: the enrollment is complete, and for both kinds the identity on disk is what brings the agents back.

Updates

--update re-renders the deployment that is recorded on the host with the erectl you are running, refreshes the image, and restarts only when something moved. It changes nothing about the shape, so every shape flag is refused by name -- including --project on eem, whose project comes from the record.

bash
sudo erectl deploy agent eim --update
sudo erectl deploy agent eem --update

An update over an unchanged stack re-renders nothing and restarts nothing, so it is safe to run on a schedule. --restart forces the bounce anyway; --no-restart applies the changes and leaves the stack running as it is.

An update sizes shared memory from the host's /dev/shm as it is now unless the deploy passed --shm-size, so a host whose /dev/shm changed gets the new size and the stack restarts to take it.

Teardown

bash
# Inference stack:
sudo erectl deploy agent eim --down              # stop and remove the stack
sudo erectl deploy agent eim --down --revoke-key # also revoke the join key

# Execution agent:
sudo erectl deploy agent eem --down
sudo erectl deploy agent eem --down --revoke-key

Each kind tears down only its own stack; the other keeps running. Teardown keeps what is expensive or shared: model caches and agent identities under the data directory, the env file, the sysctl entries, the transparent hugepage settings, and the MIG layout (with the unwind recipe printed). For the execution agent it also keeps the enrollment record beside its keys. Re-deploying after a teardown picks the identities back up without re-enrolling.

Common Options

These belong to the deploy agent group and mean the same thing for both kinds:

FlagMeaning
--regionJoin-key region (required when minting)
--join-key, --join-key-fileBring your own key (- for stdin; prefer the file form, a value is visible in ps)
--join-key-nameDisplay name for the minted key
--enroll-endpointAgent HTTP enrollment URL (default: the key's own claim, else the base_url origin)
--allow-insecure-enrollPermit a plain-HTTP, non-loopback enrollment endpoint
--enroll-timeoutSeconds to wait for the enrollment gate
--serviceInstall and enable the kind's systemd unit
--runtimepodman | docker | baremetal (default: podman, then docker; baremetal needs --service)
--install-methodBare metal: brew | deb | rpm | binary (default: reuse, then the host's package manager, then brew, then the release binary)
--deploy-dirGenerated-file home (default /etc/erebine/deploy for eim, /etc/erebine/deploy-eem for eem). This and every directory flag take an absolute path without spaces, quotes, backslashes, % or $: they are written into unit files as they are
--registry, --image-versionImage source override and tag pin; on bare metal --image-version pins the agent release and --registry is refused
--max-concurrent, --log-levelAgent tunables
--no-startStop after artifacts, units and pull; the printed recipe finishes the deploy
--down, --revoke-keyTeardown, and revoking the recorded key with it
--update, --restart, --no-restartRe-render the recorded deployment; force or suppress the restart
--lock-wait, --yes, --dry-runLocking, confirmation, preview

EIM Options

On top of the common options, the inference kind takes:

FlagMeaning
--countAgents to deploy (default 1; 1 to 64). Above one, every agent gets EREBINE_AGENT_HOST_AGENT_COUNT and sizes its default KV offload from that share of host RAM, so the agents do not each claim the whole host's RAM; an EREBINE_AGENT_KV_OFFLOAD_SIZE_GB set in agents.local.env is used as set
--hardware-type / --auto-hardware-detectExactly one is required
--gpu-strategypinned | shared | mig
--mig, --mig-profile, --mig-repartitionMIG shorthand, profile name, layout replacement
--tierTier slug (repeatable; default self_hosted -- user-owned agents; list slugs with erectl tiers)
--router-addrRouter ZeroMQ address (repeatable; defaults to tcp://<base_url host>:5555)
--data-dirData root (default /data/erebine)
--shm-sizeShared memory (/dev/shm) per agent container. Default: the size of the host's own /dev/shm, read at deploy time and rounded down to whole GiB (256g for a 256 GiB host /dev/shm; under 1 GiB, whole MiB). /dev/shm is tmpfs, so the size caps what each container can write there and sets nothing aside. Takes a compose size such as 64g, 512m or 1.5g. The agent's own KV cache offload (EREBINE_AGENT_KV_OFFLOAD_SIZE_GB) uses process RAM, not /dev/shm. Only a --kv-offloading-size passed straight to vLLM through EREBINE_AGENT_VLLM_ARGS keeps its cache here: one file of the full size per data-parallel engine, so it needs size x DP + 1 GiB + DP x 424 MiB, and the agent clamps it at launch. The plan shows the size and where it came from, warns when agents.local.env passes such an offload that does not fit at DP=2 (with the DP=1 and DP=2 limits), and falls back to 90g with a warning when the host's /dev/shm cannot be read. A single agent, MIG aside, on docker or on a podman-compose release after 1.6.0 shares the host's IPC namespace and /dev/shm (ipc: host), so no shm_size is written there and the plan shows n/a; podman-compose 1.6.0 and earlier ignore ipc: host and get the size
--gpu-memory-utilization, --max-model-lenInference tunables
--overprovision-factorAdmit slots per safe slot, at least 1.0 (default 1.5). Agents bound to an anchor are routed by safe slots
--no-sysctl, --no-thp, --force-chownStep controls (bare metal refuses --no-sysctl and --shm-size)
--vllm-pathBare metal: the vllm executable, or the Python environment holding bin/vllm (default: vllm on PATH)

Troubleshooting

  • Join key expired: keys live one hour. The command minimizes the window by pulling first; a very slow first pull can still cross it. Re-running is safe -- every step converges and a fresh key is minted.
  • Enrollment timeout with many agents: enrollment is rate limited per host. The command paces starts above five agents automatically; the timeout scales with --count and can be raised with --enroll-timeout.
  • Agents enroll but never serve: the router-health gate catches this -- it means the ZeroMQ path to the join key's router addresses (--router-addr, or the default derived from base_url) is unreachable from the host. Check routing and firewalls toward the router's agent port, and pass --router-addr explicitly if the agent plane is not on the base_url host.
  • The join key was minted for the other kind: a key carries the role it was minted for, and a deploy of the wrong kind is refused before the host is touched. Mint the key from the matching dashboard tab, or drop --join-key-file and let the deploy mint the right one.
  • The execution agent enrolled but published no manifest: it advertises no tools, so the third gate fails the deploy. Its tools are compiled into the agent: check that --domains and --disable-domains leave some enabled, and read journalctl -u erebine-eem.
  • The execution agent will not enroll again: it redeems a join key once and keeps the enrollment beside its keys, in agent-enrollment.json under its config directory. Remove that file to enroll again from a fresh key.
  • MIG devices unresolvable after a driver update: run systemctl restart erebine-mig-setup to re-carve and refresh the CDI spec.
  • Another deploy holds the lock: the message names its pid, start time, and last completed step; pass --lock-wait to queue behind it. The lock is per host, so the two kinds queue behind each other rather than racing.
  • A recorded deployment belongs to the other kind: an explicit --deploy-dir that points at the other kind's directory is refused, naming the kind it found and the flag that fixes it.
  • No container runtime found: install podman or docker, or re-run with --runtime baremetal --service.
  • Another unit runs the agent: an enabled or running unit, such as the one a deb or rpm package ships, runs erebine-eim-agent or erebine-eem-agent and would be a second agent of the same kind. Run the systemctl disable --now command the error names, then re-run the deploy.
  • Bare metal: no vLLM found: under sudo, PATH is root's. Pass --vllm-path with the environment vLLM is installed in.
  • Bare metal: an agent instance is failed: five failed starts inside ten minutes park an instance so a bad join key is not retried forever. Read journalctl -u erebine-agents@<n>, fix the cause, then systemctl reset-failed and systemctl restart erebine-agents.