EIM Advanced
Knobs for the operators who need them. Quantization, speculative decoding, flow control, lease windows, vLLM startup timeouts, and per-tenant cache salts. Touch these when a symptom forces your hand.
Quantization
The EIM node supports both automatic and manual quantization to fit large models into limited GPU VRAM. Auto-quantization inspects the model size and available VRAM at startup and selects the best method automatically.
Quantization Override
Quantization is determined by the database value set when the endpoint is created. The agent CLI can override this with an explicit method:
# Force a specific quantization method
EREBINE_AGENT_VLLM_QUANTIZATION=bitsandbytes
Quantization Environment Variables
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_VLLM_QUANTIZATION |
- | Force a specific quantization method (overrides database value). Supported values: bitsandbytes, bitsandbytes-fp4, fp8, awq, gptq, and more... |
Pre-Quantized Models: If the model is already quantized (AWQ, GPTQ format on HuggingFace), the EIM node detects this and passes the appropriate flag to vLLM. No additional configuration is needed.
Speculative Decoding
Speculative decoding drafts several tokens ahead and has the model verify them in one forward pass, so a response streams faster without changing what the model samples. It is a setting on each model in the catalog, not on the node: set the method, the token count and a draft model on the model (see Model Management: Speculative Decoding) and every EIM node that serves the model applies it when it starts the engine. The node syncs the draft model like any model, verifies it, and passes its local path to the engine. Nothing is downloaded, placed or configured on the node by hand.
Which setting a launch uses
- Node override. When
EREBINE_AGENT_SPECULATIVE_ENABLEDis set to a boolean, theEREBINE_AGENT_SPECULATIVE_*variables below are the whole setting for every model this node serves, and the models' catalog settings are ignored. The block is taken as a unit, never merged with a catalog setting. - Endpoint. An endpoint's speculative decoding setting of
offnever speculates;onspeculates with the model's method, or with the model's own multi-token-prediction layers when the model names no method. - Model. On
auto, a model whose setting is enabled speculates. - Workload profile. A model not enabled speculates only on an endpoint whose workload profile is
latency.
When to use the override: a node with no catalog to read -- the macOS app's built-in agent, or a node that serves EREBINE_AGENT_INITIAL_MODEL_PATH without a router -- and as a per-node switch: EREBINE_AGENT_SPECULATIVE_ENABLED=0 keeps every engine on the node from speculating. The other EREBINE_AGENT_SPECULATIVE_* variables do nothing without EREBINE_AGENT_SPECULATIVE_ENABLED; the node warns about them at startup.
Override Environment Variables
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_SPECULATIVE_ENABLED |
unset | Unset: each model's catalog setting applies. 1, true or yes: speculate on every launch with the variables below. 0, false or no: never speculate on this node. |
EREBINE_AGENT_SPECULATIVE_METHOD |
auto | The method, from the table below. Unset: native MTP when the model's checkpoint ships MTP layers. |
EREBINE_AGENT_SPECULATIVE_TOKENS |
method default | Tokens drafted per step. Higher values pay off only while most drafted tokens are accepted. |
EREBINE_AGENT_SPECULATIVE_NGRAM_FALLBACK |
disabled | Set to 1 or true to speculate with ngram when the model's checkpoint ships no MTP layers, instead of not speculating. Applies when the method is unset or native MTP. |
EREBINE_AGENT_SPECULATIVE_DRAFT_MODEL_PATH |
- | Local directory of the draft model, for the methods that run one. Only the override reads it: a draft set in the catalog is synced and located by the node itself. |
Override Methods
| Method | Draft Model | Description |
|---|---|---|
mtp |
not used | Native Multi-Token Prediction from the model's own MTP layers. Default tokens per step: 1. |
deepseek_mtp, qwen3_next_mtp, mimo_mtp |
not used | Accepted spellings of mtp; the engine runs every native MTP method as mtp. Default tokens per step: 1, and 2 for qwen3_next_mtp. |
ngram |
not used | Prompt lookup: continues an n-gram already in the context. Default tokens per step: 5. |
eagle, eagle3 |
required | An EAGLE or EAGLE-3 head trained for the model. Default tokens per step: 3. |
dflash |
required | A DFlash drafter trained for the model. Default tokens per step: 15. |
draft_model |
required | A smaller model that shares the model's tokenizer. Default tokens per step: 5. |
medusa, mlp_speculator |
required | Medusa or MLP speculator heads trained for the model. Default tokens per step: 3. |
suffix |
not used | Suffix decoding. Not available: the node images do not install the library it needs, so the node launches without speculative decoding and logs why. |
Example: DFlash on a Node Without a Catalog
EREBINE_AGENT_SPECULATIVE_ENABLED=1
EREBINE_AGENT_SPECULATIVE_METHOD=dflash
EREBINE_AGENT_SPECULATIVE_TOKENS=7
EREBINE_AGENT_SPECULATIVE_DRAFT_MODEL_PATH=/srv/models/my-model-dflash
Example: Native MTP With an N-Gram Fallback
EREBINE_AGENT_SPECULATIVE_ENABLED=1
# Method unset: native MTP when the checkpoint ships MTP layers,
# n-gram when it does not.
EREBINE_AGENT_SPECULATIVE_NGRAM_FALLBACK=1
While an engine speculates: the model's HF config overrides are not passed to the engine, because the engine does not apply them to the draft (vllm-project/vllm#37435); and a thinking budget is accepted only from a request that is greedy, sends top_p 1.0, or carries a top_k (vllm-project/vllm#58231). A model's catalog setting can supply a default top_k for that; the override cannot. See Sampling defaults.
MoE Kernel Tuning
For Mixture-of-Experts (MoE) models, the EIM node can automatically generate and apply optimized kernel tuning configurations. This improves expert dispatch performance on your specific GPU hardware.
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_MOE_CONFIG_ENABLED |
enabled | Enable automatic MoE kernel tuning config generation. Set to 0 or false to disable. |
EREBINE_AGENT_MOE_CONFIG_PATH |
auto | Custom path for MoE tuned config files. When unset, the EIM node stores configs alongside the model cache. |
Non-MoE Models: These settings have no effect on dense (non-MoE) models. The EIM node detects whether the loaded model uses MoE architecture and only generates configs when applicable.
Auto-Configuration
The EIM node includes a dynamic auto-configuration system that inspects the model architecture, GPU hardware, and available memory at startup to select optimal vLLM parameters. This covers tensor parallelism, quantization, context length, and CUDA graph settings.
Auto-Configuration Environment Variables
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_AUTO_CONFIG |
enabled | Master toggle for dynamic auto-configuration. Set to 0 or false to disable all auto-tuning and use only explicit settings. |
EREBINE_AGENT_AUTO_CONFIGURE_GPU |
enabled | Auto-configure the parallel layout (tensor, data, and pipeline parallelism) from the detected GPU count and NVLink topology. Set to 0 or false to use the explicit EREBINE_AGENT_TENSOR_PARALLEL_SIZE value. |
EREBINE_AGENT_DATA_PARALLEL_SIZE |
auto | Explicit data-parallel replica count. When unset, the agent runs one replica per NVLink island whenever the model fits inside an island. On an MoE model the expert-parallel derivation takes precedence when it produces its own replica count, and the decision log records the value it superseded. |
EREBINE_AGENT_PIPELINE_PARALLEL_SIZE |
auto | Explicit pipeline-parallel stage count. When unset, the agent spreads a model that does not fit one NVLink island across islands as pipeline stages. Applies to dense and MoE models alike. |
EREBINE_AGENT_VLLM_ARGS |
empty | Extra vLLM arguments appended to the launch, space separated. The launch caps a single request's prefill per scheduler step at a quarter of the batched-token budget, never below 2048 tokens and omitted when the budget itself is 2048 tokens or less; pass --long-prefill-token-threshold here to override it. A --tensor-parallel-size, --data-parallel-size or --pipeline-parallel-size passed here reaches the layout planner as well, so the agent plans the layout the engine runs: --data-parallel-size 2 alone on four GPUs yields tensor parallelism 2. A vision or video endpoint at tensor parallelism above one launches with --mm-encoder-tp-mode data, which keeps the full image encoder on every rank and splits the batched images across the ranks instead of reducing after every encoder layer; vLLM falls back to weights on its own for a model whose encoder does not support it. Pass the flag here to override. |
EREBINE_AGENT_AUTO_CUDA_MITIGATION |
enabled | Auto-apply CUDA graph mitigations for known GPU issues (A30, A40, L40 with TP>1). Set to 0 or false to disable. |
Disabling Auto-Configuration
To take full manual control of vLLM parameters, disable all auto-configuration:
# Disable all auto-configuration
EREBINE_AGENT_AUTO_CONFIG=0
EREBINE_AGENT_AUTO_CONFIGURE_GPU=0
EREBINE_AGENT_AUTO_CUDA_MITIGATION=0
# Then set explicit values (the default for GPU_MEMORY_UTILIZATION is 0.95;
# the value below is an example override for memory-constrained deployments).
# Two NVLink pairs on one host: one replica per pair.
EREBINE_AGENT_TENSOR_PARALLEL_SIZE=2
EREBINE_AGENT_DATA_PARALLEL_SIZE=2
EREBINE_AGENT_GPU_MEMORY_UTILIZATION=0.85
EREBINE_AGENT_MAX_MODEL_LEN=8192
Boolean parsing is case-sensitive: The EREBINE_AGENT_AUTO_CONFIG, EREBINE_AGENT_AUTO_CONFIGURE_GPU, EREBINE_AGENT_AUTO_CUDA_MITIGATION, EREBINE_AGENT_VLLM_DISABLE_CUDA_GRAPHS, and EREBINE_AGENT_VLLM_DISABLE_CUSTOM_ALL_REDUCE env vars compare against the literal strings 0, 1, false, and true. Values such as True, TRUE, FALSE, off, or yes are silently ignored. Use the exact lower-case form.
CUDA Graph Workarounds
Some GPU models (A30, A40, L40) experience CUDA graph capture failures under specific tensor parallelism configurations. The EIM node detects these cases and applies mitigations automatically when EREBINE_AGENT_AUTO_CUDA_MITIGATION is enabled.
For manual control of CUDA graph behavior:
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_VLLM_DISABLE_CUDA_GRAPHS |
disabled | Force disable CUDA graphs (enforce eager execution). Set to 1 or true if experiencing CUDA graph capture failures. |
EREBINE_AGENT_VLLM_DISABLE_CUSTOM_ALL_REDUCE |
disabled | Force disable custom all-reduce optimization. Set to 1 or true for multi-GPU P2P issues on A30/A40 GPUs. |
Flow Control
The EIM node implements credit-based flow control for streaming inference to prevent buffer overflow on the router side. Flow control ensures the EIM node does not send chunks faster than the router can forward them to clients.
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_STREAMING_SOCKET_ENABLED |
true | Enable the dedicated streaming DEALER socket. When enabled, inference chunks use a separate socket to avoid head-of-line blocking on the control channel. |
EREBINE_AGENT_FLOW_CONTROL_ENABLED |
true | Enable credit-based flow control for streaming inference. Set to 0 or false to disable backpressure. |
EREBINE_AGENT_FLOW_CONTROL_WINDOW_BYTES |
65536 | Initial credit window size in bytes per stream. The EIM node can send up to this many bytes before waiting for the router to replenish credits. |
EREBINE_AGENT_FLOW_CONTROL_REPLENISH_THRESHOLD |
0.5 | Fraction of the window that triggers a credit replenishment from the router (0.0-1.0). |
EREBINE_AGENT_FLOW_CONTROL_PAUSE_THRESHOLD_BYTES |
262144 | Total unacknowledged bytes across all streams before the EIM node pauses sending. Acts as a global safety valve. |
EREBINE_AGENT_FLOW_CONTROL_TIMEOUT_SECONDS |
30 | Seconds to wait for credit recovery before considering the stream stalled. |
When to adjust: Touch these only if the EIM node logs show streaming stalls or excessive backpressure pauses.
Lease Configuration
The EIM node maintains a lease with the router mesh via periodic heartbeats. If the router does not receive a renewal within the lease duration window, it marks the EIM node as expired and stops routing requests to it.
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_LEASE_RENEWAL_INTERVAL_MS |
10000 | How often (in milliseconds) the EIM node sends a lease renewal heartbeat to the router. Default: 10 seconds. |
EREBINE_AGENT_LEASE_DURATION_MS |
30000 | Requested lease duration in milliseconds. If the router does not receive a renewal within this window, the EIM node is marked expired. Default: 30 seconds. |
Keep the ratio safe: The lease duration should be at least 2-3x the renewal interval. A tight margin increases the risk of false lease expirations during transient network issues.
Model Pull Retry
When the EIM node downloads a model from the router or a remote registry, it uses exponential backoff retries on failure. These settings control retry behavior.
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_MODEL_PULL_MAX_ATTEMPTS |
5 | Maximum number of retry attempts for a failed model pull before giving up. |
EREBINE_AGENT_MODEL_PULL_BASE_DELAY_MS |
1000 | Base delay in milliseconds between retries. Each subsequent retry doubles the delay (exponential backoff). |
EREBINE_AGENT_MODEL_PULL_MAX_DELAY_MS |
15000 | Maximum delay in milliseconds between retries. The backoff delay is capped at this value. |
Retry Sequence: With defaults, retries occur at approximately 1s, 2s, 4s, 8s, 15s (capped). Increase EREBINE_AGENT_MODEL_PULL_MAX_ATTEMPTS for unreliable networks.
Enrollment Options
Additional options for the enrollment and startup process.
Insecure Enrollment
By default, enrollment requires HTTPS endpoints. For development or trusted network environments, you can allow non-HTTPS enrollment:
# Via CLI flag
EREBINE_AGENT_JOIN_KEY=ejk_abc123... erebine-eim-agent enroll --insecure
# Via environment variable
EREBINE_AGENT_ALLOW_INSECURE=1 EREBINE_AGENT_JOIN_KEY=ejk_abc123... erebine-eim-agent enroll
Security Risk: Insecure enrollment transmits your join key and agent credentials over an unencrypted connection. Only use this in isolated development environments. Never use --insecure in production.
Initial Model Path
Preload a local model on startup instead of waiting for the router to stream one. This is useful when you have already downloaded a model to disk:
# Load a model from a local path on startup
EREBINE_AGENT_INITIAL_MODEL_PATH=/data/models/my-model
When set, the EIM node starts vLLM with this model immediately after registration. The model must already exist at the specified path in a format vLLM can load (safetensors or equivalent).
Tenant Cache Isolation
In multi-tenant deployments where a single EIM node serves requests from different tenants, the node gives every tenant its own cache salt, derived from a server secret, so one tenant's requests never reuse another tenant's cached prefixes in the KV cache. When this env var is unset, the node derives the secret from its enrollment: from the project for a dedicated node, from the node's own identity for a shared one. With no enrollment identity it uses a random secret that changes on every restart. Tenants stay separated in every case, but a derived secret comes from identifiers rather than from anything only you hold, so its salts can be worked out by anyone who knows them. Set your own in production:
# Server secret for tenant-isolated cache keys
EREBINE_AGENT_VLLM_SALT_SECRET=your-secret-string-here
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_VLLM_SALT_SECRET |
derived from the enrollment | Server secret for generating per-tenant cache isolation salts. When unset, the EIM node derives it from its enrollment (the project for a dedicated node, the node itself for a shared one), or uses a random secret per process when it has no enrollment identity. Every tenant gets its own salt either way; a derived secret's salts can be worked out from identifiers, and a random one's change on every restart. Production deployments must set this to a unique, high-entropy secret. |
EREBINE_AGENT_ALLOW_INSECURE |
disabled | Set to 1 or true to allow non-HTTPS enrollment. Development only. |
EREBINE_AGENT_INITIAL_MODEL_PATH |
- | Filesystem path to a local model to preload on startup. EIM node starts vLLM with this model immediately after registration. |
vLLM Runtime Tuning
Knobs that control vLLM startup behavior, KV cache backend selection, and prefix caching policy. These exist for slow-disk, large-model, and multi-tenant deployments.
Startup Timeouts
Large models on slow storage can exceed the default vLLM startup window. The EIM node exposes two timeout knobs plus a degraded-state ceiling. Defaults track the value vLLM itself respects at the time the EIM node was built; run erebine-eim-agent --help to print the current effective ceiling:
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_VLLM_STARTUP_TIMEOUT_SECONDS |
vLLM-defined | Hard ceiling (in seconds) on how long the EIM node waits for vLLM to finish loading a model before marking the start attempt as failed. Increase for large models or slow disks. |
EREBINE_AGENT_VLLM_STARTUP_INACTIVITY_GRACE_SECONDS |
vLLM-defined | Inactivity grace window during vLLM startup. If vLLM produces no progress output for this long, the EIM node considers the start attempt stuck. |
EREBINE_AGENT_DEGRADED_TO_FAILED_TIMEOUT_SECONDS |
vLLM-defined | How long an EIM node may remain in the degraded state before being promoted to failed. Operators on flaky networks may want a longer window. |
KV Cache Backend
The EIM node supports two KV cache backends, selected with EREBINE_AGENT_KV_CACHE_BACKEND (or --kv-cache-backend): native, lmcache, or none. The default native backend uses vLLM's built-in paged attention with CPU offload (see KV Cache Offload Issues). The lmcache backend runs the LMCache server beside vLLM and adds a disk tier, so a prefix survives an engine restart and, with a remote adapter, can be shared across nodes; see LMCache. LMCache is applied on NVIDIA CUDA and AMD ROCm nodes only; on other accelerators the setting is ignored and the node runs without a KV cache backend.
Prefix Caching
Always enabled: Prefix caching is enabled unconditionally on every EIM node. There is no env-var toggle to turn it off; the EIM node always emits --enable-prefix-caching when starting vLLM. This is intentional and matches the behavior validated for hybrid SSM and dense models alike.
Native MTP Is Opt-In
No family speculates by default. Native Multi-Token Prediction runs only when something asks for speculative decoding: the model's catalog setting, the endpoint, the latency workload profile, or the node override (see Speculative Decoding). To run a model's MTP layers, choose method mtp on the model.
ROCm Short-Sequence Workaround (Qwen3.5 GDN)
On AMD ROCm GPUs, Qwen3.5 models that use GDN (Generalized Delta Networks) layers crash when the sequence length is below 64 tokens. The EIM node's bundled erebine-vllm fork applies a zero-pad workaround to short sequences so these prompts no longer hit the crash path. No operator configuration is required; the fix is active whenever the EIM node detects an affected model on an AMD ROCm device.
LMCache
With EREBINE_AGENT_KV_CACHE_BACKEND=lmcache the EIM node starts an lmcache server process before vLLM and connects the engine to it. The server keeps two tiers. L1 lives in CPU memory and is sized by EREBINE_AGENT_KV_OFFLOAD_SIZE_GB, the same budget the native backend uses for offload. L2 lives on disk under the cache directory (EREBINE_AGENT_LMCACHE_DISK_PATH, default ~/.cache/erebine/lmcache, which is /var/lib/inference/.cache/erebine/lmcache in the container images). Everything below is derived by the node at launch; the environment variables exist to override it.
LMCache Disk Tier
The disk tier uses LMCache's native filesystem adapter (fs_native): a pool of I/O worker threads, byte-accurate usage tracking, and LRU eviction that the server performs itself once the tier reaches 80% of its capacity, reclaiming 20% per cycle. The node sizes the tier from the host every time it launches:
- Capacity is derived from the volume that holds the cache directory. It never exceeds 75% of the volume. It leaves a reserve of 10% of the volume (at most 256 GiB) for whatever shares the disk, normally the model store. It is bounded above by eight times the L1 size, or 12.5% of the volume if that is larger, because the disk tier only holds what spills out of L1. When the volume is nearly full the tier keeps 2% of the volume, or whatever is actually free if that is less. Objects the cache already holds count as reclaimable, so a restart with a populated cache gets the same capacity a fresh start would. Below one GiB the disk tier is not configured and the server runs L1-only.
- Workers are half the host's physical cores, rounded, between 4 and 32.
- Direct I/O is off by default. The node turns
use_odirecton only when the filesystem's block size divides the 4096-byte buffer alignment and a test write withO_DIRECTunder the cache directory succeeds. Filesystems that refuse direct I/O keep buffered writes. - Launch housekeeping runs before the server starts: objects written under a different model, tensor-parallel degree, KV cache dtype, block size or chunk size cannot be read back and are removed, and the remaining objects are trimmed oldest-first to the capacity. A cache written by an earlier version of the node is kept as long as the launch identity is unchanged.
The resolved adapter is logged once at launch together with the disk path, worker count, direct I/O decision and capacity, so the numbers a node chose are always in its log.
LMCache Environment Variables
| Variable | Default | Description |
|---|---|---|
EREBINE_AGENT_KV_CACHE_BACKEND |
native |
native, lmcache, or none. Also --kv-cache-backend. |
EREBINE_AGENT_KV_OFFLOAD_SIZE_GB |
25% of RAM, 4 to 128 | The L1 (CPU memory) tier in GiB, total across tensor-parallel ranks. With both this and the disk budget at zero the backend is refused. |
EREBINE_AGENT_LMCACHE_DISK_PATH |
~/.cache/erebine/lmcache |
Cache directory for the disk tier. Created at mode 0700. Put it on its own disk when you can; see deployment. |
EREBINE_AGENT_LMCACHE_DISK_SIZE_GB |
derived | Disk tier capacity in GiB. Replaces the derived value; still capped at 75% of the volume. 0 disables the disk tier. Also --lmcache-disk-size-gb. |
EREBINE_AGENT_LMCACHE_CHUNK_SIZE |
half of max batched tokens | KV chunk size in tokens. Hybrid-attention models are given a chunk size that is a multiple of the engine's block size automatically; an explicit value is raised to that floor when it is below it. |
EREBINE_AGENT_LMCACHE_MP_PORT |
6555 |
Port the engine uses to reach the LMCache server on the node. Bound to localhost. |
EREBINE_AGENT_LMCACHE_HTTP_PORT |
8090 |
The server's HTTP frontend, bound to 127.0.0.1. Its /status endpoint is what the node polls for liveness and what you read when inspecting. |
EREBINE_AGENT_LMCACHE_L2_ADAPTER |
built-in fs_native |
Replaces the whole --l2-adapter JSON, for a remote or shared tier (Redis, S3, and the other LMCache adapters). A JSON object is one adapter; a JSON array is a cascade, one adapter per element in that order. Nothing is merged in and the node does not validate it; the disk sizing above does not apply. An empty array, [], configures no L2 tier at all and the server runs L1-only; EREBINE_AGENT_LMCACHE_DISK_SIZE_GB=0 does the same with the built-in adapter. See custom tiers. |
EREBINE_AGENT_LMCACHE_EXTRA_ARGS |
- | Whitespace-separated flags appended to the end of the lmcache server command line. A flag that repeats one the node emits wins, because the server keeps the last occurrence. No quoting: a value cannot contain a space. The node already sizes --max-gpu-workers to the engine's GPU rank count, tensor times pipeline times data parallel, so that flag needs no passthrough. |
EREBINE_AGENT_LMCACHE_EXTRA_CONFIG |
- | JSON object merged into the engine-side connector configuration. |
EREBINE_AGENT_LMCACHE_LOG_LEVEL |
agent log level | Log level of the LMCache server. At debug the server writes one eviction line per second, so pin it lower when debugging the node itself. |
EREBINE_AGENT_LMCACHE_LOOKUP_HASH_LOG_DIR |
off | A truthy value writes lookup-hash records under <disk path>/lookup_hashes; an absolute path writes them there. The records fingerprint prompt content, so the directory is created at mode 0700. Rotation is tuned with EREBINE_AGENT_LMCACHE_LOOKUP_HASH_LOG_ROTATION_INTERVAL (seconds, 21600), EREBINE_AGENT_LMCACHE_LOOKUP_HASH_LOG_ROTATION_MAX_SIZE (bytes, 104857600) and EREBINE_AGENT_LMCACHE_LOOKUP_HASH_LOG_MAX_FILES (10). |
EREBINE_AGENT_LMCACHE_PATH |
/usr/local/bin/erebine-lmcache in the images |
Path of the executable that starts the cache server. The images set the launcher. Unset, the node finds lmcache in the active Python environment, then beside the vLLM binary, then on PATH. |
Custom Tiers and the Launcher
The container images start the cache server through erebine-lmcache, a launcher that runs the stock lmcache command line unchanged and adds two adapter keys that LMCache 0.5.4 does not have on its own:
max_capacity_gbonfs. The plain filesystem adapter tracks what it stores and can delete, but declares no capacity, so the server never evicts from it and the tier only grows. With a capacity set, the server evicts from the tier live, LRU, at the adapter'sevictionwatermark, and objects already on disk are counted at startup so a restart resumes from the true usage.s3_bucketons3. The S3 adapter addresses the bucket as a hostname. Object stores that serve S3 without wildcard DNS, a wildcard certificate or a storage domain, OpenStack Swift among them, cannot answer that. With a bucket named, every request goes to/<bucket>/<key>on the plain endpoint host instead.
A JSON array in EREBINE_AGENT_LMCACHE_L2_ADAPTER configures a cascade. Each element becomes one --l2-adapter in array order. The server stores every chunk to every tier and reads a chunk from the first tier, by position, that holds it. Each tier evicts on its own eviction settings. Nothing moves a chunk from a later tier back to an earlier one, and a chunk stays pinned in L1 until the slowest tier has stored it, so a remote second tier should be measured for upload rate from the node before it is relied on.
EREBINE_AGENT_LMCACHE_DISK_PATH=/var/lib/inference/.cache/erebine/lmcache
EREBINE_AGENT_LMCACHE_DISK_SIZE_GB=150
EREBINE_AGENT_LMCACHE_L2_ADAPTER=[{"type":"fs","base_path":"/var/lib/inference/.cache/erebine/lmcache","max_capacity_gb":150,"eviction":{"eviction_policy":"LRU","trigger_watermark":0.8,"eviction_ratio":0.2}},{"shared":true,"type":"s3","s3_endpoint":"swift.example.net","s3_bucket":"kv-cache","s3_region":"REGION","s3_prefer_http2":false,"aws_access_key_id":"...","aws_secret_access_key":"...","max_capacity_gb":2000,"eviction":{"eviction_policy":"LRU","trigger_watermark":0.8,"eviction_ratio":0.2}}]
Keep EREBINE_AGENT_LMCACHE_DISK_PATH equal to the fs tier's base_path and EREBINE_AGENT_LMCACHE_DISK_SIZE_GB equal to its capacity: the launch housekeeping above trims that directory to that size, and the two ceilings should agree. Write the variable on one line without surrounding quotes. Set "shared": true only on a tier that several nodes write, such as one bucket; a local disk is not shared. The other LMCache adapters, raw_block for a whole block device among them, take their upstream keys unchanged.
LMCache Deployment
The disk tier lives wherever the cache directory does. In the container images that is /var/lib/inference/.cache/erebine/lmcache, which is on the container's writable layer unless you mount it. Mount it, and prefer a disk of its own: the capacity rule keeps a reserve for a shared volume, but a dedicated disk means the cache and the model store never compete, and the tier can use the whole 75%.
volumes:
- /data/erebine/models:/var/lib/inference/.cache/erebine/models
- /data/erebine/vllm-cache:/var/lib/inference/.cache/erebine/vllm
- /data/erebine/config:/var/lib/inference/.config/erebine
# LMCache disk tier: a dedicated disk mounted at /data/erebine/lmcache
- /data/erebine/lmcache:/var/lib/inference/.cache/erebine/lmcache
environment:
EREBINE_AGENT_KV_CACHE_BACKEND: "lmcache"
# Optional. Leave unset to size the disk tier from the volume.
# EREBINE_AGENT_LMCACHE_DISK_SIZE_GB: "400"
Nothing else changes: the server is started, watched and stopped by the node, and it listens on localhost only. The cache survives container restarts and node upgrades. It is emptied automatically when the served model, its tensor-parallel degree, KV cache dtype, block size or chunk size changes, because objects written under the old identity cannot be read back.
Inspecting LMCache
On the node, the server's status endpoint reports the tiers and the adapter it is running:
curl -s http://127.0.0.1:8090/status | jq '.store_controller, .worker_liveness'
The built-in disk adapter appears as FSNativeL2Adapter with its base_path, num_workers and use_odirect; a custom tier appears under its own type, FSL2Adapter or S3L2Adapter, and a cascade lists its tiers in order, with each tier's capacity and usage under l2_eviction_controller.adapters. worker_liveness.tracked_instances lists the engine workers attached to the cache; the node watches the same field and restarts the engine when a server that still answers HTTP has silently lost its workers. The launch log line for the cache server carries the resolved adapter and capacity.
| Symptom | Solution |
|---|---|
| The log says the disk tier is not configured | The volume holding the cache directory is too small or too full to give the tier one GiB within the rules above, or could not be measured. Mount a disk for the cache directory, or set EREBINE_AGENT_LMCACHE_DISK_SIZE_GB explicitly. |
| The capacity is smaller than expected | The tier is bounded by eight times the L1 size. Raise EREBINE_AGENT_KV_OFFLOAD_SIZE_GB, or set EREBINE_AGENT_LMCACHE_DISK_SIZE_GB to the size you want, up to 75% of the volume. |
use_odirect is false |
The filesystem refused a direct write, or its block size does not divide 4096. Buffered I/O is correct there; nothing needs fixing unless you move the directory to a filesystem that supports direct I/O. |
| The cache is empty after a restart | The launch identity changed (model, tensor-parallel degree, KV cache dtype, block size or chunk size) and the orphaned objects were removed on purpose. A restart with the same identity keeps the cache. |
Single-Node Queuing
When a project has exactly one compatible EIM node and that node is at capacity, the router automatically queues incoming requests instead of returning an immediate 503 error. This improves the experience for small deployments running a single EIM node.
How It Works
- A request arrives and the router finds exactly one compatible EIM node.
- That node is at its concurrent request limit.
- Instead of failing, the router parks the request in a waiting queue.
- When the node completes an in-flight request and frees a slot, the queued request is dispatched.
- If the queue timeout expires before capacity becomes available, the request fails with a 503.
Queue Limits
| Parameter | Default | Description |
|---|---|---|
| Max queued per node | 10 | Maximum number of requests that can wait for a single node to free capacity. Compile-time constant. |
| Queue timeout | 30 | Maximum seconds a request waits in the queue before receiving a 503 response. The deployment sets it on the router. |
Router-Side Setting: The queue timeout is set on the router, not on the EIM node; no node setting changes the wait window.
This behavior is transparent to clients. During queuing, the router holds the HTTP connection open. If the request is dispatched successfully, the client receives a normal response. The X-Request-ID response header can be used to trace queued requests in logs.
AMD CPU Deployment (ZenDNN)
Run inference on AMD EPYC CPUs without a GPU using vLLM with ZenDNN optimization. The ZenDNN CPU image is prebuilt and published for you; there is nothing to build locally.
Prebuilt image: The ZenDNN CPU image is published at ghcr.io/erebine/container-agents/eim-vllm-zendnn:latest and runs with compose/compose.agent-amd-cpu-zendnn.yaml from the erebine/container-agents repository.
Hardware Requirements
| Component | Minimum | Recommended |
|---|---|---|
| CPU | AMD EPYC with AVX-512 | AMD EPYC 9454 (Genoa) or newer |
| System RAM | 64GB | 96GB+ (scales with model size) |
| Disk Space | 100GB SSD | 500GB+ NVMe SSD |
| CPU Cores | 16 cores | 24+ cores |
Memory Requirements by Model Size
CPU inference requires significantly more system RAM than GPU VRAM:
| Model Size | System RAM Required |
|---|---|
| Sub-1B parameters | ~32GB |
| 3-4B parameters | ~64GB |
| 7-8B parameters | ~96GB |
CPU-Specific Environment Variables
| Variable | Default | Description |
|---|---|---|
VLLM_PLUGINS |
zentorch | Enable ZenDNN optimization plugin |
VLLM_CPU_KVCACHE_SPACE |
auto | KV cache size in GB (auto: totalRAM - model - overhead) |
VLLM_CPU_OMP_THREADS_BIND |
auto | CPU core binding range (auto: 0-{nproc-1}) |
VLLM_CPU_NUM_OF_RESERVED_CPU |
auto | CPUs reserved for OS operations (auto: 1) |
Auto-Tuning: The EIM node automatically computes VLLM_CPU_KVCACHE_SPACE, VLLM_CPU_OMP_THREADS_BIND, and VLLM_CPU_NUM_OF_RESERVED_CPU at startup based on system resources. No manual calculation is required. Override via EREBINE_AGENT_VLLM_ENV if the defaults are unsuitable.
Docker Compose for CPU EIM Node
The CPU compose file is published in the erebine/container-agents repository under compose/:
git clone https://github.com/erebine/container-agents.git
cd erebine-public/compose
export EREBINE_AGENT_JOIN_KEY=ejk_your_key_here
docker compose -f compose.agent-amd-cpu-zendnn.yaml up -d
Memory Tuning: If the auto-computed VLLM_CPU_KVCACHE_SPACE causes out-of-memory errors, override it to a lower value via EREBINE_AGENT_VLLM_ENV="VLLM_CPU_KVCACHE_SPACE=<GB>".
Performance Considerations
- Concurrency: CPU inference supports fewer concurrent requests than GPU. Start with
EREBINE_AGENT_MAX_CONCURRENT=5and adjust based on model size. - Data Type: Use
--dtype=bfloat16for optimal performance on AMD EPYC with AVX-512 VNNI. - Model Selection: Smaller models (1-8B parameters) work best for CPU inference. Larger models will have significantly higher latency.
- Memory Bandwidth: Inference performance is often memory-bandwidth limited. Ensure your system has adequate memory channels populated.
Troubleshooting
Symptoms operators hit at deploy time, and the specific knob or command that resolves each one.
Agent Fails to Start
| Symptom | Solution |
|---|---|
| Join key expired | Generate a new join key from the Agents dashboard |
| Connection refused | Verify network connectivity to Erebine router mesh |
| Invalid join key format | Ensure the complete key is provided without truncation |
GPU Not Detected
# Verify NVIDIA driver
nvidia-smi
# Verify Container Toolkit installation
docker run --rm --gpus all nvidia/cuda:12.1-base-ubuntu22.04 nvidia-smi
# Check Docker runtime configuration
docker info | grep -i runtime
Model Loading Fails
| Symptom | Solution |
|---|---|
| Out of disk space | Increase disk allocation or reduce cache size |
| Model not found | Verify the model name selected for the endpoint exists on HuggingFace, or that EREBINE_AGENT_INITIAL_MODEL_PATH points to a valid local model directory. |
Permission Denied Errors
# Fix host directory permissions
sudo chown -R 5152:5152 /data/erebine
# Verify permissions
ls -la /data/erebine
Out of Memory (OOM)
- Reduce
EREBINE_AGENT_GPU_MEMORY_UTILIZATIONto 0.85 or lower - Reduce
EREBINE_AGENT_MAX_CONCURRENTto limit concurrent requests - Reduce
EREBINE_AGENT_MAX_MODEL_LENfor shorter context windows - Use a smaller model or add more GPUs
KV Cache Offload Issues
The agent ships with vLLM native CPU KV cache offloading enabled by default (a sliding share of system RAM, 25% on a 16 GiB host rising to 75% at 1 TiB). See erebine-eim-agent --help under --kv-offload-size-gb for tuning. Setting the environment variable EREBINE_AGENT_KV_OFFLOAD_SIZE_GB=0 disables offload.
The size is the total for the node's engine: with data parallelism each engine gets an equal share. The node's offload lives in process memory, not in /dev/shm. A --kv-offloading-size passed straight to vLLM with --vllm-arg is different: vLLM then keeps the cache in one /dev/shm file per data-parallel engine, each of the full size, so the engine needs the size times the data-parallel size of /dev/shm. The node lowers such a size to what fits, less headroom for NCCL and vLLM's own queues, and logs KV offload clamped to fit /dev/shm; below 1 GiB it removes the flag.
| Symptom | Solution |
|---|---|
| Host memory pressure or OOM | Lower EREBINE_AGENT_KV_OFFLOAD_SIZE_GB below the 25% default, or set it to 0 to disable offload entirely. A node takes its default from its share of the host's RAM, the RAM divided by EREBINE_AGENT_HOST_AGENT_COUNT. erectl deploy sets that count when a host runs several nodes; set it yourself when you run several nodes on one host another way |
The log shows KV offload clamped to fit /dev/shm, or the engine fails with Insufficient space in /dev/shm or Bad address |
A passed-through --kv-offloading-size times the data-parallel size does not fit /dev/shm. Raise the container's shm size to the size times the data-parallel size plus the headroom_gib the log names, or lower the size |
| No TTFT improvement on repeated prefixes | Confirm offload is enabled (non-zero EREBINE_AGENT_KV_OFFLOAD_SIZE_GB) and that requests actually share prefixes |
Common Commands
| Command | Description |
|---|---|
docker-compose logs -f agent |
View agent logs |
docker-compose restart agent |
Restart the agent |
docker-compose down |
Stop all services |
nvidia-smi |
Monitor GPU utilization |
docker stats |
Monitor container resource usage |
Frequently Asked Questions
How do I get a join key?
Navigate to the Agents page in your dashboard and click "Generate Join Key". Configure the region and expiration, then copy the generated key. The full key is only shown once.
Can I run multiple models on one GPU?
The agent loads one model at a time per vLLM instance. To serve multiple models, deploy multiple agents on separate GPUs or use time-sharing (not recommended for production).
How do I update the agent?
Pull the latest image and restart: docker-compose pull && docker-compose up -d. Your model cache and configuration persist through updates.
What models are supported?
Any model compatible with vLLM, including most HuggingFace Transformers models. Check the vLLM supported models list for compatibility.
How much VRAM do I need?
The EIM node uses a component-based estimator (weights + GQA-aware KV cache + an activation budget of roughly 20% of weights + a fixed 1.5GB CUDA overhead + a 5% safety margin), not a flat per-parameter multiplier. Rough fp16 ballpark figures including this overhead: 7B around 17-18GB, 13B around 30-32GB, 70B around 150GB+ (typically split across multiple GPUs). Models with strong grouped-query attention (GQA) and FP8 KV cache quantization run lower than this; long context windows and high concurrency run higher. Pre-quantized AWQ/GPTQ/bitsandbytes models reduce the weight component but do not eliminate the CUDA overhead or KV-cache budget. The EIM node logs its own estimate at startup; rely on that when sizing hardware.
Can I use AMD GPUs?
Yes. AMD ROCm GPU support is available. You can also run inference on AMD EPYC CPUs using vLLM with ZenDNN optimization. See the AMD CPU Deployment section for details.
Is my data secure?
EIM nodes only receive requests scoped to your project, and all node-to-router connections use CURVE-encrypted ZMQ. In a fully self-hosted topology (your EIM node and your own Erebine router), inference data stays within infrastructure you operate. When connecting a self-hosted EIM node to the shared Erebine router mesh, prompts and responses transit the shared router process during dispatch; choose a topology that matches your data-handling requirements.
What happens if my EIM node goes offline?
Requests are automatically routed to other available EIM nodes. If you have fallback enabled, requests can be served by shared infrastructure. Otherwise, they queue until your node reconnects.
How does KV cache offload work?
Native CPU KV cache offload reduces Time-to-First-Token (TTFT) for repeated prompt prefixes. See the KV Cache Offload Issues section for the default size, the tuning knob, and how to disable it; the same guidance applies here.
// jump to section
Type the number, then Enter. Esc to dismiss.