erebus, the primordial dark

Everything we removed is the feature. No contracts, no proprietary formats, no reserved-capacity hostage math. An OpenAI-compatible router over a mesh that answers to no pantheon.

Empty space
is the product.

3 regions 8 agents online p99 231us limb: live
Start free Read the routing spec credits granted every cycle

One curl

// one curl: the whole pipeline in one frame route + stream + ledger
api.erebine.ai/proj_ABC123/llama worked example
$ export OPENAI_BASE_URL=https://api.erebine.ai/proj_ABC123/llama
$ curl -s $OPENAI_BASE_URL/v1/chat/completions \
    -H "Authorization: Bearer ere_myproject_****" \
    -d '{"model":"deepseek-r1-distill-llama-70b",
         "messages":[{"role":"user","content":"why erebine?"}],
         "stream":true}'

route  eim-sjc3-04  composite 0.9142  decision 412us
cache  prefix hit - 912 tok reused

Erebus existed before the gods. That is the whole point: infrastructure that answers to no pantheon. Your models, your hardware, any silicon.

done   1,284 tok - ledger committed
this request
node eim-sjc3-04
decision 412us
wire CURVE · msgpack
prefix cache hit

Sixteen signals scored, A-Res sampled, leased, streamed, billed. The panel is the pipeline.

A worked example: the fields are what every request returns.

Routing spec

Measured claims

231us (1)

Router decision, p99. The math runs in the request path and still beats your ORM.

(1) live p99 - rolling 5,000-decision window, worst active router
0 pickle (2)

CURVE auth and MessagePack framing on every hop. The CVE class is deleted, not patched.

(2) cve-2024-9053, cve-2025-32444
16 signals (3)

Model-cache affinity, prefix-cache match, EWMA latency, KV pressure, and twelve more, scored per worker on every request. The decision is arithmetic, not round robin.

(3) composite score inputs - specified in docs/router

What ships in the router.

six entries · each links to its documentation
sovereign · answers to no pantheon

The whole stack,
in your datacenter.

The same router, mesh, and workspace platform that runs erebine.ai, licensed to run inside your network and on your accelerators. Full capability: nothing is removed, throttled, or held back for the private build. It is not a lesser edition of the hosted service; it is the hosted service, on your hardware.

frontend node / month
router node / month
agent / month
Operational license, container images, internal documentation, 24x7x365 operations, and direct access to engineering. Six-month minimum. Licensed per node on a monthly rate. Node rates on request.

Credits in, tokens out.

full pricing breakdown

A credit is the unit you spend, metered per token. Every plan grants a balance each cycle, and the grant, not the price, is what you spend. Cached tokens bill under the standard per-token rate and priority routing bills over it; erepress trims tokens before they are billed at all.

Metered: tokens in and out, stored model bytes per GB per month, and hourly EIM reservations on dedicated hardware. Plans differ by grant size and per-token multiplier. Compute tiers differ by silicon and isolation.

// plans what a cycle grants
  • Free
  • Starter
  • Hot-shot
  • Extra
  • Double Extra Deluxe
// compute tiers where a credit is spent
  • Free
  • CPU AMD Optimized
  • GPU NVIDIA Shared NVL
  • GPU NVIDIA Shared
  • GPU AMD Shared
  • Self-Hosted

Figures are published to signed-in accounts. Start free to read them, or ask for a quote.

contact@erebine.ai