Everything we removed is the feature. No contracts, no proprietary formats, no reserved-capacity hostage math. An OpenAI-compatible router over a mesh that answers to no pantheon.
Everything we removed is the feature. No contracts, no proprietary formats, no reserved-capacity hostage math. An OpenAI-compatible router over a mesh that answers to no pantheon.
$ export OPENAI_BASE_URL=https://api.erebine.ai/proj_ABC123/llama
$ curl -s $OPENAI_BASE_URL/v1/chat/completions \
-H "Authorization: Bearer ere_myproject_****" \
-d '{"model":"deepseek-r1-distill-llama-70b",
"messages":[{"role":"user","content":"why erebine?"}],
"stream":true}'
route eim-sjc3-04 composite 0.9142 decision 412us
cache prefix hit - 912 tok reused
Erebus existed before the gods. That is the whole point: infrastructure that answers to no pantheon. Your models, your hardware, any silicon.
done 1,284 tok - ledger committed
Sixteen signals scored, A-Res sampled, leased, streamed, billed. The panel is the pipeline.
A worked example: the fields are what every request returns.
Routing specRouter decision, p99. The math runs in the request path and still beats your ORM.
(1) live p99 - rolling 5,000-decision window, worst active routerCURVE auth and MessagePack framing on every hop. The CVE class is deleted, not patched.
(2) cve-2024-9053, cve-2025-32444Model-cache affinity, prefix-cache match, EWMA latency, KV pressure, and twelve more, scored per worker on every request. The decision is arithmetic, not round robin.
(3) composite score inputs - specified in docs/routerThe same router, mesh, and workspace platform that runs erebine.ai, licensed to run inside your network and on your accelerators. Full capability: nothing is removed, throttled, or held back for the private build. It is not a lesser edition of the hosted service; it is the hosted service, on your hardware.
A credit is the unit you spend, metered per token. Every plan grants a balance each cycle, and the grant, not the price, is what you spend. Cached tokens bill under the standard per-token rate and priority routing bills over it; erepress trims tokens before they are billed at all.
Metered: tokens in and out, stored model bytes per GB per month, and hourly EIM reservations on dedicated hardware. Plans differ by grant size and per-token multiplier. Compute tiers differ by silicon and isolation.
Figures are published to signed-in accounts. Start free to read them, or ask for a quote.
contact@erebine.ai