262K native context
Every open-weight model we serve exposes its full native context window, up to 262,144 tokens, so long documents and long-running agents don't get truncated.

Enterprise-grade, OpenAI-compatible inference for the open-weight model ecosystem. One API, transparent per-token pricing, no lock-in.
We are onboarding the major open-weight model families behind a single OpenAI-compatible API. New models and families ship as they release.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
Multiple sizes, quantizations, and context windows. More families onboarding soon.
+ 6 more families on the roadmap, with new models added as they ship.
The details that separate a toy endpoint from infrastructure you can route real traffic through.
Every open-weight model we serve exposes its full native context window, up to 262,144 tokens, so long documents and long-running agents don't get truncated.
Automatic prefix caching at the inference layer means repeated system prompts and conversation history are billed at a fraction of the cost and return far faster.
OpenAI-compatible by design. The Vercel AI SDK, LangChain, n8n, and other standard tooling work against a single endpoint, so the open ecosystem plugs in without rewrites.
First-class function calling, streamed tool calls, toggleable reasoning mode, structured outputs, and multimodal (text + image) input on the models that support it.
Inference is spread across a global fleet of high-performance GPUs with automatic failover. One region going down never takes your requests down with it.
We don't train on your prompts. A zero-data-retention, no-log serving path is available for workloads that require it, a property of the build and not a promise.
From a single request to a fleet of GPUs, the same open-weight model served reliably wherever you are.
A single base URL for the full model catalog. Auth, validation, and streaming are handled at the edge of the network.
Requests are routed to the best-fit healthy GPU based on model, region, and load, with warm-session affinity for cache reuse.
Models run on optimized vLLM deployments across our global fleet, tuned for tokens-per-second and time-to-first-token.
Responses stream immediately with accurate per-token usage and cached-token accounting, so you bill what you actually use.
One OpenAI-compatible endpoint, reachable through the platforms you already use. No lock-in.