Lightning fast inferenceoptimized for you
Access to DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3, and Qwen3.8 through OpenAI-compatible and Anthropic-compatible APIs.
02 / Lightning fast inference
DeepSeek V4.1
550tok/s
up to 5.8× faster
95–247tok/s
Prism Engine
$ prism chat --model deepseek-v4.1
›
Major Providers
$ chat --model deepseek-v4.1
›
03 / Serverless Model APIs
Run open source models through a single API.No infrastructure to manage.
One platform for the full AI inference stack
Model APIs, GPU infrastructure, and agent runtimes. All in one platform.
Model APIs
GPU infrastructure
Agent runtimes
Top Open Source models
DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen on one key.
Better price-performance
Up to 50% less than major cloud providers. Not because we cut corners, because we built the infrastructure.
Scale with your workload
Start small and scale seamlessly from APIs to dedicated clusters.
Serverless API
Dedicated endpoints
GPU control
Built for production reliability
Stable infrastructure with low latency, high throughput, and reliable uptime at scale.
99.99%
Uptime SLA
<50ms
P50 latency
Global
PoPs
04 / Questions
What is Prism?
Prism is an OpenAI- and Anthropic-compatible inference API for coding agents. It serves open-weight frontier models so you can run the primary agent loop without standing up your own GPU stack.
Which models can I call?
The public lineup is DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen. Currently available model ids are deepseek-v4.1-flash, deepseek-v4-flash, glm-5.3, kimi-k3, qwen-3.8, and qwen.
Which API formats are supported?
Prism supports OpenAI Chat Completions at https://api.prisminference.com/v1 and Anthropic Messages at https://api.prisminference.com. Keep your existing client and change the base URL, API key, and model id.
How do I get an API key?
Create a Prism account, then create an API key from your account settings. Store it as PRISM_API_KEY and send it with your requests.
How is inference priced?
Usage is billed per token. There is no per-seat fee. See Pricing for serverless, elastic, dedicated, and batch options.
Which coding agents work with Prism?
Any client or gateway that can send OpenAI Chat Completions or Anthropic Messages, including Cursor, Claude Code, and configurable custom loops. Clients that require the OpenAI Responses API need a compatibility gateway.
Does Prism retain prompts or train on them?
No. Prism applies strict zero data retention to API inputs and outputs and never uses them for training. See the Privacy Policy for details.
Can I get a private deployment?
Yes. Book a call for dedicated capacity in our cloud or yours.
05 / Get a key