Backed byY CombinatorCombinator

Lightning fast inferenceoptimized for you

Access to DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3, and Qwen3.8 through OpenAI-compatible and Anthropic-compatible APIs.

Start Building
Review this pull request for concurrency bugs.

02 / Lightning fast inference

DeepSeek V4.1

550tok/s

up to 5.8× faster

95–247tok/s

Prism Engine

$ prism chat --model deepseek-v4.1

›

—

TTFT

—

 

Major Providers

$ chat --model deepseek-v4.1

›

—

TTFT

—

 

03 / Serverless Model APIs

Run open source models through a single API.No infrastructure to manage.

One platform for the full AI inference stack

Model APIs, GPU infrastructure, and agent runtimes. All in one platform.

Top Open Source models

DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen on one key.

DeepSeek
GLM
Kimi
Qwen
Qwen
Kimi
GLM
DeepSeek

Better price-performance

Up to 50% less than major cloud providers. Not because we cut corners, because we built the infrastructure.

Scale with your workload

Start small and scale seamlessly from APIs to dedicated clusters.

Serverless API

Dedicated endpoints

GPU control

Built for production reliability

Stable infrastructure with low latency, high throughput, and reliable uptime at scale.

99.99%

Uptime SLA

<50ms

P50 latency

Global

PoPs

Operational

04 / Questions

Frequently Asked Questions

Models, compatibility, pricing, and how to call the API.

Contact support
What is Prism?

Prism is an OpenAI- and Anthropic-compatible inference API for coding agents. It serves open-weight frontier models so you can run the primary agent loop without standing up your own GPU stack.

Which models can I call?

The public lineup is DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen. Currently available model ids are deepseek-v4.1-flash, deepseek-v4-flash, glm-5.3, kimi-k3, qwen-3.8, and qwen.

Which API formats are supported?

Prism supports OpenAI Chat Completions at https://api.prisminference.com/v1 and Anthropic Messages at https://api.prisminference.com. Keep your existing client and change the base URL, API key, and model id.

How do I get an API key?

Create a Prism account, then create an API key from your account settings. Store it as PRISM_API_KEY and send it with your requests.

How is inference priced?

Usage is billed per token. There is no per-seat fee. See Pricing for serverless, elastic, dedicated, and batch options.

Which coding agents work with Prism?

Any client or gateway that can send OpenAI Chat Completions or Anthropic Messages, including Cursor, Claude Code, and configurable custom loops. Clients that require the OpenAI Responses API need a compatibility gateway.

Does Prism retain prompts or train on them?

No. Prism applies strict zero data retention to API inputs and outputs and never uses them for training. See the Privacy Policy for details.

Can I get a private deployment?

Yes. Book a call for dedicated capacity in our cloud or yours.

05 / Get a key

The cheapest and fastest inference on the market.