Access models from the top open-source labs through a single, unified Inference API. Pay only for the tokens you use.
Showing DeepSeek-V4.1-Flash, item 1 of 6.
Model 01Available
DeepSeek-V4.1-Flash
deepseek-v4.1-flash
Model 02Available
DeepSeek-V4-Flash
deepseek-v4-flash
Model 03In progress
GLM-5.3
glm-5.3
Model 06In progress
Qwen3.6 35B-A3B
qwen
Model profile · Available
DeepSeek's 552B successor for coding, reasoning, and tool use, with a 1M-token context window for long agent runs.
Input
$0.17
/ M tokens
Cache read
$0.006
/ M tokens
Output
$0.63
/ M tokens
Context
1M
tokens
Max output
384K
tokens
Serving
Serverless
Pay per token
02 / Production workloads
Built for production agent workloads
Prism serves open-weight models through one reliable inference API built for long tool traces, large context windows, and production-scale agent workloads.
Low latency
Speculative decoding and batching shaped around code generation, not general chat.
High throughput
Batched serving across a GPU fleet so a customer agent can stay in a tight loop.
Reliable uptime
Redundant serving paths keep inference available when individual workers or GPUs go offline.