Access models from the top open-source labs through a single, unified Inference API. Pay only for the tokens you use.
Showing Qwen3.8 2.4T-A95B, item 5 of 6.
Model 05In progress
Qwen3.8 2.4T-A95B
qwen-3.8
Model 06In progress
Qwen3.6 35B-A3B
qwen
Model 01Available
DeepSeek-V4.1-Flash
deepseek-v4.1-flash
Model 04In progress
Kimi K3
kimi-k3
Model profile · In progress
Qwen's 2.4T-parameter open-weight MoE with 95B active, built for coding, research, and long-horizon agent tasks.
Input
$2.00
/ M tokens
Cache read
$0.25
/ M tokens
Output
$6.00
/ M tokens
Context
256K
tokens
Max output
128K
tokens
Serving
Serverless
Pay per token
02 / Production workloads
Built for production agent workloads
Prism serves open-weight models through one reliable inference API built for long tool traces, large context windows, and production-scale agent workloads.
Low latency
Speculative decoding and batching shaped around code generation, not general chat.
High throughput
Batched serving across a GPU fleet so a customer agent can stay in a tight loop.
Reliable uptime
Redundant serving paths keep inference available when individual workers or GPUs go offline.