Access models from the top open-source labs through a single, unified Inference API. Pay only for the tokens you use.
Showing GLM-5.3, item 3 of 6.
Model 05In progress
Qwen3.8 2.4T-A95B
qwen-3.8
Model 02Available
DeepSeek-V4-Flash
deepseek-v4-flash
Model 03In progress
GLM-5.3
glm-5.3
Model 04In progress
Kimi K3
kimi-k3
Model profile · In progress
Z.AI’s strongest GLM-5 agentic coder — the same base as GLM-5.2, with every gain coming from post-training.
Input
$1.40
/ M tokens
Cache read
$0.26
/ M tokens
Output
$4.40
/ M tokens
Context
1M
tokens
Max output
128K
tokens
Serving
Serverless
Pay per token
02 / Production workloads
Built for production agent workloads
Prism serves open-weight models through one reliable inference API built for long tool traces, large context windows, and production-scale agent workloads.
Low latency
Speculative decoding and batching shaped around code generation, not general chat.
High throughput
Batched serving across a GPU fleet so a customer agent can stay in a tight loop.
Reliable uptime
Redundant serving paths keep inference available when individual workers or GPUs go offline.