Pay as you go
Serverless endpoints
Per-token access to frontier open source models. Call an endpoint and go: zero setup.
- Frontier Open Source Models
- No minimums
- Community support
01 / Plans
Frontier open-weight models on an OpenAI-compatible API. No per-seat fee.
Pay as you go
Per-token access to frontier open source models. Call an endpoint and go: zero setup.
Early access
Private, optimized endpoints priced per million tokens, with per-workload tuning.
Reserved capacity
Reserved GPUs sized to your roadmap, with negotiated latency SLAs.
Lowest rate
High-throughput offline jobs at the lowest per-token rate, on spare fleet capacity.
02 / Elastic Endpoints
Pay only for what you use. No minimums, no rate limits.
Price per 1M tokens
| Model | Input | Output | Cache read |
|---|---|---|---|
| GLM-5.3 | $1.40 | $4.40 | $0.26 |
| Kimi K3 | $3.00 | $15.00 | $0.30 |
| DeepSeek-V4-Flash | $0.14 | $0.28 | $0.07 |
| Qwen3.6 35B-A3B | $0.15 | $1.00 | $0.05 |
03 / Questions
Serverless endpoints are billed per token with no minimums. Call an endpoint and go: zero setup.
Private, optimized endpoints priced per million tokens, with per-workload tuning and dedicated support.
Reserved GPUs sized to your roadmap, with negotiated latency SLAs and dedicated support.
High-throughput offline jobs at the lowest per-token rate, on spare fleet capacity. Built for millions of requests per job, evals, and embeddings.
No. Every option uses strict zero data retention for API inputs and outputs, and Prism never uses them for training. See the Privacy Policy for details.
04 / Deploy