Models / Open weights

Open models for every agent

Power reasoning, tool use, and long-running agent workloads through OpenAI- and Anthropic-compatible APIs.

01 / Model catalog

Access models from the top open-source labs through a single, unified Inference API. Pay only for the tokens you use.

Showing Qwen3.8 2.4T-A95B, item 5 of 6.
Model 05In progress
Qwen

Qwen3.8 2.4T-A95B

qwen-3.8

Model profile · In progress

Qwen's 2.4T-parameter open-weight MoE with 95B active, built for coding, research, and long-horizon agent tasks.

Input

$2.00

/ M tokens

Cache read

$0.25

/ M tokens

Output

$6.00

/ M tokens

Context

256K

tokens

Max output

128K

tokens

Serving

Serverless

Pay per token

02 / Production workloads

Built for production agent workloads

Prism serves open-weight models through one reliable inference API built for long tool traces, large context windows, and production-scale agent workloads.

INPUTMODELTOKENS

Low latency

Speculative decoding and batching shaped around code generation, not general chat.

010203GPU BATCH

High throughput

Batched serving across a GPU fleet so a customer agent can stay in a tight loop.

REQUESTPRIMARY FLEETFAILOVER FLEET

Reliable uptime

Redundant serving paths keep inference available when individual workers or GPUs go offline.

03 / Deploy

The cheapest and fastest inference on the market.