Efficient MoE model with 3B active parameters, strong price-to-performance, and fast responses for coding agents that iterate through frequent tool calls.
Input
$0.15
/ M tokens
Cache read
$0.05
/ M tokens
Output
$1.00
/ M tokens
Context
256K
tokens
Max output
64K
tokens
Serving
Serverless
Pay per token