Ultra-fast, ultra-cheap reasoning model for high-throughput coding and tool-use workloads, with a 1M-token context window for long agent runs.
Input
$0.17
/ M tokens
Cache read
$0.028
/ M tokens
Output
$0.21
/ M tokens
Context
1M
tokens
Max output
384K
tokens
Serving
Serverless
Pay per token