# Prism Inference: Full Agent Context Prism is a paid inference API for coding agents. It serves open-weight models through OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages. ## Start here - [Human and agent setup](https://docs.prisminference.com/guides/coding-agents) - [Agent Skill](https://prisminference.com/skills/prism-inference/SKILL.md) - [Documentation index](https://docs.prisminference.com/llms.txt) - [Quickstart](https://docs.prisminference.com/quickstart) - [Authentication](https://docs.prisminference.com/authentication) - [SDKs and clients](https://docs.prisminference.com/sdks) - [Models](https://docs.prisminference.com/models) - [Rate limits](https://docs.prisminference.com/rate-limits) - [Errors](https://docs.prisminference.com/errors) ## API contract - OpenAI base URL: https://api.prisminference.com/v1 - Anthropic base URL: https://api.prisminference.com - Bearer credential: PRISM_API_KEY - Chat Completions: POST /v1/chat/completions - Responses: POST /v1/responses - Messages: POST /v1/messages - Count Messages tokens: POST /v1/messages/count_tokens - List models: GET /v1/models - Retrieve model: GET /v1/models/{model} Callable models: 1. `deepseek-v4.1-flash`: text and base64 image input, 1M context. 2. `gemma-4-31b`: text and base64 image input, 32K context, served at BF16. Always query `GET /v1/models`. Do not send requests to models absent from the response. Context limits: `deepseek-v4.1-flash` accepts up to 1,000,000 tokens of context; `gemma-4-31b` accepts up to 32,768. Prism does not truncate silently. Large requests return a `context_length_exceeded` error instead of a partial result. ## Errors, retries, and rate limits Errors use an OpenAI-compatible envelope with a stable machine-readable `code`, a `retryable` flag, a `fix` hint, and a `docs_url`: ```json { "error": { "message": "API key rate limit exceeded.", "code": "rate_limit_exceeded", "retryable": true, "fix": "Wait for the Retry-After interval, then retry with bounded exponential backoff.", "docs_url": "https://docs.prisminference.com/errors" } } ``` Retry `408`, `429`, `500`, `502`, `503`, and `504` when the request is safe to repeat. Do not retry other `4xx` responses until the request or credential changes. A `429` carries `Retry-After` in seconds; honor it, then use bounded exponential backoff with jitter. See https://docs.prisminference.com/rate-limits. ## Prompt-first setup The coding agent creates the account, receives its restricted key immediately, and stores it without exposing it. The human only supplies the emailed OTP; they never see or re-enter the key. Required sequence: 1. Read the matching agent guide. 2. Inspect existing config without exposing secrets. 3. Collect the owner email and username, then send `human_email` and `username` to `POST https://prisminference.com/api/agent-signups`. 4. Immediately pass the returned `api_key` into the harness-specific native secret mechanism. Never repeat it to the human or include it in the final response. 5. Ask the human only for the emailed OTP and call `POST /api/agent-signups/verify` with the key as Bearer authentication. 6. Show a redacted preview and make the smallest targeted provider change. 7. Validate the configured provider and model using the agent's native command and live model list. Native key destinations: - Codex: the `PRISM_API_KEY` environment selected by `env_key`. - OpenCode: its auth store, written through `/connect` with custom provider ID `prism`. - Hermes: `~/.hermes/.env`, written by `hermes config set PRISM_API_KEY`. - OpenClaw: its auth store, written by piping the key to `openclaw models auth paste-api-key --provider prism` on stdin. Agent guides: - Codex: https://docs.prisminference.com/guides/codex - Cursor: https://docs.prisminference.com/guides/cursor - OpenCode: https://docs.prisminference.com/guides/opencode - Hermes: https://docs.prisminference.com/guides/hermes - OpenClaw: https://docs.prisminference.com/guides/openclaw ## Compatibility boundaries Chat Completions supports streaming, function tools, structured outputs, reasoning controls, and supported base64 images on DeepSeek V4.1 Flash. Responses is stateless. Prism does not store responses, chain `previous_response_id`, run background jobs, or provide hosted web search, file search, computer, shell, local shell, or native apply-patch tools. Anthropic Messages supports streaming, tools, thinking controls, and token counting. Prism accepts Bearer or `x-api-key` authentication but Bearer is the recommended common format. ## Billing and security Prism is prepaid. Configuration and model-catalog validation do not consume inference credit. Any live model request does. The human funds the workspace in the browser. Agents must not access payment methods, enable auto-top-up, or expose credentials. Prism applies zero data retention to API inputs and outputs and does not use them for model training. - Privacy: https://prisminference.com/privacy - Terms: https://prisminference.com/terms - Billing: https://prisminference.com/app/settings/billing - API keys: https://prisminference.com/app/settings/api-keys ## Rate limits Per-key rate limits return `429` with a `Retry-After` header in seconds. Honor `Retry-After`, then retry with bounded exponential backoff and jitter. See https://docs.prisminference.com/rate-limits. ## Machine-readable - [OpenAPI 3.1](https://docs.prisminference.com/openapi.yaml) - [API catalog](https://prisminference.com/.well-known/api-catalog) - [Security contact](https://prisminference.com/.well-known/security.txt) - [Pricing](https://prisminference.com/pricing.json) - [Status](https://status.prisminference.com) - [Documentation index](https://docs.prisminference.com/llms.txt)