Models
Requests run through the codegraff gateway, which proxies these model aliases to their upstream providers. Prices are per million tokens (USD).
Available models
| Model | Provider | Context | Input | Output | Cached in |
|---|---|---|---|---|---|
gemini-3.7-flash | 1M | $0.75 | $3.75 | $0.075 | |
claude-opus-4.8 | Anthropic | 1M | $5.00 | $25.00 | $0.50 |
gpt-5.5 | OpenAI | 400K | $5.00 | $30.00 | $0.50 |
claude-sonnet-4.6 | Anthropic | 1M | $3.00 | $15.00 | $0.30 |
muse-spark-1.2 | Meta | 1M | $1.25 | $4.25 | $0.15 |
muse-spark-1.2-contributor | Meta (Contributor) | 1M | $0.10 | $0.20 | $0.002 |
deepseek-v4-pro | DeepSeek | 1M | $0.435 | $0.87 | $0.0036 |
deepseek-v4-flash | DeepSeek | 1M | $0.14 | $0.28 | $0.0028 |
deepseek-v4-flash-fast | DeepSeek | 1M | $0.294 | $0.588 | $0.0735 |
grok-build | xAI | 256K | $1.00 | $2.00 | $0.20 |
kimi-k3 | Moonshot | 1M | $3.60 | $18.00 | $0.36 |
kimi-k2.7-code | Moonshot | 256K | $1.14 | $4.80 | $0.23 |
kimi-k2.7-code-highspeed | Moonshot | 1M | $2.28 | $9.60 | $0.46 |
minimax-m3 | MiniMax | 1M | $0.72 | $2.88 | $0.144 |
glm-5.2 | Zhipu | 128K | $1.68 | $5.28 | $0.312 |
mimo-v2.5-pro-ultraspeed | Xiaomi | 1M | $1.305 | $2.61 | $0.0108 |
mimo-v2.5-pro | Xiaomi | 1M | $0.435 | $0.87 | $0.0036 |
stealth-ox-alpha | Stealth | 1M | Free | Free | Free |
mimo-v2.5 | Xiaomi | 1M | $0.14 | $0.28 | $0.0028 |
Muse Spark contributor data use
Requests sent to
muse-spark-1.2-contributor may be used by Meta to train future models. Use muse-spark-1.2 when prompts and completions must not be used for Meta model training.Prompt caching
Cached input tokens are billed at the lower cached-input rate shown above. The gateway meters cache hits automatically for every provider, including Claude and Gemini, so repeated context in multi-turn and tool-calling loops is discounted with no code changes. Gemini 3.7 Flash uses the Interactions API with implicit caching (minimum 4,096 input tokens). Pass
previous_interaction_id to continue a stored thread. Its built-in Google Search, Maps, computer-use, file-search, code-execution, and URL-context tools are blocked — they bill outside the token meter. Your own function tools still work. Prices above are introductory through 2026-12-31.Using a model
Pass any alias as --model on the CLI, as the model option in the SDKs, or as the model field of an OpenAI-compatible request. The SDKs also accept a per-turn model override.