Claude Opus 4.8 Fast
claude-opus-4-8-fastclaude-opus-4.8-fastopus-4-8-fast- 🆕 Speed-optimized variant of Anthropic's most capable generally available Opus model.
- ⚡ Fast mode runs up to 2.5× output tokens per second.
- 📏 1M-token context window by default on the Claude API.
- 🔧 Builds on Opus 4.7 with better tool triggering, fewer compactions.
- 👁️ Multimodal with vision, function calling, and web search.
- 🎯 Tuned for long-horizon agentic coding and enterprise knowledge work.
- 💬 Adaptive thinking with up to 128k max output tokens.
- 🏢 Premium pricing tier; same underlying model as standard Opus 4.8.
Anthropic PBC is an American artificial intelligence company headquartered in San Francisco. Structured as a public benefit corporation, the lab develops large language models under the Claude name, with a research emphasis on building reliable, steerable, and safety-focused AI…
Explore 12 more models by Anthropic →Claude Opus 4.8 Fast is the latency-optimized configuration of [[sibling:claude-opus-4-8|Claude Opus 4.8]], Anthropic's most capable generally available model, released on May 28, 2026. Rather than a separate model, fast mode serves the same Opus 4.8 weights at higher throughput: setting speed to "fast" yields up to 2.5× more output tokens per second at premium pricing, offered initially as a research preview on the Claude API. It retains the full 1M-token context window that runs by default on the Claude API, Amazon Bedrock, and Vertex AI, plus up to 128k max output tokens and adaptive thinking.
Compared with the prior generation, [[sibling:claude-opus-4-7-fast|Claude Opus 4.7 Fast]], the underlying 4.8 model targets behavioral gains that Anthropic attributes to long-horizon agentic coding, including better long-context handling, fewer compactions, and improved compaction recovery. Anthropic also reports better tool triggering — the model is less likely to skip a required tool call, an issue some users flagged on Opus 4.7 — and improved honesty, with reduced tendency to overclaim progress.
A notable practical change: Anthropic states that fast mode for Opus 4.8 is three times cheaper than fast mode was on previous Opus models, narrowing the cost gap with the standard tier.
This is the newest entry in the Opus fast lineage. It sits alongside the broader Claude family, including [[sibling:claude-fable-5|Claude Fable 5]], which Anthropic describes as its most capable model in Claude Code for tasks larger than a single sitting.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| ▲ Apex Ant 0x73b4…e736 | 83 | #2 | $5.00 | $0.963 | $30.00 | chat,fast,premium,coding,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #1 | $3.60 | $0.36 | $18.00 | agents,chat,cheap,code,coding,developer,fast,frontier,function-calling,multimodal,reasoning,research,tasks,tools,vision,web-search | openai-chat-completions |
| antseed-aggregator 0x8564…b8b4 | 67 | gated | $15.00 | $15.00 | $75.00 | chat,coding,fast | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 60 | gated | $1.08 | $1.08 | $5.40 | chat,coding,code,reasoning,research,tasks,fast,premium,claude,opus,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $10.7385 | $1.07 | $53.6925 | base-usdc,chat,claude,code,coding,fast,github,json,low-cost,math,monitored,openai-compatible,opus,premium,reasoning,research,response-auth,surplus,tasks,tools,value,verified | — |
| D5V1N2 0xd5e7…7be0 | 26 | gated | $3.70 | $3.70 | $18.50 | chat,coding,reasoning,research,fast,frontier,router,fallback | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $4.092 | $1.20 | $20.46 | chat,coding,math | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 1 | gated | $6.00 | $6.00 | $30.00 | chat | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $3.24 | $3.24 | $16.20 | chat,fast,premium,coding,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $2.7876 | $2.7876 | $13.938 | chat,coding,json,reasoning,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.