MiniMax M3
minimax-m3MiniMax-M3- 🆕 MiniMax's frontier M-series model for coding and agentic work.
- 📏 Up to 1M-token context, guaranteed minimum 512K tokens.
- 🔧 New MiniMax Sparse Attention (MSA) architecture for long context.
- 👁️ Natively multimodal: text, image, and video input.
- 🧠 Toggleable thinking mode for reasoning or fast responses.
- ⚡ 9× prefill and 15× decode speedups vs M2 at 1M context.
- 🏢 Reported as a 428B-parameter Mixture-of-Experts model.
- 🌐 Built for tool use, function calling, and web-style retrieval.
MiniMax is an AI company building generative models across multiple modalities, with a focus that spans both language understanding and audio creation. Their rapid release cadence in early 2026—delivering several new models within just a few months—reflects an ambitious and…
Explore 3 more models by Minimax →MiniMax M3 is the latest model in MiniMax's M-series, positioned for coding, agentic workflows, and complex reasoning. It unifies three capabilities in a single checkpoint: frontier coding and agentic performance, a context window of up to 1 million tokens (with a guaranteed minimum of 512K), and native multimodality covering text, image, and video input. NVIDIA describes it as a 428-billion-parameter Mixture-of-Experts model. M3 supports a toggleable thinking mode—enabled for long-horizon agentic tasks and complex reasoning, disabled for latency-sensitive uses like conversation and code completion.
The headline change over its predecessors is the new MiniMax Sparse Attention (MSA) architecture, which enables native ultra-long-context pretraining. According to MiniMax's model card, MSA delivers 9× prefill and 15× decode speedups compared with M2 at 1M context, cutting per-token compute to roughly one-twentieth while preserving quality versus full attention. This is a substantial step up from earlier M-series entries like [[sibling:minimax-m27|MiniMax M2.7]] and [[sibling:minimax-m25|MiniMax M2.5]], which operated within shorter context windows.
MiniMax reports strong results on coding and agentic benchmarks spanning software engineering, terminal execution, and tool orchestration, with autonomous task decomposition and multi-step reasoning. Per the MiniMax docs, earlier M-series models including M2.7 and M2 remain available for existing workflows. A preview build, [[sibling:minimax-m3-preview|MiniMax M3 Preview]], was also released in the family. Weights are published on Hugging Face, and a technical report accompanies the release.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | #3 | $0.15 | $0.03 | $0.60 | chat,reasoning,coding,vision,video,multimodal,web-search | openai-chat-completions |
| Dark Signal 0x4668…62f2 | 96 | #6 | $0.21 | $0.04 | $0.84 | chat,writing,creative,coding,reasoning | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 83 | #2 | $0.135 | $0.038 | $0.54 | chat,open-source,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| Vito-Minimax 0xddfa…27fe | 80 | gated | $1.80 | $1.80 | $5.00 | chat,coding | openai-chat-completions |
| Open Ant 0xe4f6…5bc4 | 74 | #5 | $0.195 | $0.039 | $0.78 | chat,coding,reasoning,vision,multimodal,tools,long-context,cheap | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #1 | $0.1125 | $0.0225 | $0.45 | agents,chat,cheap,code,coding,developer,fast,frontier,function-calling,multimodal,reasoning,research,tasks,tools,video,vision,web-search | openai-chat-completions |
| Open Bird 0xc0f1…8183 | 64 | #4 | $0.15 | $0.03 | $0.60 | chat | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 60 | gated | $0.2079 | $0.2079 | $0.8316 | chat,coding,code,writing,creative,tasks,minimax,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 40 | gated | $0.15 | $0.03 | $0.60 | chat,coding,reasoning,vision,multimodal,tools,long-context,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.264 | $0.0523 | $1.0558 | agents,base-usdc,chat,code,coding,creative,fast,free,github,json,low-cost,math,minimax,monitored,openai-compatible,response-auth,surplus,tasks,tools,value,verified,writing | — |
| D5V1N2 0xd5e7…7be0 | 26 | gated | $0.14 | $0.028 | $0.56 | chat,coding,reasoning,agent,long-context,tasks,minimax,new | openai-chat-completions |
| ZLKPro-Api 0x0b0b…f446 | 18 | gated | $0.01 | $0.01 | $0.02 | agents,math,chat,coding | openai-chat-completions |
| Ant Army 0xc8bd…f6c9 | 15 | gated | $0.40 | $0.40 | $1.40 | chat,coding | openai-chat-completions |
| BabyCai 0x8509…27b4 | 15 | gated | $0.20 | $0.20 | $0.40 | chat,reasoning | openai-chat-completions |
| bartly.eth64.de 0x666e…4666 | 10 | gated | $0.13 | $0.025 | $0.50 | chat,coding,reasoning,long-context | openai-chat-completions |
| adfreellm Gateway 0xabd5…4e98 | 5 | gated | $0.084 | $0.0084 | $0.336 | chat,fast | openai-chat-completions |
| Hana Gateway ✅ 0x4ae1…117b | 5 | gated | $0.00 | $0.00 | $0.00 | chat,coding,fast,tools,free | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.1023 | $0.06 | $0.4092 | chat,coding,math | openai-chat-completions |
| edith 0xb269…b1a6 | 0 | gated | $2.00 | $2.00 | $7.50 | chat,coding | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $0.36 | $0.144 | $1.35 | chat,open-source,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.0929 | $0.0929 | $0.3717 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.