- 🧠 Trillion-parameter Mixture-of-Experts, 32B active per token.
- 🆕 Native multimodal: continual pretraining on ~15T mixed vision-text tokens.
- 📏 256K-token context window for long documents and codebases.
- 👁️ Adds vision and cross-modal reasoning over text-only Kimi K2.
- 🔧 Instant and thinking modes for speed or deeper reasoning.
- ⚡ Native INT4 quantization for efficient inference.
- 🎯 Provider-reported HLE with tools: 51.8 text, 39.8 image.
- 🌐 Open-weight with OpenAI/Anthropic-compatible API.
Moonshot is an AI research lab known for developing the Kimi family of large language models. The organization has gained recognition for building capable reasoning-oriented models, with the Kimi line representing its flagship series of text generation systems.
Explore 4 more models by Moonshot →Kimi K2.5, released January 2026 by Beijing-based Moonshot AI, is the multimodal evolution of the Kimi K2 line. It keeps the family's trillion-parameter Mixture-of-Experts design — with roughly 32 billion active parameters per token and a 256K context window — while adding a vision pathway. Per Moonshot's model card, K2.5 was built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base, so vision and language develop together rather than as bolted-on features.
The clearest gain over the text-only predecessor Kimi K2 is native multimodality: K2.5 understands images, supports cross-modal reasoning, and grounds agentic tool use in visual inputs. It also formalizes dual operation — an instant mode for speed and a thinking mode for deeper reasoning. On Moonshot's reported Humanity's Last Exam, K2.5 scores 31.5 text and 21.3 image without tools, rising to 51.8 text and 39.8 image with tools.
K2.5 became the architectural template for its successors. [[sibling:kimi-k2-6|Kimi K2.6]] reuses the same trillion-parameter, 32B-active topology, changing the post-training recipe, while [[sibling:kimi-k2-7-code|Kimi K2.7 Code]] specializes the same MoE backbone for long-horizon software engineering. All ship open-weight with native INT4 quantization and an OpenAI/Anthropic-compatible API.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 99 | #5 | $0.28 | $0.11 | $1.75 | chat,reasoning,coding,vision,multimodal,web-search | openai-chat-completions |
| ▲ Apex Ant 0x73b4…e736 | 83 | #1 | $0.0457 | $0.0457 | $0.2079 | chat,open-source,coding,privacy,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| Open Forge 0x1d90…b0aa | 75 | #4 | $0.36 | $0.07 | $1.65 | math,coding,study,multimodal | openai-chat-completions |
| surplusintelligence.ai 0x0e49…8927 | 68 | #2 | $0.168 | $0.066 | $1.05 | agents,anon,chat,cheap,code,coding,developer,frontier,function-calling,multimodal,reasoning,research,tasks,tools,vision,web-search | openai-chat-completions |
| Open Bird 0xc0f1…8183 | 64 | #3 | $0.21 | $0.035 | $1.05 | chat,long-context | openai-chat-completions |
| NovaRoute AI 0xc50d…ed7b | 60 | gated | $0.0508 | $0.0508 | $0.231 | chat,coding,code,reasoning,tasks,kimi,value,surplus,openai-compatible,low-cost,verified,github,response-auth,base-usdc,monitored | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 40 | gated | $0.28 | $0.11 | $1.75 | chat,coding,reasoning,vision,multimodal,tools,cheap | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.4977 | $0.198 | $3.1103 | base-usdc,chat,code,coding,github,json,kimi,low-cost,math,monitored,openai-compatible,reasoning,response-auth,surplus,tasks,tools,value,verified | — |
| Meridian AI 0x8c8c…06f5 | 22 | gated | $0.0798 | $0.0798 | $0.4309 | chat,coding | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 3 | gated | $0.191 | $0.22 | $1.1935 | chat,coding,math | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 1 | gated | $0.28 | $0.28 | $1.75 | chat | openai-chat-completions |
| Skeffo Inference 0x1af8…e2b5 | 0 | gated | $0.60 | $0.60 | $3.00 | chat,vision,reasoning,agent,function-calling | openai-chat-completions |
| Leftermute 0x388b…5389 | 0 | gated | $0.2828 | $0.2828 | $1.7675 | chat,coding,json,tools | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $0.0815 | $0.0315 | $0.3703 | chat,open-source,coding,privacy,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.