MiMo-V2.5

CodeAudioVisionReasoningWeb searchFunction calling
Advertised as mimo-v2-5mimo-v2.5xiaomi-mimo-v2-5
Quick reference
MiMo-V2.5 — TLDR
  • 🆕 Xiaomi's native omnimodal model: text, image, video, audio.
  • 🧠 Sparse MoE backbone, 310B total with 15B active parameters.
  • 📏 Context window extends up to 1 million tokens.
  • 👁️ Unified architecture for multimodal perception and reasoning.
  • 🔧 Function calling, web search, and code-oriented capabilities.
  • 💬 Accepts audio input alongside text, image, and video.
  • 🔒 Open weights released under the MIT license.
  • ⚡ Distributed in FP8 quantization on Hugging Face.
💰 Best price on AntSeed
$0.042 / $0.084
per 1M · cheapest in / out
📏 Context
1M tokens
🐜 Sellers
9
advertising on AntSeed
Provider

XiaomiMiMo is the large language model initiative from Xiaomi, the Chinese electronics and technology company, dedicated to developing capable open language models under the MiMo name. The effort reflects Xiaomi's broader push into foundational AI research alongside its consumer…

About this model

MiMo-V2.5 is Xiaomi's native omnimodal model, designed to understand text, images, video, and audio within a single unified architecture. It uses a sparse Mixture-of-Experts backbone with 310B total parameters and roughly 15B active per token, and supports a context window of up to 1 million tokens. Xiaomi released the model in 2026 and open-sourced the weights and tokenizer, along with a separate Base checkpoint, under the MIT license on Hugging Face.

Beyond perception, the model is oriented toward agentic and developer workflows. Its documented capabilities include reasoning, function calling, web search, and code-focused use, with audio accepted as a native input modality alongside text, images, and video. The weights are distributed in FP8 quantization, which lowers the memory footprint for serving the large MoE network.

MiMo-V2.5 belongs to Xiaomi's broader MiMo series of open models, and an accompanying Base variant is published for further fine-tuning and research. As an omnimodal release with a long-context MoE design, it extends the family's focus toward unified multimodal understanding and agentic tool use rather than text-only generation. Because the catalog lists no sibling models here, this entry is described from the model's own card and configuration rather than direct head-to-head family comparisons.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
huggingface.coXiaomiMiMo/MiMo-V2.5 · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
8.08M
input + output
Requests
203
settled calls
Buyers
7
distinct, on this model
Sellers used
3
of 9 advertising
Settled
$12.04
gross USDC, this model
Sellers serving MiMo-V2.5 (9)compare on the network explorer →
SellerReputationRoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.