Kimi K3 Fast
kimi-k3-fast-api- 🏢 Moonshot AI's latest K-series generation, listed here as a separate endpoint.
- 🧠 Open-weight multimodal reasoning model, per Moonshot's Hugging Face model card.
- 📏 One-million-token context for whole repositories and large document sets.
- 👁️ Accepts images alongside text.
- 🆕 New attention stack: Kimi Delta Attention plus Attention Residuals, per the model card.
- 🔧 Tool calling, web search and long-horizon agentic coding workflows.
- 🎯 Aimed at repository navigation, debugging and iteration against logs and tests.
- 🔒 Weights published publicly by Moonshot AI on Hugging Face.
Moonshot is an AI research lab known for developing the Kimi family of large language models. The organization has gained recognition for building capable reasoning-oriented models, with the Kimi line representing its flagship series of text generation systems.
Explore 4 more models by Moonshot →Kimi K3 Fast is listed in this catalog as a separate speed-oriented endpoint of Moonshot AI's K3 model, the open-weight multimodal reasoning system the company publishes on Hugging Face; provider documentation specific to the Fast variant is not available, so its serving characteristics are not described here. The underlying K3 model, according to Moonshot's model card, accepts a one-million-token context window and takes images as well as text, and the catalog description positions it for complex coding, knowledge work and long-horizon agentic workflows rather than short chat turns.
The clearest generational change from the K2 line — including [[sibling:kimi-k2-5|Kimi K2.5]], [[sibling:kimi-k2-6|Kimi K2.6]] and the coding-focused [[sibling:kimi-k2-7-code|Kimi K2.7 Code]] — is the attention stack. According to Moonshot's model card, K3 introduces Kimi Delta Attention, a hybrid linear attention mechanism, together with Attention Residuals, which allow representations to be retrieved selectively across depth. Both changes are directed at sustaining long contexts and extended tool-use loops. The standard [[sibling:kimi-k3|Kimi K3]] endpoint exposes the same released model family.
Typical uses described by the provider include navigating large codebases, cross-file refactors, tool- and terminal-driven agent runs, and research across long document sets, with screenshots, logs, test results and runtime feedback usable as inputs the model iterates against. Weights are released under Moonshot's own license terms rather than a standard open-source license.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $4.0438 | $0.4044 | $20.2188 | chat | — |
| antseed-opal-badger-2580 0xc85d…2580 | 2 | gated | $2.25 | $2.25 | $11.25 | chat | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.