- 🏢 Tencent's native multimodal text-to-image model in the Hunyuan family.
- 📏 80B total parameters with roughly 13B activated via Mixture-of-Experts.
- 🧠 Unified autoregressive framework instead of the prevalent DiT design.
- 👁️ World-knowledge reasoning elaborates sparse prompts into richer scenes.
- 🆕 Instruct variant adds reasoning and image-to-image creative editing.
- ⚡ Distilled checkpoint enables efficient roughly 8-step sampling.
- 🎯 Photorealistic output with strong prompt adherence and fine detail.
- 📚 Handles long, detailed prompts spanning complex multi-subject scenes.
Tencent is a Chinese multinational technology conglomerate headquartered in Shenzhen. One of the highest-grossing multimedia companies in the world by revenue, Tencent has built a vast portfolio spanning gaming, social media, fintech, and cloud services. In recent years, the…
Explore 1 more model by Tencent →Hunyuan Image 3.0 is Tencent's text-to-image generator and the latest entry in the company's Hunyuan image line. Tencent describes it as a native multimodal model that unifies multimodal understanding and generation within a single autoregressive framework, with the image-generation module released openly. The architecture pairs a Mixture-of-Experts design with roughly 80 billion total parameters and about 13 billion activated during inference, using a Transfusion-style approach to bind text and image tokens.
The most concrete change from its same-family predecessor is structural. According to the technical report, version 3.0 moves beyond the prevalent DiT-based architectures to a unified autoregressive framework that models text and image modalities more directly, which the team links to more contextually rich generation. It also adds world-knowledge reasoning, automatically expanding sparse prompts with contextually appropriate detail.
On the feature side, the model emphasizes photorealistic imagery, strong prompt adherence, and fine-grained detail, alongside support for long, detailed prompts spanning multiple subjects and lighting parameters. The arXiv technical report details the data curation and post-training reinforcement learning behind these behaviors.
Tencent has since shipped additional checkpoints: an Instruct release adding reasoning-based prompt enhancement and image-to-image editing, plus a distilled variant tuned for efficient deployment with roughly 8-step sampling. Weights and code are available through Tencent's Hugging Face repository for self-hosting.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | $ / img | Categories | API |
|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | #2 | $0.045 | image,creative | openai-images |
| ▲ Apex Ant 0x73b4…e736 | 91.01 | #1 | $0.04 | image,media | openai-images |
| D5V1N2 0xd5e7…7be0 | 33.17 | gated | $0.1045 | image,creative,router,fallback | openai-images |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.