GLM 4.7 Flash Heretic
glm-4.7-flash-hereticolafangensan-glm-4-7-flash-hereticolafangensan-glm-4.7-flash-heretic- 🆕 Uncensored "abliterated" variant of the GLM-4.7-Flash base model.
- 🧠 Built on a 30B-A3B mixture-of-experts base (~3B active).
- 🔧 Decensored via the Heretic method, credited to Olafangensan.
- 📏 Catalog context window of 200K tokens; FP8 quantized.
- ⚡ Tuned for fast inference and unfiltered creative writing.
- 🎯 Supports reasoning, function calling, and web search.
- 🔒 MIT-licensed community release, distributed via Hugging Face.
- 📚 Roughly 3,600 downloads on Hugging Face at writing.
Community represents the broader ecosystem of independent creators, fine-tuners, and open-source contributors who build and share models outside any single corporate lab. Rather than a formal organization, this category collects specialized models developed by individual…
Explore 4 more models by Community →GLM 4.7 Flash Heretic is a community-modified version of GLM-4.7-Flash, a roughly 30-billion-parameter mixture-of-experts model that activates about 3 billion parameters per token for lightweight deployment. The "Heretic" suffix denotes abliteration — an automated decensoring technique from the open-source Heretic project — applied here by the contributor Olafangensan to strip refusal behavior while aiming to preserve the underlying model's capabilities.
The practical difference from the unmodified GLM-4.7-Flash is behavioral rather than architectural: the variant is designed to answer prompts without the typical "I cannot help with that" responses, targeting unfiltered dialogue and creative writing. It retains the base model's reasoning traces, function-calling, and tool-use support. Downstream quant makers have anecdotally observed that the decensoring process can shorten the model's reasoning blocks and "focus" outputs, though these are informal notes rather than measured results.
This catalog entry lists a 200K-token context window and FP8 quantization. It carries an MIT license and is distributed through Hugging Face for local use via tools such as vLLM, SGLang, or Ollama.
As a hobbyist-oriented release from the broader open community rather than a vendor flagship, there are no official benchmark figures specific to this Heretic variant; users should treat it as an experimental, uncensored derivative of the base GLM-4.7-Flash.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Venice.ai Proxy 0x1f22…18c9 | 100.00 | gated | $0.07 | $0.07 | $0.40 | chat,reasoning,web-search,uncensored | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 40.35 | gated | $0.0519 | $0.026 | $0.2966 | — | — |
| Apex TEE Test 0xe672…7955 | 0.51 | gated | $0.063 | $0.021 | $0.36 | chat,uncensored,open-source,privacy,reasoning,long-context,agents | openai-chat-completions |
| antseed-neon-puma-944e 0x6650…944e | 0.00 | gated | $0.0477 | $0.0477 | $0.2728 | chat,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 0.00 | gated | $0.0283 | $0.0283 | $0.1616 | chat,coding,json,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.