Coder3101 releases Gemma 4 QAT Unquantized Heretic with altered refusal
Unquantized variant offers alternative alignment with distinct divergence and refusal behavior.
Deep Dive
Developer coder3101 has released an unquantized variant of the Gemma 4 heretic collection, intentionally diverging in refusal and alignment patterns from the original. The model now requires 4-bit quantization—someone needs to do it—so users can try it as an alternative to the original.
Key Points
- Released by Reddit user coder3101 as an FP16 unquantized variant of Gemma 4 Heretic
- Intentionally different divergence and refusal compared to the original Heretic collection
- Requires 4-bit quantization for consumer GPU deployment; currently needs high VRAM
Why It Matters
Offers a new alignment option for local LLM deployment and fine-tuning research with modified safety behavior.