Gemma 4 beats Gemini and Claude Opus on practical instruction tasks
User reports Gemma 4 nails nuance while Claude Opus 5 sounds 'AI-ish' and sloppy
A Reddit post by u/MaxDev0 is going viral after claiming that Google's Gemma 4 (26B A4B), a relatively small MoE model, consistently outperforms much larger and more marketed models like Gemini 3.5 Flash and Claude Opus 5 in practical, real-world tasks. The user ran side-by-side tests focused on instruction following and nuance, not typical math or coding benchmarks. In one test, Gemini failed to follow email composition instructions and missed tone entirely. Claude Opus 5 produced 'smart' but verbose, AI-sounding output. Gemma 4, however, nailed the subtlety and understood layers of intent without being blunt.
The second test involved prompt refinement and engineering. Gemini took instructions too literally, outputting clunky, explicit commands like 'Do X, Y, Z and don't do A, B, C.' Gemma 4 naturally crafted prompts that steered away from unwanted elements without explicitly mentioning them. The user argues that leaderboards like Artificial Analysis don't align with what average users or developers actually need: models that listen, understand nuance, and don't hallucinate simple instructions. The post asks the community for better benchmarks that measure real-world utility and invites others to share similar experiences, sparking a broader conversation about the disconnect between raw technical metrics and everyday model usability.
- Gemma 4 (26B A4B) beat Gemini 3.5 Flash and Claude Opus 5 in email composition and prompt refinement tests
- Claude Opus 5 output was described as overly verbose and 'AI-ish' despite being smart
- User questions Artificial Analysis leaderboard metrics, calling for benchmarks focused on instruction following and nuance
Why It Matters
Benchmark scores may mislead developers; real-world instruction following and nuance are critical for practical AI deployment.