Which local LLMs survive the hype? Reddit reveals the long-term favorites
Users ditch flashy newcomers for consistent performers like Llama 3 and Mistral.
Deep Dive
A Reddit user shares that despite checking out each new local model, they always end up returning to the same few they already use, and asks the community which models they have kept using long-term and why—citing speed, writing style, VRAM use, long context, or other features—as well as which models initially impressed but later became annoying.
Key Points
- Llama 3 (8B/70B) and Mistral 7B are the most commonly kept models due to speed and low VRAM (4–8 GB).
- Qwen 2.5 (7B/14B) praised for long-context handling (up to 32K tokens) and consistent tone.
- Phi-3-mini and Gemma 2 lost users over time due to short context windows and awkward writing styles.
Why It Matters
Real-world performance beats benchmarks: professionals prioritize speed, VRAM efficiency, and consistent output over raw scores.