Local LLM Showdown: Best Models for 3x RTX 3090 Rigs
MiniMax and Step models run fast in Q3, but Gemma-4 12B is still missing.
Deep Dive
A Reddit user (jacek2023) picked local LLMs that run on three RTX 3090 GPUs, so no 300B models are included, and the user suggests skipping 200B models too—though MiniMax and Step are noted as fast in Q3. Gemma-4 12B is still missing.
Key Points
- 3x RTX 3090 (48GB total) is the baseline for 'local' models in this comparison.
- MiniMax and Step models perform well in Q3 quantization, making them viable for memory-constrained setups.
- Gemma-4 12B is notably absent from the list; its future inclusion may change rankings.
Why It Matters
Helps AI practitioners pick cost-effective, private models that run on local hardware without cloud costs.