Reddit survey reveals best local LLMs for 8GB to 48GB VRAM
Community benchmarks show which quantized models run on your GPU's memory budget.
Deep Dive
A Reddit user asked the community to share their experiences with kv cache and context, performance, hardware, and what they use their models for, noting how fast the field moves and wanting to congeal everyone’s experiences.
Key Points
- 8GB VRAM users run 7B models at Q4_K_M with 4K-8K context, achieving 10-20 tokens/second using KV cache offloading.
- 16GB setups handle 13B Q4 models with 8K-16K context at 15-25 tokens/second, or 70B Q2 models at slower speeds.
- 48GB (e.g., dual 3090s) runs 70B Q4 models with 64K+ context, prioritizing privacy over cloud APIs.
Why It Matters
Professionals can now choose the optimal local LLM setup for private, low-latency AI assistants without cloud dependency.