Open Source

Free AI Models Now Rival Paid Giants, But They're 28x Slower

Free AI Models Now Rival Paid Giants, But They're 28x Slower

⚡You can run powerful AI at home, but it might take hours instead of minutes.

Deep Dive

Someone ran a multi-day, 469-question domain-specific eval of Qwen3.8-27B fine-tunes — all run locally with llama.cpp and the same thinking settings — and also ran the full set on Opus 5.5 and Astra, plus partially on Qwen3.8-Flash-Next via OpenRouter. No model got 100% accuracy.

Opus 5.5 got close and topped the list at 99.6%. Of the local models, mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4_K_M had the lower token usage and time to completion of tasks while achieving higher accuracy — a Q4_K_M fine-tune performing better than other L or XL models, and a combo of AgentionAI's Signal-3.8-27B and Shockem's Terse-Coder LoRA.

It hit the 16,384-token cap once, versus 4 cut-offs for Dirk and 4 for Unsloth. Astra had the lowest token usage, though its provider appears to mask reasoning, and OpenRouter TTFTs aren't worth reading into because of caching. The time taken to solve problems compared to frontier models is massive, especially on a potato; the poster's current feasible model's KPI over the full eval set is 28x slower, though the gap could be more forgiving

Key Points
  • Free AI models can now match paid ones in accuracy, with one scoring 99.6%.
  • But they are much slower: 9 hours vs 19 minutes for the same tasks.
  • Running AI locally gives you full privacy, while paid models send data to servers.

Why It Matters

You can save money and keep data private with free AI, but expect slower results unless you upgrade your computer.

📬 Get the top 10 AI stories daily