New Free AI Engine Makes Your PC Run Chatbots 2.6x Faster
This free software could save you thousands on cloud AI and speed up your work.
A developer known as jesdga95 has built Basalt, a specialized engine that makes running large AI models on your own computer dramatically faster. On a setup with two Nvidia graphics cards (a 5090 and a 5060 Ti), it generates text at 665 tokens per second for structured tasks and 354 for prose—about 2.6 times faster than the Strata engine it's based on. Even a single 5090 gets 585 tokens per second, which is still blazing fast.
Why does this matter? Running AI locally means you don't pay cloud fees, your data stays private, and you're not dependent on internet connections. Basalt supports up to eight simultaneous users, making it suitable for small teams or families. It also includes a custom vision encoder that's up to 3x faster than the standard one, so it can process images quickly. The engine is open source under MIT license, and the developer provides pre-packaged model files for different hardware budgets.
The big catch: Basalt only works on Nvidia's newest Blackwell architecture (like the 5090 and 5060 Ti) and requires Linux. If you have an older Nvidia card, AMD, or Intel GPU, or use Windows or Mac, you're out of luck. The developer says he only supports what he can personally test. Also, the model it runs (Qwen3.8 Flash-Next) is a specific, smaller AI, so it won't match the biggest cloud models in raw capability.
Still, for hobbyists and small businesses with the right hardware, this is a big leap. It shows how fast local AI is becoming, and it's all free to use and modify. The developer even jokes that it's 'vibe coded slop' but fast slop—highlighting a new trend of AI-assisted coding. If you have a high-end gaming PC and want to experiment with AI without cloud costs, this could be a game-changer once it supports more hardware.
- Basalt is a free tool that makes AI models run up to 2.6x faster on Nvidia's newest gaming cards.
- It can handle 8 users at once, making it great for small teams or families to share a local AI.
- The catch: it only works on Linux and the latest Nvidia Blackwell chips, so most computers can't run it yet.
Why It Matters
Faster local AI means lower costs, better privacy, and less reliance on cloud services for those with the right hardware.