Alibaba's New AI Runs Blazing Fast on a Handful of Home GPUs
One hobbyist setup juggled five AI chats at once — each faster than you can read.
Someone's been putting Qwen3.8-Flash-Next through a batch of different agentic coding use-cases, mostly on a pi harness, and reports it has seriously exceeded their expectations. Both speed and quality were a surprise: 3-5 concurrent streams running at roughly 100 tokens per second generation each, while a single stream easily hits 150+ tokens per second, with prefill at 10k+ tokens per second. They also shared the model link they used and a metric dashboard covering the last few days.
- A hobbyist ran a new Alibaba AI model on four AMD graphics cards and got speeds over 150 words per second — far faster than you can read
- He could run three to five separate AI conversations at once, each still fast, which is unusual for home hardware
- It's an early, unofficial test using a compressed version of the model, so treat the numbers as a promising hint rather than a proven result
Why It Matters
Suggests powerful, private AI could soon run on hardware you own instead of rented cloud servers.