Open Source

Qwen 3.6 27B scores 1.79% on DeepSWE, ranking 18th among 20 models

Open-source model trails closed rivals despite 70 hours of compute and 44k tokens per task.

Deep Dive

Alibaba's Qwen 3.6 27B, running on a single RTX 6000 Pro Blackwell GPU via RunPod, scored just 1.79% on the DeepSWE software engineering benchmark, placing 18th out of 20 models. The evaluation took 70 hours with an average of 32 minutes per task and 44k output tokens per task. The model used FP8 precision with BF16 KV cache, a 262k context window on VLLM, and the mini-swe agent harness. Only one rollout per task was used (instead of the official four) to save time, meaning no score range is shown. The benchmark was orchestrated by Codex 5.5xhigh.

The results reveal a stark divide: even the best open-source model, Kimi-k2.6, remains far behind leading closed-source systems. The author notes that Qwen 3.6 27B is a "local poor man's SOTA"—accessible to run locally but severely underperforming. As models become competitive, they tend to go closed-source quickly. The commentary suggests that local models are losing ground, and the race may be unwinnable for open-source advocates without massive compute and proprietary data.

Key Points
  • Qwen 3.6 27B's 1.79% DeepSWE score ranks 18th out of 20, above only Haiku 4.5 and Minimax M2.7.
  • Benchmark ran on single RTX 6000 Pro Blackwell GPU using FP8 inference and VLLM over 70 hours.
  • Best open-source model Kimi-k2.6 still lags far behind closed-source leaders, highlighting local model limitations.

Why It Matters

Local open-source models continue to fall behind closed-source rivals, raising doubts about viability for high-level software engineering tasks.

📬 Get the top 10 AI stories daily