Viral Wire

Alibaba's Qwen3.7-Plus beats GPT-5.4 on visual agent benchmarks, costs 6x less

This multimodal agent reads screens, navigates apps, and writes code autonomously.

Deep Dive

Alibaba's Qwen team released Qwen3.7-Plus on June 2, a multimodal agent model that combines visual perception, graphical user interface (GUI) control, and code generation into a single autonomous loop. The model accepts text, images, and video as input but outputs only text, enabling it to read screens, navigate apps, write code from visual templates, and invoke external tools without human intervention. In a demonstration, it autonomously developed an English vocabulary learning application over 11 hours, generating more than 10,000 lines of code across 1,000-plus agent calls.

Pricing is aggressive: $0.40 per million input tokens and $1.60 per million output tokens — roughly 6x cheaper on input than Qwen3.7-Max. It supports a 1-million-token context window, up to 65,536 output tokens, and an internal chain-of-thought reasoning budget of 256,000 tokens. A 'preserve_thinking' parameter maintains reasoning state across tool calls. Benchmarks show leadership in screen understanding (ScreenSpot Pro: 79.0 vs. GPT-5.4's 67.4; AndroidWorld: 81.0 vs. Gemini 3.1 Pro's 70.7), but it trails on pure reasoning tasks like GPQA Diamond (90.3 vs. Claude Opus 4.6 Max's 91.3). Independent evaluation ranked it #53 out of 164 models, with slow output speed (~52.9 tokens/sec) and verbosity issues. Available via Alibaba Cloud's Bailian platform as an API-only offering.

Key Points
  • Qwen3.7-Plus scores 79.0 on ScreenSpot Pro vs. GPT-5.4's 67.4 and 81.0 on AndroidWorld vs. Gemini 3.1 Pro's 70.7.
  • Priced at $0.40/M input tokens — 6x cheaper than Qwen3.7-Max.
  • Autonomous demo: built an English vocabulary app with 10,000+ lines of code over 1,000 agent calls.

Why It Matters

Alibaba challenges Western AI leaders with cheaper, visual-capable agents that can automate GUI interactions and code generation.

📬 Get the top 10 AI stories daily