Viral Wire

Alibaba's Qwen3.7-Plus sees screens and codes autonomously at 6x lower cost

Alibaba's new multimodal agent beat GPT-5.4 on screen understanding benchmarks.

Deep Dive

Alibaba's Qwen team released Qwen3.7-Plus on June 2, a proprietary multimodal agent model that combines visual perception, graphical user interface control, and autonomous code generation within a single loop. The model accepts text, images, and video as input but outputs only text, enabling it to read screens, navigate apps, and write code from visual templates without human intervention. In a demo, the agent autonomously built an English vocabulary learning app over 11 hours, generating over 10,000 lines of code across 1,000+ agent calls. It also recreated Apple's Stocks app by parsing UI screenshots and generating SwiftUI code. The model is available via Alibaba Cloud's Bailian platform through API, with pricing at $0.40 per million input tokens — roughly 6x cheaper than its text-only sibling Qwen3.7-Max. It supports a 1-million-token context window with up to 65,536 output tokens and internal chain-of-thought reasoning.

On benchmarks, Qwen3.7-Plus leads on GUI-grounding tasks: 79.0 on ScreenSpot Pro (vs. GPT-5.4's 67.4) and 81.0 on AndroidWorld (vs. Gemini 3.1 Pro's 70.7). It also excels on long-context retrieval (MRCR-v2 128k: 91.7 vs. Claude Opus 4.6 Max's 84.0). However, performance drops on pure reasoning: GPQA Diamond STEM score of 90.3 trails Claude Opus 4.6 Max's 91.3, and SWE-Bench Pro coding score of ~57.6 falls short of Qwen3.7-Max's 60.6. Independent testing by Artificial Analysis ranked the model #53 out of 164 on its Intelligence Index, with slow output (52.9 tokens/sec) and a verbosity issue (110 million output tokens vs. 29 million median). Despite these limitations, Alibaba's aggressive pricing and strong GUI capabilities position Qwen3.7-Plus as a serious contender in the autonomous agent space, competing directly with OpenAI, Anthropic, and Google.

Key Points
  • Leads GUI benchmarks: ScreenSpot Pro 79.0 vs GPT-5.4's 67.4, AndroidWorld 81.0 vs Gemini 3.1 Pro's 70.7
  • Priced at $0.40/1M input tokens, ~6x cheaper than Qwen3.7-Max, with 1M token context window
  • API-only, no open weights; can autonomously build apps from screenshots (e.g., recreated Apple Stocks app)

Why It Matters

Alibaba enters the autonomous agent race with a cost-effective model that can navigate GUIs and write code.

📬 Get the top 10 AI stories daily