Viral Wire

Alibaba Qwen3.7-Max challenges US models with speed and agentic focus

Qwen3.7-Max delivers 208 tokens/sec and 23% hallucination rate, but abstains from half of prompts.

Deep Dive

Alibaba has unveiled Qwen3.7-Max, its flagship closed-weight large language model designed for long-running agentic work and text-only applications like coding and scientific discovery. The model can process up to 1 million input tokens and generate up to 64,000 output tokens at a speed of 208 tokens per second, tying with Gemini 3.5 Flash for third-fastest overall. On the Artificial Analysis Intelligence Index, Qwen3.7-Max ranks seventh (reasoning score 56.6), trailing top US models from OpenAI, Anthropic, and Google but leading all Chinese competitors. It also achieves the lowest hallucination rate among frontier models at 23%, though this comes partly from declining to answer more than half of prompts. Pricing is set at $2.50/M input tokens, $0.25/M cached input, and $7.50/M output via Alibaba Cloud Model Studio.

Alibaba highlights Qwen3.7-Max's agentic capabilities through an internal test where it autonomously optimized an attention kernel on unfamiliar hardware, making 1,158 tool calls and 432 kernel evaluations over 35 hours to produce code that ran 10x faster than a standard reference. However, these claims lack independent benchmark validation yet. The model also supports tool use, prompt caching, and native compatibility with OpenAI and Anthropic API specs. Qwen3.7-Max continues Alibaba's trend of closing top-tier models (alongside Qwen3.6-Max-Preview and Qwen3.6-Plus) while keeping smaller variants like Qwen3.6-27B open source. This shift, following leadership changes in the Qwen team, suggests Alibaba is prioritizing revenue over reach with its most capable models.

Key Points
  • Qwen3.7-Max processes 1M input tokens, outputs at 208 tokens/sec, ranks 7th on the AI Index
  • Lowest hallucination rate (23%) among frontier models, but abstains from >50% of prompts
  • Closed-weight model; Alibaba also released multimodal Qwen3.7-Plus-Preview

Why It Matters

Qwen3.7-Max is the smartest Chinese LLM, fastest overall, but its closed source marks a strategic revenue pivot.

📬 Get the top 10 AI stories daily