Alibaba Qwen3.7-Max challenges US models with speed and agentic focus
Qwen3.7-Max delivers 208 tokens/sec and 23% hallucination rate, but abstains from half of prompts.
Alibaba has unveiled Qwen3.7-Max, its flagship closed-weight large language model designed for long-running agentic work and text-only applications like coding and scientific discovery. The model can process up to 1 million input tokens and generate up to 64,000 output tokens at a speed of 208 tokens per second, tying with Gemini 3.5 Flash for third-fastest overall. On the Artificial Analysis Intelligence Index, Qwen3.7-Max ranks seventh (reasoning score 56.6), trailing top US models from OpenAI, Anthropic, and Google but leading all Chinese competitors. It also achieves the lowest hallucination rate among frontier models at 23%, though this comes partly from declining to answer more than half of prompts. Pricing is set at $2.50/M input tokens, $0.25/M cached input, and $7.50/M output via Alibaba Cloud Model Studio.
Alibaba highlights Qwen3.7-Max's agentic capabilities through an internal test where it autonomously optimized an attention kernel on unfamiliar hardware, making 1,158 tool calls and 432 kernel evaluations over 35 hours to produce code that ran 10x faster than a standard reference. However, these claims lack independent benchmark validation yet. The model also supports tool use, prompt caching, and native compatibility with OpenAI and Anthropic API specs. Qwen3.7-Max continues Alibaba's trend of closing top-tier models (alongside Qwen3.6-Max-Preview and Qwen3.6-Plus) while keeping smaller variants like Qwen3.6-27B open source. This shift, following leadership changes in the Qwen team, suggests Alibaba is prioritizing revenue over reach with its most capable models.
- Qwen3.7-Max processes 1M input tokens, outputs at 208 tokens/sec, ranks 7th on the AI Index
- Lowest hallucination rate (23%) among frontier models, but abstains from >50% of prompts
- Closed-weight model; Alibaba also released multimodal Qwen3.7-Plus-Preview
Why It Matters
Qwen3.7-Max is the smartest Chinese LLM, fastest overall, but its closed source marks a strategic revenue pivot.