Senior SWE Bench tests AI on underspecified, real-world feature tasks
Real-world coding with fuzzy specs—AI struggles to ask for clarification.
Deep Dive
A Reddit user submitted a post with a link and comments.
Key Points
- Benchmark comprises 200+ underspecified tasks drawn from real GitHub feature requests
- Top models (GPT-4, Claude) achieve <25% success rate when ambiguity is present
- Tasks require active clarification seeking and iterative problem-solving, not just code completion
Why It Matters
For software engineers, this reveals AI's core weakness—handling ambiguous real-world requirements—and signals where tool improvements are needed.