GPT-5.6 Sol scores perfect 100% on brutal music theory benchmark
A 12-question chord spelling test designed to stump LLMs—GPT-5.6 Sol aced every one.
Waldvogel, a LessWrong user, designed a 12-question chord spelling test to benchmark LLMs on undergraduate music theory. The questions were deliberately cruel—asking for chords like a "iiø6/5 of ♭VI in F major" or a "7#5 chord built on the supertonic of A-sharp minor," the latter containing a rare triple sharp. All questions were submitted in a single prompt, forcing models to process an entire set of complex musical notation at once. The prompt demanded theoretically correct spellings, banning enharmonic shortcuts.
Results shocked the author. GPT-5.6 Sol (the only premium model tested) scored a flawless 100%, while Claude Sonnet 5 hit 91% and GPT-5.5 scored 83%. The free models doubled the author's expectations. Older models crumbled: Claude Sonnet 4 base scored 0%, GPT-4.1 just 16%, and Gemini 2.5 Pro even outperformed its newer 3.1 Pro. With these results, Waldvogel concedes the test has no future as a benchmark—LLMs have already surpassed it. The hardest question (Q8) stumped 67% of models, but top-tier performance signals that advanced symbolic reasoning in niche domains is now a solved problem for frontier AI.
- GPT-5.6 Sol scored a perfect 100% on the 12-question chord spelling test
- Claude Sonnet 5 scored 91%, while older Claude Sonnet 4 base scored 0% and GPT-4.1 scored just 16%
- The hardest question—a 7#5 chord on the supertonic of A-sharp minor—was answered correctly by only 33% of all LLMs tested
Why It Matters
Frontier LLMs now master niche expert domains like music theory, forcing benchmark designers to constantly raise the bar.