Audio & Speech

Hawley's 2.55M-parameter Swin V2 world model boosts music AI key detection to .70

Self-supervised model runs in 0.6s on Apple MPS, enabling real-time collaborative music agents.

Deep Dive

Scott Hawley presents a hierarchical self-supervised world model for symbolic music, designed to let AI co-creation agents 'listen' with rich representations while keeping humans in control. The 2.55M-parameter Swin V2 encoder trains on MIDI piano-roll images using JEPA-style objectives—pitch/time-shift equivariance, masked embedding prediction, and a distributional regularizer—with zero labels or music-theory vocabulary. Probing reveals that musical properties decode at matching time scales: phrase boundaries appear at coarse levels, note density and harmonic details at fine levels.

A conditional flow-matching decoder, following the Representation AutoEncoder paradigm, reconstructs target windows with pixel F1 of 0.996 and enables masked inpainting via per-level conditioning dropout. A small chord-supervision head boosts chord recovery to .54, and key detection—never directly supervised—jumps to .70. The pipeline generates suggestions in 2.8 seconds on CPU, 0.6 seconds on Apple MPS, and is demonstrated in a live demo. Combined with an LLM-based 'brain,' this forms the core of a collaborative music agent that serves, not replaces, human agency.

Key Points
  • JEPA-style self-supervised training extracts temporal and phrase structure without music-theory labels
  • Key detection accuracy improves from .16 to .70 with a small chord-supervision head that also lifts chord recovery to .54
  • Runs in 2.8s on CPU and 0.6s on Apple MPS; conditional flow-matching enables inpainting and controlled variations at pixel F1 0.996

Why It Matters

Real-time, label-efficient music AI that augments human creativity—not replaces it—could redefine digital audio workstations and live jamming.

📬 Get the top 10 AI stories daily