MazzikaAI steers Google Lyria for real-time Arabic maqam accompaniment
Subsecond latency, six maqamat, and zero fine-tuning—this AI listens and adapts
Generative music models have long ignored Arabic maqam, with its microtonal intervals, modal complexity, and ornamented call-and-response structures. MazzikaAI, developed by Jiaxin Du and colleagues, tackles this by using natural language as a real-time control interface. Instead of retraining a model, the system compiles live MIDI, gesture data, and inferred harmony into continuously updated text prompts that steer Google's streaming text-to-music generator, Lyria RealTime. This knowledge-based layer embeds expert rules for six core maqamat, characteristic ornaments, and ensemble dynamics, achieving subsecond key-to-audible-update latency.
The system's empirical results show significantly more off-grid quartertone content compared to baseline generations, confirming that deterministic knowledge-based rules can reliably ground foundation models in non-Western musical traditions—no fine-tuning required. MazzikaAI demonstrates a scalable blueprint for real-time human-AI co-creation: an AI partner that listens, adapts dynamically, and respects idiomatic microtonal structures. Beyond Arabic maqam, this architecture opens a generalizable path for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.
- MazzikaAI compiles live MIDI, gestures, and harmony into text prompts for Google Lyria RealTime, avoiding model fine-tuning
- Covers six core maqamat with subsecond key-to-audible-update latency for real-time responsiveness
- Dynamically generated prompts significantly increase off-grid quartertone content, grounding outputs in microtonal scales
Why It Matters
MazzikaAI proves foundation models can serve non-Western music in real time, expanding AI co-creation beyond Western traditions.