Audio & Speech

VIOLET's neural violin synthesis rivals top commercial virtual instruments

39 hours of training data lets VIOLET nail dynamics and bow techniques

Deep Dive

VIOLET, developed by Baotong Tian, Cynthia Lu, Vincent K.M. Cheung, Ting-Kang Wang, Jonathan Churchill, and Zhiyao Duan, tackles a long-standing gap in neural music synthesis: continuous articulation. While piano synthesis has advanced rapidly, instruments like the violin require expressive control over pitch slides, bow pressure, vibrato, and dynamics. VIOLET addresses this with a latent diffusion framework built on a Diffusion Transformer (DiT) with rectified flow, enabling it to generate audio conditioned on MIDI note sequences, note-level technique labels, and continuous dynamics curves.

To train the model, the authors curated CSV-TD, a dataset containing 39 hours of 48kHz synthetic violin audio with time-aligned annotations for MIDI notes, techniques, and dynamics. In objective and subjective evaluations, VIOLET achieved high technique adherence, accurate pitch/timing alignment, and precise dynamics control. The system outperformed the current state-of-the-art neural violin synthesizer and approached the quality of a leading commercial virtual instrument in clarity, naturalness, and dynamics tracking. Code and demo audio are publicly available, marking a significant step toward fully neural, controllable violin performance synthesis.

Key Points
  • VIOLET uses a Diffusion Transformer with rectified flow to synthesize violin audio from MIDI, playing techniques, and continuous dynamics
  • The new CSV-TD dataset provides 39 hours of 48kHz audio with fine-grained annotations for training and evaluation
  • Accepted at ISMIR 2026; VIOLET outperforms prior neural violin systems and rivals top commercial virtual instruments

Why It Matters

Neural synthesis could replace massive sample libraries, making expressive virtual violins more realistic and controllable for composers and producers.

📬 Get the top 10 AI stories daily