MusicLM
FreeA model by Google Research for generating high-fidelity music from text descriptions.
About MusicLM
MusicLM by Google Research generates high-fidelity music at 24 kHz from text descriptions such as 'a calming violin melody backed by a distorted guitar riff'. It uses a hierarchical sequence-to-sequence modeling approach to produce music that remains consistent over several minutes. The model outperforms previous systems in both audio quality and adherence to text descriptions. MusicLM can be conditioned on both text and a melody, allowing it to transform whistled and hummed melodies according to a text caption. It also supports 'story mode' where sequential text prompts influence the continuation of semantic tokens. To support future research, the authors publicly released MusicCaps, a dataset of 5.5k music-text pairs with rich descriptions by human experts.
Key Features
Pros & Cons
- Produces high-fidelity (24 kHz) audio with musical coherence
- Maintains consistency over several minutes of music
- Flexible conditioning: text alone or text with melody input
- Outperforms prior state-of-the-art systems in quality and text adherence
- Includes publicly available benchmark dataset (MusicCaps)