MusicLM logo

MusicLM

Paid

Generate original music with text input, create music from hummed or whistled melodies, and explore the MusicCaps dataset.

5.0
Inputs: text
Type
Saas
Company
Google Research

About MusicLM

MusicLM is an innovative music generation service that revolutionizes the way people create music. It enables users to generate high-quality, consistent music at 24 kHz over several minutes, based on text descriptions or whistled and hummed melodies. MusicLM outperforms previous systems in audio quality, as well as its ability to adhere to the text description. With MusicCaps, a dataset of 5.5k music-text pairs provided by human experts, users are able to explore the full potential of MusicLM. Suitable for both novice and experienced music creators alike, MusicLM is a powerful tool to bring music to life.

Key Features

Generate original music with text input: MusicLM lets users create high-quality audio by inputting text descriptions.
Create music from hummed or whistled melodies: MusicLM produces audio in 24 kHz quality from hummed or whistled melodies.
Explore MusicCaps dataset: With the 5.5k music-text pairs, users can explore the full potential of MusicLM.

Pros & Cons

Pros
  • Produces high-fidelity 24 kHz audio
  • Maintains musical coherence over several minutes
  • Supports both text and melody conditioning for flexible control
  • Publicly available dataset (MusicCaps) for research and evaluation
  • Demonstrates state-of-the-art audio quality and text adherence

Best For

Generate original music with text input: MusicLM lets users create high-quality audio by inputting text descriptions.Create music from hummed or whistled melodies: MusicLM produces audio in 24 kHz quality from hummed or whistled melodies.Explore MusicCaps dataset: With the 5.5k music-text pairs, users can explore the full potential of MusicLM.

Alternatives to MusicLM

FAQ

What is MusicLM?
MusicLM is a generative model from Google Research that creates high-fidelity music from text descriptions, whistled or hummed melodies, and other conditioning inputs.
What is MusicCaps?
MusicCaps is a publicly released dataset of 5,500 music-text pairs with rich, expert-written descriptions, designed to support research in music generation.
Can MusicLM generate long music pieces?
Yes, MusicLM generates music at 24 kHz that remains consistent and coherent over several minutes.
How does MusicLM handle melody conditioning?
MusicLM can be conditioned on both a text prompt and a melody (from whistling or humming) to produce audio that follows the melody while adhering to the text description.
What is Story Mode?
Story Mode allows users to provide a sequence of text prompts that influence how the model continues the semantic tokens derived from previous captions, enabling narrative-driven music generation.