OptionalCreativeVersion 1.0.0

HeartMuLa: Open-Source Music Generation from Lyrics and Tags

HeartMuLa: Suno-like song generation from lyrics + tags.

Written by Neura Market from the official Hermes Agent documentation for Heartmula. Commands, paths, and version numbers are reproduced from the source unchanged.

Read the official documentation

HeartMuLa is an open-source family of music foundation models (Apache-2.0) that generates full songs conditioned on lyrics and tags, with multilingual support. It is comparable to Suno but runs locally or offline, making it a strong choice for developers, musicians, and AI researchers who want to generate music without relying on a cloud API.

What it does

HeartMuLa takes a text file of lyrics (with optional structural tags like [Verse], [Chorus]) and a text file of comma-separated style tags (e.g., "piano,happy,wedding") and produces a 48kHz stereo MP3 song. The generation is conditioned on both inputs, with lyrics generally dominating the output. The model family includes:

  • HeartMuLa (3B/7B) - the core music language model
  • HeartCodec - a 12.5Hz music codec for high-fidelity audio reconstruction
  • HeartTranscriptor - a Whisper-based lyrics transcription tool
  • HeartCLAP - an audio-text alignment model

In practice, you write your lyrics with bracketed section markers, pick a few descriptive tags, and run a single Python script. The generation takes about as long as the song itself (real-time factor ~1.0), so a 4-minute song takes roughly 4 minutes on a GPU.

Before you start

Hardware requirements

  • Minimum: 8GB VRAM with --lazy_load true (models load/unload sequentially)
  • Recommended: 16GB+ VRAM for comfortable single-GPU usage
  • Multi-GPU: Use --mula_device cuda:0 --codec_device cuda:1 to split across GPUs
  • The 3B model with lazy_load peaks at ~6.2GB VRAM

Platform notes

  • Triton is not available on macOS; GPU acceleration requires Linux/CUDA.
  • RTX 5080 incompatibility has been reported in upstream issues.
  • No GPU? You can run on CPU with --mula_device cpu --codec_device cpu, but expect generation to take 30-60+ minutes per song versus ~4 minutes on GPU. CPU mode also needs ~12GB+ free RAM. If you lack an NVIDIA GPU, consider a cloud GPU service (Google Colab free tier with T4, Lambda Labs) or the online demo at https://heartmula.github.io/.

Installation

1. Clone the repository

cd ~/  # or desired directory
git clone https://github.com/HeartMuLa/heartlib.git
cd heartlib

2. Create a virtual environment (Python 3.10 required)

uv venv --python 3.10 .venv
. .venv/bin/activate
uv pip install -e .

3. Fix dependency compatibility issues

IMPORTANT: As of Feb 2026, the pinned dependencies have conflicts with newer packages. Apply these fixes:

# Upgrade datasets (old version incompatible with current pyarrow)
uv pip install --upgrade datasets

# Upgrade transformers (needed for huggingface-hub 1.x compatibility)
uv pip install --upgrade transformers

4. Patch source code (required for transformers 5.x)

Patch 1 - RoPE cache fix in src/heartlib/heartmula/modeling_heartmula.py:

In the setup_caches method of the HeartMuLa class, add RoPE reinitialization after the reset_caches try/except block and before the with device: block:

# Re-initialize RoPE caches that were skipped during meta-device loading
from torchtune.models.llama3_1._position_embeddings import Llama3ScaledRoPE
for module in self.modules():
    if isinstance(module, Llama3ScaledRoPE) and not module.is_cache_built:
        module.rope_init()
        module.to(device)

Why: from_pretrained creates model on meta device first; Llama3ScaledRoPE.rope_init() skips cache building on meta tensors, then never rebuilds after weights are loaded to real device.

Patch 2 - HeartCodec loading fix in src/heartlib/pipelines/music_generation.py:

Add ignore_mismatched_sizes=True to ALL HeartCodec.from_pretrained() calls (there are 2: the eager load in __init__ and the lazy load in the codec property).

Why: VQ codebook initted buffers have shape [1] in checkpoint vs [] in model. Same data, just scalar vs 0-d tensor. Safe to ignore.

5. Download model checkpoints

cd heartlib  # project root
hf download --local-dir './ckpt' 'HeartMuLa/HeartMuLaGen'
hf download --local-dir './ckpt/HeartMuLa-oss-3B' 'HeartMuLa/HeartMuLa-oss-3B-happy-new-year'
hf download --local-dir './ckpt/HeartCodec-oss' 'HeartMuLa/HeartCodec-oss-20260123'

All 3 can be downloaded in parallel. Total size is several GB.

GPU / CUDA

HeartMuLa uses CUDA by default (--mula_device cuda --codec_device cuda). No extra setup is needed if you have an NVIDIA GPU with PyTorch CUDA support installed.

  • The installed torch==2.4.1 includes CUDA 12.1 support out of the box
  • torchtune may report version 0.4.0+cpu, this is just package metadata, it still uses CUDA via PyTorch
  • To verify GPU is being used, look for "CUDA memory" lines in the output (e.g. "CUDA memory before unloading: 6.20 GB")
  • No GPU? You can run on CPU with --mula_device cpu --codec_device cpu, but expect generation to be extremely slow (potentially 30-60+ minutes for a single song vs ~4 minutes on GPU). CPU mode also requires significant RAM (~12GB+ free). If the user has no NVIDIA GPU, recommend using a cloud GPU service (Google Colab free tier with T4, Lambda Labs, etc.) or the online demo at https://heartmula.github.io/ instead.

Usage

Basic generation

cd heartlib
. .venv/bin/activate
python ./examples/run_music_generation.py \
  --model_path=./ckpt \
  --version="3B" \
  --lyrics="./assets/lyrics.txt" \
  --tags="./assets/tags.txt" \
  --save_path="./assets/output.mp3" \
  --lazy_load true

Input formatting

Tags (comma-separated, no spaces):

piano,happy,wedding,synthesizer,romantic

or

rock,energetic,guitar,drums,male-vocal

Lyrics (use bracketed structural tags):

[Intro]

[Verse]
Your lyrics here...

[Chorus]
Chorus lyrics...

[Bridge]
Bridge lyrics...

[Outro]

Key parameters

ParameterDefaultDescription
--max_audio_length_ms240000Max length in ms (240s = 4 min)
--topk50Top-k sampling
--temperature1.0Sampling temperature
--cfg_scale1.5Classifier-free guidance scale
--lazy_loadfalseLoad/unload models on demand (saves VRAM)
--mula_dtypebfloat16Dtype for HeartMuLa (bf16 recommended)
--codec_dtypefloat32Dtype for HeartCodec (fp32 recommended for quality)

Performance

  • RTF (Real-Time Factor) ≈ 1.0, a 4-minute song takes ~4 minutes to generate
  • Output: MP3, 48kHz stereo, 128kbps

When not to use it

  • If you need a quick, polished song without local setup, the online demo at https://heartmula.github.io/ is faster.
  • If you lack an NVIDIA GPU and cannot tolerate 30+ minute generation times, use a cloud GPU service or the demo instead.
  • If you need real-time or low-latency generation, HeartMuLa's ~1.0 RTF may be too slow.

Limits and gotchas

  1. Do NOT use bf16 for HeartCodec, degrades audio quality. Use fp32 (default).
  2. Tags may be ignored, known issue (#90). Lyrics tend to dominate; experiment with tag ordering.
  3. Triton not available on macOS, Linux/CUDA only for GPU acceleration.
  4. RTX 5080 incompatibility reported in upstream issues.
  5. The dependency pin conflicts require the manual upgrades and patches described above.

Related skills

HeartMuLa pairs well with:

  • audiocraft-audio-generation, for other music generation approaches
  • songwriting-and-ai-music, for generating lyrics and song structures

Links

Skills the docs pair this with

More Creative skills