Preprint
Large Language Models

Foundation Models for Music

Ying-Chao Ma, Anders Oland, Anton Ragni, B. D. Sette, C. Saitis, Chris Donahue, Chenghua Lin, Christos Plachouras, Emmanouil Benetos, Elio Quinton, Elona Shatri, Fabio Morreale, Ge Zhang, Gyorgy Fazekas, Gus G. Xia, Huan Zhang, Ilaria Manco, Jiawen Huang, Julien Guinot, Liwei Lin, Luca Marinelli, Max W. Y. Lam, Megha Sharma, Qiuqiang Kong, R. Dannenberg, Ruibin Yuan, Shangda Wu, Shih-Lun Wu, Shu-yuan Dai, Shunwei Lei, Shiyin Kang, Simon Dixon, Wenhu Chen, Wehhao Huang, Xin Du, Xingwei Qu, Xu Tan, Yizhi Li, Zeyue Tian, Zhi-Xin Wu, Zhizheng Wu, Ziyang Ma, Ziyu Wang
August 26, 2024arXiv.org61 citations

61

Citations

1

Influential Citations

arXiv.org

Venue

2024

Year

Abstract

In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This comprehensive review examines state-of-the-art (SOTA) pre-trained models and foundation models in music, spanning from representation learning, generative learning and multimodal learning. We first contextualise the significance of music in various industries and trace the evolution of AI in music. By delineating the modalities targeted by foundation models, we discover many of the music representations are underexplored in FM development. Then, emphasis is placed on the lack of versatility of previous methods on diverse music applications, along with the potential of FMs in music understanding, generation and medical application. By comprehensively exploring the details of the model pre-training paradigm, architectural choices, tokenisation, finetuning methodologies and controllability, we emphasise the important topics that should have been well explored, like instruction tuning and in-context learning, scaling law and emergent ability, as well as long-sequence modelling etc. A dedicated section presents insights into music agents, accompanied by a thorough analysis of datasets and evaluations essential for pre-training and downstream tasks. Finally, by underscoring the vital importance of ethical considerations, we advocate that following research on FM for music should focus more on such issues as interpretability, transparency, human responsibility, and copyright issues. The paper offers insights into future challenges and trends on FMs for music, aiming to shape the trajectory of human-AI collaboration in the music realm.

Analysis

Why This Paper Matters

This paper arrives at a critical juncture in AI research, where foundation models are transforming domains beyond text and images. Music, with its rich temporal and multimodal structure, presents unique challenges and opportunities. The review consolidates a rapidly growing body of work, offering a structured map of the landscape. For practitioners, it clarifies which music representations are well-served by current FMs and which are neglected, guiding resource allocation and research focus.

Moreover, the paper emphasizes the lack of versatility in prior music AI systems, which were often task-specific. By framing the discussion around foundation models, it aligns music AI with the broader trend toward general-purpose models that can handle multiple tasks with minimal adaptation. This shift is crucial for practical deployment, where a single model could serve generation, understanding, and even therapeutic applications.

Technical Contributions

  • Taxonomy of Music FMs: Categorizes models into representation learning, generative learning, and multimodal learning, providing a clear framework for understanding the field.
  • Analysis of Pre-training Paradigms: Discusses various pre-training objectives, architectural choices (e.g., transformers, diffusion), and tokenization strategies specific to music (e.g., MIDI, audio tokens).
  • Controllability and Finetuning: Highlights methods for controlling generation (e.g., text prompts, melody conditioning) and finetuning approaches like instruction tuning and in-context learning, which are underexplored in music.
  • Music Agents: Introduces a section on music agents, which are AI systems that can interact with users and tools to accomplish music-related tasks, a novel area for FMs.
  • Ethical Framework: Proposes a set of ethical considerations, including interpretability, transparency, and copyright, which are often overlooked in technical reviews.

Results

As a review paper, it does not present new experimental metrics. Instead, it synthesizes findings from 61 cited works, identifying trends and gaps. For instance, it notes that while audio and symbolic representations are common, other modalities like lyrics and video are less explored. It also points out that scaling laws and emergent abilities, well-studied in NLP, have not been systematically investigated for music FMs. The paper's contribution is the comprehensive mapping of the field, which serves as a benchmark for future research.

Significance

The paper's broader impact lies in its potential to steer the music AI community toward more unified and versatile models. By highlighting underexplored areas like instruction tuning and long-sequence modeling, it encourages researchers to adopt techniques from NLP and other domains. The emphasis on ethical considerations is timely, given the legal and creative implications of AI-generated music. For Neura Market's audience, this review offers a strategic overview of where the field is heading, helping practitioners identify opportunities for innovation and collaboration.