Fugatto Nvidia logo

Fugatto Nvidia

Paid
4.5
Inputs: text, audioOutputs: audio
Type
Saas
Company
NVIDIA

About Fugatto Nvidia

Fugatto is a foundational generative audio transformer model developed by NVIDIA Research. It is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. The model introduces a specialized dataset generation approach optimized for producing a wide range of audio generation and transformation tasks, ensuring meaningful relationships between audio and language. To achieve compositional abilities—such as combining, interpolating, or negating instructions—Fugatto leverages ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance. This enables seamless and flexible composition of instructions, leading to highly customizable audio outputs outside the training distribution. Evaluations across diverse tasks demonstrate that Fugatto performs competitively with specialized models while ComposableART enhances its sonic palette and control over synthesis. Notably, the framework can synthesize emergent sounds—sonic phenomena that transcend conventional audio generation—unlocking new creative possibilities.

Key Features

Follows free-form text instructions with optional audio inputs
Specialized dataset generation approach for audio tasks
ComposableART inference-time technique for compositional guidance
Enables seamless combination, interpolation, or negation of instructions
Synthesizes emergent sounds beyond conventional audio generation
Competitive performance with specialized models across diverse tasks

Pros & Cons

Pros
  • Versatile: handles both audio generation and transformation tasks
  • Supports free-form text prompts for easy control
  • Compositional abilities allow complex instruction combinations
  • Competitive with specialized models in various evaluations
  • Capable of synthesizing novel emergent sounds
Cons
  • Currently a research prototype, not available as a commercial product or service

Best For

Audio generation and transformation from natural language descriptionsCustomizable sound design for creative projectsComposing audio outputs by combining multiple instructionsProducing emergent and novel sonic phenomena

Alternatives to Fugatto Nvidia

FAQ

What is Fugatto?
Fugatto is a foundational generative audio transformer model from NVIDIA Research that can synthesize and transform audio based on free-form text instructions, optionally using audio inputs as context.
How does Fugatto achieve compositional control?
Fugatto uses ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance, allowing instructions to be combined, interpolated, or negated for customizable outputs.
Can Fugatto generate novel sounds?
Yes, the framework can synthesize emergent sounds—sonic phenomena that go beyond conventional audio generation—enabling new creative possibilities.
Is Fugatto available for public use?
The website describes a research publication (ICLR 2025) with a demo website. No commercial availability or API is mentioned on this page.