Sketch2Sound Adobe logo

Sketch2Sound Adobe

Paid

Controllable audio generation via time-varying signals and sonic imitations

4.5
Inputs: audio, textOutputs: audio
Type
Saas
Company
Adobe

About Sketch2Sound Adobe

Sketch2Sound is a generative audio model developed by Adobe Research and Northwestern University that can create high-quality sounds from interpretable time-varying control signals—loudness, brightness, and pitch—combined with text prompts. The model can also synthesize arbitrary sounds from sonic imitations such as vocal imitations or reference sound-shapes. Built on top of any text-to-audio latent diffusion transformer (DiT), Sketch2Sound is lightweight, requiring only about 40,000 steps of fine-tuning and a single linear layer per control signal, making it more efficient than methods like ControlNet. By applying random median filters to control signals during training, the model allows users to prompt it with controls that have flexible levels of temporal specificity. Sketch2Sound enables sound artists to blend the semantic flexibility of text prompts with the expressivity and precision of sonic gestures, making it particularly useful for creating sound effects synced to video through vocal imitations.

Key Features

Generates high-quality sounds from time-varying control signals (loudness, brightness, pitch) and text prompts
Synthesizes arbitrary sounds from sonic imitations (vocal imitations or reference sound-shapes)
Implemented on top of any text-to-audio latent diffusion transformer (DiT)
Lightweight: requires only ~40k steps of fine-tuning and a single linear layer per control signal
Uses random median filters during training for flexible temporal specificity of controls
Combines semantic flexibility of text prompts with expressivity of sonic gestures
Designed for sound artists to create sounds synced to video via vocal imitations

Pros & Cons

Pros
  • Lightweight compared to ControlNet – only 40k fine-tuning steps and a single linear layer per control
  • Allows flexible temporal specificity via random median filters on control signals
  • Enables control over three fundamental audio attributes (loudness, brightness, pitch) in addition to text
  • Can synthesize sounds from vocal imitations, making it accessible for non-technical users
  • Maintains audio quality comparable to text-only baselines while following control signal gist
Cons
  • Requires a sonic imitation (vocal or reference sound) as input – not purely text-to-audio
  • Control signals are limited to loudness, brightness, and pitch; may not capture all desired sound attributes
  • Currently a research model from an academic paper – not yet available as a standalone commercial product
  • May require technical expertise to set up and fine-tune the model for custom use

Best For

Creating sound effects for video through vocal imitationsSound design where precise temporal control (loudness, brightness, pitch) is needed alongside text promptsGenerating ambient sounds (e.g., forest ambience) where control signal bursts translate to specific events (e.g., bird chirps)Rhythmic sound generation (e.g., bass drum and snare drum patterns) using pitch and unpitched regions from controls

Alternatives to Sketch2Sound Adobe

FAQ

What inputs does Sketch2Sound accept?
Sketch2Sound can take either a set of time-varying control signals (loudness, brightness, pitch) and a text prompt, or a sonic imitation (vocal imitation or reference sound-shape) from which the control signals are extracted.
How does Sketch2Sound work?
It extracts three control signals from an input sonic imitation: loudness, spectral centroid (brightness), and pitch probabilities. These signals are encoded and added to the latents of a text-to-audio latent diffusion transformer (DiT). During training, random median filters are applied to the control signals to allow flexible temporal specificity.
What makes Sketch2Sound different from other text-to-audio models?
Sketch2Sound adds direct control over loudness, brightness, and pitch via time-varying signals, in addition to text prompts. It is also lightweight, requiring only about 40k fine-tuning steps and a single linear layer per control, unlike heavier methods like ControlNet.
Can Sketch2Sound generate sounds synchronized to video?
Yes, the model is designed to create sound effects that are synced to video through vocal imitations. The examples in the demo video demonstrate this capability.