Audio and Speech
23 June 2026
Meta Unveils SAM Audio: A Model for On-Demand Sound Separation
N
New-ZZZ desk
AI at Meta Blog · 1 month ago
The company introduced SAM Audio—the first unified multimodal model for sound separation. This model allows users to easily isolate any sound from a complex audio recording using intuitive prompts, whether text, visual cues, or timestamps. At the core of SAM Audio is the PE-AV technical engine, which ensures advanced performance, making audio segmentation as accessible as object segmentation in images.
Why it matters
- —This is the first unified approach to audio segmentation that works using multimodal prompts (text, video, time).
- —The model allows users to isolate specific sounds (e.g., a guitar or traffic noise) from a complex recording with a single click, mimicking real-life scenarios.
- —The company also introduced new tools, including the SAM Audio-Bench benchmark and the automatic SAM Audio Judge.
Key facts
- SAM Audio is a unified multimodal model for sound separation.
- The model uses intuitive prompts: text, visual cues, or timestamps.
- It is powered by the PE-AV technical engine, ensuring high performance.
- New tools were presented: the SAM Audio-Bench benchmark and the automatic SAM Audio Judge.
The full text is in the original source. Here we provide a brief summary and key facts.