Home Knowledge Base Sound Source Localization

Sound Source Localization is the multimodal task of identifying the spatial location in a visual scene that corresponds to an observed sound — using audio-visual correlation to generate heatmaps or bounding boxes over video frames that pinpoint where a sound is originating from, such as localizing a speaking person, a playing instrument, or a barking dog by jointly analyzing audio spectral features and visual motion patterns.

What Is Sound Source Localization?

Why Sound Source Localization Matters

Sound Source Localization Methods

MethodSupervisionLocalization OutputTraining DataKey Innovation
Attention & ActivateSelf-supervisedHeatmapUnlabeled videoAV attention maps
LVSContrastiveHeatmapUnlabeled videoHard negatives
Mix-and-LocalizeSelf-supervisedPer-source heatmapMixed audioSource separation
EZ-VSLSelf-supervisedHeatmapUnlabeled videoPseudo-labels
SLAVCSelf-supervisedHeatmap + segmentsUnlabeled videoSemantic grouping

Sound source localization is the spatial grounding task of audio-visual AI — pinpointing where sounds originate in visual scenes through learned cross-modal correlations between audio spectral features and visual spatial features, enabling applications from robotics and surveillance to augmented reality that require machines to understand the spatial relationship between what they see and what they hear.

sound source localizationmultimodal ai

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.