Acoustic Space Mapping: A Machine Learning Approach to Sound Source Separation and Localization

Antoine Deleforge 1
1 PERCEPTION - Interpretation and Modelling of Images and Videos
Inria Grenoble - Rhône-Alpes, LJK - Laboratoire Jean Kuntzmann, INPG - Institut National Polytechnique de Grenoble
Abstract : In this thesis, we address the long-studied problem of binaural (two microphones) sound source separation and localization through supervised learning. To achieve this, we develop a new paradigm referred to as acoustic space mapping, at the crossroads of binaural perception, robot hearing, audio signal processing and machine learning. The proposed approach consists in learning a link between auditory cues perceived by the system and the emitting sound source position in another modality of the system, such as the visual space or the motor space. We propose new experimental protocols to automatically gather large training sets that associate such data. Obtained datasets are then used to reveal some fundamental intrinsic properties of acoustic spaces and lead to the development of a general family of probabilistic models for locally-linear high- to low-dimensional space mapping. We show that these models unify several existing regression and dimensionality reduction techniques, while encompassing a large number of new models that generalize previous ones. The properties and inference of these models are thoroughly detailed, and the prominent advantage of proposed methods with respect to state-of-the-art techniques is established on different space mapping applications, beyond the scope of auditory scene analysis. We then show how the proposed methods can be probabilistically extended to tackle the long-known cocktail party problem, i.e., accurately localizing one or several sound sources emitting at the same time in a real-word environment, and separate the mixed signals. We show that resulting techniques perform these tasks with an unequaled accuracy. This demonstrates the important role of learning and puts forwards the acoustic space mapping paradigm as a promising tool for robustly addressing the most challenging problems in computational binaural audition.
Complete list of metadatas

Cited literature [107 references]  Display  Hide  Download

https://tel.archives-ouvertes.fr/tel-00913965
Contributor : Team Perception <>
Submitted on : Wednesday, December 4, 2013 - 4:16:04 PM
Last modification on : Wednesday, April 11, 2018 - 1:59:11 AM
Long-term archiving on: Saturday, April 8, 2017 - 4:00:19 AM

Identifiers

  • HAL Id : tel-00913965, version 1

Collections

Citation

Antoine Deleforge. Acoustic Space Mapping: A Machine Learning Approach to Sound Source Separation and Localization. Machine Learning [cs.LG]. Université de Grenoble, 2013. English. ⟨tel-00913965⟩

Share

Metrics

Record views

1686

Files downloads

6218