Perceptual-based 3D sound localization is the application of knowledge of the human auditory system to develop 3D sound localization technology.
Motivation and applications Human listeners combine information from two ears to localize and separate sound sources originating in different locations in a process called binaural hearing. The powerful signal processing methods found in the neural systems and brains of humans and other animals are flexible, environmentally adaptable, and take place rapidly and seemingly without effort. Emulating the mechanisms of binaural hearing can improve recognition accuracy and signal separation in DSP algorithms, especially in noisy environments. Furthermore, by understanding and exploiting biological mechanisms of sound localization, virtual sound scenes may be rendered with more perceptually relevant methods, allowing listeners to accurately perceive the locations of auditory events. One way to obtain the perceptual-based sound localization is from the sparse approximations of the anthropometric features. Perceptual-based sound localization may be used to enhance and supplement robotic navigation and environment recognition capability. In addition, it is also used to create virtual auditory spaces which is widely implemented in hearing aids.
Problem statement and basic concepts While the relationship between human perception of sound and various attributes of the sound field is not yet well understood, DSP algorithms for sound localization are able to employ several mechanisms found in neural systems, including the interaural time difference (ITD, the difference in arrival time of a sound between two locations), the interaural intensity difference (IID, the difference in intensity of a sound between two locations), artificial pinnae, the precedence effect, and head-related transfer functions (HRTF). When localizing 3D sound in spatial domain, one could take into account that the incoming sound signal could be reflected, diffracted and scattered by the upper torso of the human which consists of shoulders, head and pinnae. Localization also depends on the direction of the sound source.
HATS: Head and Torso Simulator
Brüel's & Kjær's Head And Torso Simulator (HATS) is a mannequin prototype with built-in ear and mouth simulators that provides a realistic reproduction of the acoustic properties of an average adult human head and torso. It is designed to be used in electro-acoustics tests, for example, headsets, audio conference devices, microphones, headphones and hearing aids. Various existing approaches are based on this structural model.
Existing approaches
Particle based tracking It is essential to be able to analyze the distance and intensity of various sources in a spatial domain. We can track each such sound source, by using a probabilistic temporal integration, based on data obtained through a microphone array and a particle filtering tracker. Using this approach, the Probability Density Function (PDF) representing the location of each source is represented as a set of particles to which different weights (probabilities) are assigned. The choice of particle filtering over Kalman filtering is further justified by the non-gaussian probabilities arising from false detections and multiple sources.
… excerpt ends here. Continue reading the full article.

