IP Library Granted Patent US 12,560,670
Granted Patent B2
US 12,560,670 · App. 18/276,860 · Granted Feb 24, 2026

Model learning device, direction of arrival estimation device, model learning method, direction of arrival estimation method, and program

Inventor: Masahiro Yasuda (Tokyo, JP)
Assignee: NTT, Inc.
G01S3/8083G10K11/1752
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,560,670
App. No.
18/276,860
Granted
Feb 24, 2026
Kind
B2
Abstract

A method and a device for estimating a direction-of-arrival of a sound source and for learning a model are described. Operations for the estimating comprises estimating a reverberation component of an acoustic intensity vector, extracting an angle mask, estimating a time-frequency mask for noise suppression and sound source separation. A first direction-of-arrival of a sound source is estimated by applying the time-frequency mask to the acoustic intensity vector. A second direction-of-arrival of the sound source is estimated by applying the angle mask to the acoustic intensity vector. A cost function of a model is calculated based on the first and second direction-of-arrivals of the sound source and a label. Embodiments further describe updating a parameter of the model based on the cost function for learning the model.

Claims (28)

1 . A model learning device comprising:

processing circuitry configured to:

receive a real number spectrogram extracted from a complex spectrogram of acoustic data having a label indicating a sound source direction-of-arrival (“DOA”) for each time when the sound source direction-of-arrival is known, and an acoustic intensity vector (“IV”) extracted from the complex spectrogram as inputs, and output a reverberation component of the estimated acoustic intensity vector;

receive the acoustic intensity vector as an input and extract, as an angle mask, a time-frequency mask for selecting a time-frequency bin having an azimuth angle larger than an azimuth angle derived in a state where noise suppression and sound source separation are not performed;

receive the real number spectrogram, the acoustic intensity vector from which the reverberation component has been subtracted, and the angle mask as inputs, and output a time-frequency mask for noise suppression and sound source separation;

derive a first sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the time-frequency mask to the acoustic intensity vector from which the reverberation component has been subtracted;

derive a second sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the angle mask to the acoustic intensity vector from which the reverberation component has been subtracted; and

calculate a cost function of a model on a basis of the first derived sound source direction-of-arrival, the second derived sound source direction-of-arrival, and the label, and update, based on the calculated cost function of the model, a parameter of the model as training of the model, wherein the model after being trained estimates a sound source direction-of-arrival as an online operation with accuracy, the model is configured as a hybrid of IV-based DOA estimation and a deep neural network (DNN)-based estimation in the online operation, and the DNN represents a recurrent neural network comprising a unidirectional recurrent layer using forward information up to a current time, thereby estimating the sound source direction-arrival with accuracy while excluding feedback according to acoustic data from a future time.

2 . A direction of arrival estimation device comprising:

processing circuitry configured to:

receive a real number spectrogram extracted from a complex spectrogram of acoustic data, and an acoustic intensity vector extracted from the complex spectrogram as inputs, and output a reverberation component of the estimated acoustic intensity vector;

receive the acoustic intensity vector as an input and extracts, as an angle mask, a time-frequency mask for selecting a time-frequency bin having an azimuth angle larger than an azimuth angle derived in a state where noise suppression and sound source separation are not performed;

receive the real number spectrogram, the acoustic intensity vector from which the reverberation component has been subtracted, and the angle mask as inputs, and output a time-frequency mask for noise suppression and sound source separation; and

derive, by a trained model, a sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the time-frequency mask to the acoustic intensity vector from which the reverberation component has been subtracted, wherein the trained model estimates a sound source direction-of-arrival as an online operation with accuracy, the model is configured as a hybrid of IV-based DOA estimation and a deep neural network (DNN)-based estimation in the online operation, and the DNN represents a recurrent neural network comprising a unidirectional recurrent layer using forward information up to a current time, thereby estimating the sound source direction-arrival with accuracy while excluding feedback according to acoustic data from a future time.

3 . A model learning method comprising:

a step of receiving a real number spectrogram extracted from a complex spectrogram of acoustic data having a label indicating a sound source direction-of-arrival for each time when the sound source direction-of-arrival is known, and an acoustic intensity vector extracted from the complex spectrogram as inputs, and outputting a reverberation component of the estimated acoustic intensity vector;

a step of receiving the acoustic intensity vector as an input and extracting, as an angle mask, a time-frequency mask for selecting a time-frequency bin having an azimuth angle larger than an azimuth angle derived in a state where noise suppression and sound source separation are not performed;

a step of receiving the real number spectrogram, the acoustic intensity vector from which the reverberation component has been subtracted, and the angle mask as inputs, and outputting a time-frequency mask for noise suppression and sound source separation;

a step of deriving a sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the time-frequency mask to the acoustic intensity vector from which the reverberation component has been subtracted;

a step of deriving a sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the angle mask to the acoustic intensity vector from which the reverberation component has been subtracted; and

a step of calculating a cost function of a model on a basis of the derived sound source direction-of-arrival and the label, and updating, based on the calculated cost function of the model, a parameter of the model as training of the model, wherein the model after being trained estimates a sound source direction-of-arrival as an online operation with accuracy, the model is configured as a hybrid of IV-based DOA estimation and a deep neural network (DNN)-based estimation in the online operation, and the DNN represents a recurrent neural network comprising a unidirectional recurrent layer using forward information up to a current time, thereby estimating the sound source direction-arrival with accuracy while excluding feedback according to acoustic data from a future time.

4 . A direction of arrival estimation method comprising:

a step of receiving a real number spectrogram extracted from a complex spectrogram of acoustic data, and an acoustic intensity vector extracted from the complex spectrogram as inputs, and outputting a reverberation component of the estimated acoustic intensity vector;

a step of receiving the acoustic intensity vector as an input and extracting, as an angle mask, a time-frequency mask for selecting a time-frequency bin having an azimuth angle larger than an azimuth angle derived in a state where noise suppression and sound source separation are not performed;

a step of receiving the real number spectrogram, the acoustic intensity vector from which the reverberation component has been subtracted, and the angle mask as inputs, and outputting a time-frequency mask for noise suppression and sound source separation; and

a step of deriving, by a trained model, a sound source direction-of-arrival on a basis of an acoustic intensity vector obtained by applying the time-frequency mask to the acoustic intensity vector from which the reverberation component has been subtracted, wherein the trained model estimates a sound source direction-of-arrival as an online operation with accuracy, the model is configured as a hybrid of IV-based DOA estimation and a deep neural network (DNN)-based estimation in the online operation, and the DNN represents a recurrent neural network comprising a unidirectional recurrent layer using forward information up to a current time, thereby estimating the sound source direction-arrival with accuracy while excluding feedback according to acoustic data from a future time.

5 . A computer-readable non-transitory recording medium storing a computer-executable program instructions that when executed by a processor cause a computer to execute operations of the processing circuitry as configured in the model learning device according to claim 1 .

6 . A computer-readable non-transitory recording medium storing a computer-executable program instructions that when executed by a processor cause a computer to execute operations of the processing circuitry as configured in the direction of arrival estimation device according to claim 2 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2023
From: YASUDA, MASAHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 064557/0197 →
Continuity (1)
Related Publication 20240118363A1 · Apr 11, 2024
References Cited (17)
“Masahiro Yasuda, Sound Event Localization Based on Sound Intensity Vector Refined By DNN-Based Denoising and Source Separation, Feb. 14, 2020, NTT Media Intelligence Laboratories” (Year: 2020). [cited by examiner]
Yasuda, Masahiro et al. (2020) “Sound Event Localization Based on Sound Intensity Vector Refined By DNN-Based Denoising and Source Separation” ICASSP 2020, pp. 651-655. [cited by applicant]
Adavanne et al. (2019) “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks” IEEE Journal of Selected Topics in Signal Processing, vol. 13, No. 1, Mar. 2019, pp. 3… [cited by applicant]
Xu et al. (2017) “Surrey-CVSSP System for DCASE2017 Challenge TASK4” in Detection and Classification of Acoustic Scenes and Events, Nov. 16, 2017, Munich, Germany. [cited by applicant]
Lee et al. (2017) “Ensemble of convolutional neural networks for weakly-supervised sound event detection using multiple scale input”, in Tech. report of Detection and Classification of Acoustic Scenes and Events 2017 (D… [cited by applicant]
Chang et al. (2018) “Feature Extracted DOA Estimation Algorithm Using Acoustic Array for Drone Surveillance”, in Proc. of IEEE 87th Vehicular Technology Conference. [cited by applicant]
Adavanne et al. (2018) “Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network”, in Proc. of IEEE 26th European Signal Processing Conference. [cited by applicant]
Kapka et al. (2019) “Sound source detection, localization and classification using consecutive ensemble of CMN models”, inTech. report of Detection and Classification of Acoustic Scenes and Events 2019 (DCASE) Challange. [cited by applicant]
Cao et al. (2019) “Two-stage sound event localization and detection using intensity vector and generalized cross-correlation”, in Tech. report of Detection and Classification of Acoustic Scenes and Events 2019 (DCASE) C… [cited by applicant]
Noh et al. (2019) “Three-stage approach for sound event localization and detection”, in Tech. report of Detection and Classification of Acoustic Scenes and Events 2019 (DCASE) Challange. [cited by applicant]
Nguyen et a. (2019) DCASE 2019 task 3: A two-step system for sound event localization and detection, in Tech. report of Detection and Classification of Acoustic Scenes and Events 2019 (DCASE) Challange. [cited by applicant]
Schmidt (1986) “Multiple emitter location and signal parameter estimation”, IEEE Transactions on Antennas and propagation, vol. 34, pp. 276-280. [cited by applicant]
Ahonen et al. (2007)“Teleconference application and B-format microphone array for directional audio coding”, in Proc. of AES 30th International Conference: Intelligent Audio Environments. [cited by applicant]
Kitic et al. (2018) “Tramp: Tracking by a realtime ambisonic-based particle filter”, in Proc. of LOCATA Challenge Workshop, a satellite event of IWAENC. [cited by applicant]
Jarrett et al. (2010) “3d source localization in the spherical harmonic domain using a pseudointensity vector”, in Proc. of European Signal Processing Conference. [cited by applicant]
“DCASE2019 Workshop Workshop on Detection and Classification of Acoustic Scenes and Events,” [online], Oct. 25-26, 2019, [Searched on Feb. 8, 2021], Internet URL:http:/dcase.community/workshop2019/. [cited by applicant]
Yilmaz et al. (2004) “Blind separation of speech mixtures via time-frequency masking”, IEEE Trans. Signal Process., vol. 52, pp. 1830-1847, Jul. 2004. [cited by applicant]