IP Library Granted Patent US 12,451,112
Granted Patent B2
US 12,451,112 · App. 18/571,765 · Granted Oct 21, 2025

Acoustic signal enhancement device, acoustic signal enhancement method, and program

Inventors: Tomohiro Nakatani (Tokyo, JP); Rintaro Ikeshita (Tokyo, JP); Keisuke Kinoshita (Tokyo, JP); Hiroshi Sawada (Tokyo, JP); Naoyuki Kamo (Tokyo, JP); Shoko Araki (Tokyo, JP)
Assignee: NTT, Inc.
G10K11/17821G10K11/17881H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,112
App. No.
18/571,765
Granted
Oct 21, 2025
Kind
B2
Abstract

There is provided an acoustic signal enhancement device that receives, as an input, a recording sound obtained by frequency division and updates parameters, the device including: assuming that a switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes, a beamformer unit that performs beamformer processing based on a weighted spatial covariance matrix which is updated and updates an auxiliary estimation value of a target sound; a switch unit that updates the switch weight and power of a target sound based on the updated auxiliary estimation value and outputs an estimation value of the target sound; and a weighted spatial covariance estimation unit that updates the weighted spatial covariance matrix based on the updated switch weight and the power.

Claims (49)

1 . An acoustic signal enhancement device that receives, as an input, a recording sound obtained by frequency division and updates parameters, the acoustic signal enhancement device comprising:

processing circuitry configured to:

assuming that a switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes,

perform beamformer processing based on a weighted spatial covariance matrix which is updated and update an auxiliary estimation value of a target sound;

update the switch weight and power of a target sound based on the updated auxiliary estimation value and output an estimation value of the target sound; and

update the weighted spatial covariance matrix based on the updated switch weight and the power.

2 . An acoustic signal enhancement device that receives, as an input, a recording sound obtained by frequency division and updates parameters, the acoustic signal enhancement device comprising:

processing circuitry configured to:

assuming that a first switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes, and

assuming that a second switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial-temporal states where a recording sound temporally changes,

perform reverberation suppression processing on the recording sound based on a weighted spatial-temporal covariance matrix which is updated and update an auxiliary reverberation-suppressed sound of a target sound;

update the second switch weight based on the auxiliary reverberation-suppressed sound, updated power of the target sound, and an updated beamformer coefficient;

update an estimation value of the target sound, the beamformer coefficient, the power of the target sound, and the first switch weight of the target sound based on at least one of the auxiliary reverberation-suppressed sounds; and

update the weighted spatial-temporal covariance matrix based on the first switch weight, the second switch weight, and the power.

3 . The acoustic signal enhancement device according to claim 2 ,

wherein processing circuitry configured to:

perform beamformer processing based on a weighted spatial covariance matrix which is updated and update an auxiliary estimation value of the target sound;

update the first switch weight and power of the target sound based on the updated auxiliary estimation value and output the estimation value of the target sound; and

update the weighted spatial covariance matrix based on the updated first switch weight and the power.

4 . An acoustic signal enhancement device that receives, as inputs, recording sounds from a plurality of microphones, the acoustic signal enhancement device comprising:

processing circuitry configured to,

assuming that a first switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes, and

assuming that a second switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial-temporal states where a recording sound temporally changes,

update a weighted spatial covariance matrix for estimating a coefficient for obtaining a target sound of a beamformer based on the first and second switch weights, power of each sound source, and an auxiliary reverberation-suppressed sound of each sound source;

update the coefficient of the beamformer which estimates a separation sound of a separation matrix based on the weighted spatial covariance matrix and update an auxiliary estimation value of each sound source based on the updated coefficient of the beamformer and the auxiliary reverberation-suppressed sound; and

update estimation values of all the sound sources based on the first and second switch weights, update power of each sound source based on the estimation values of all the sound sources, and update the first switch weight based on the power of each sound source.

5 . The acoustic signal enhancement device according to claim 4 , further comprising:

processing circuitry configured to:

update a weighted spatial-temporal covariance matrix for estimating a filter coefficient of reverberation suppression processing based on the first and second switch weights and the power of each sound source; and

update the filter coefficient of reverberation suppression processing based on the coefficient of the beamformer and the weighted spatial-temporal covariance matrix and update the auxiliary reverberation-suppressed sound,

wherein processing circuitry configured to

update the second switch weight in addition to the first switch weight based on the power of each sound source.

6 . An acoustic signal enhancement method executed by an acoustic signal enhancement device that receives, as an input, a recording sound obtained by frequency division and updates parameters, the acoustic signal enhancement method comprising:

assuming that a switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes,

a beamformer step of performing beamformer processing based on a weighted spatial covariance matrix which is updated and updating an auxiliary estimation value of a target sound;

a switch step of updating the switch weight and power of a target sound based on the updated auxiliary estimation value and outputting an estimation value of the target sound; and

a weighted spatial covariance estimation step of updating the weighted spatial covariance matrix based on the updated switch weight and the power.

7 . An acoustic signal enhancement method executed by an acoustic signal enhancement device that receives, as an input, a recording sound obtained by frequency division and updates parameters, the acoustic signal enhancement method comprising:

assuming that a first switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial states where a recording sound temporally changes, and

assuming that a second switch weight is a weight indicating a ratio of a classification to which a recording sound at each timing belongs in classifications of spatial-temporal states where a recording sound temporally changes,

a reverberation suppression step of performing reverberation suppression processing on the recording sound, performing beamformer processing based on a weighted spatial-temporal covariance matrix which is updated, and updating an auxiliary reverberation-suppressed sound of a target sound;

a switch step of updating the second switch weight based on the auxiliary reverberation-suppressed sound, updated power of the target sound, and an updated beamformer coefficient;

a switching beamformer step of updating an estimation value of the target sound, the beamformer coefficient, the power of the target sound, and the first switch weight of the target sound based on at least one of the auxiliary reverberation-suppressed sounds; and

a weighted spatial-temporal covariance estimation step of updating the weighted spatial-temporal covariance matrix based on the first switch weight, the second switch weight, and the power.

8 . A program causing a computer to function as the acoustic signal enhancement device according to claim 1 .

9 . A program causing a computer to function as the acoustic signal enhancement device according to claim 2 .

10 . A program causing a computer to function as the acoustic signal enhancement device according to claim 3 .

11 . A program causing a computer to function as the acoustic signal enhancement device according to claim 4 .

12 . A program causing a computer to function as the acoustic signal enhancement device according to claim 5 .

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0597 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2023
From: NAKATANI, TOMOHIRO; IKESHITA, RINTARO; KINOSHITA, KEISUKE; SAWADA, HIROSHI; KAMO, NAOYUKI; ARAKI, SHOKO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 065905/0082 →
Continuity (1)
Related Publication 20240312446A1 · Sep 19, 2024
References Cited (14)
US 20110044462A1 · Yoshioka · 2011 [cited by examiner]
US 20140056435A1 · Kjems · 2014 [cited by examiner]
US 20180061432A1 · Taniguchi · 2018 [cited by examiner]
US 20220068288A1 · Nakatani · 2022 [cited by examiner]
JP 2015135437A · 2015 [cited by examiner]
Ikeshita et al. “Independent Vector Extraction for Joint Blind Source Separation and Dereverberation” arXiv <URL: https://arxiv.org/abs/2102.04696v1> Feb. 9, 2021. [cited by applicant]
Ikeshita et al. “Independent Vector Extraction for Fast Joint Blind Source Separation and Dereverberation” arXiv <URL: https://arxiv.org/abs/2102.04696v2> Apr. 22, 2021. [cited by applicant]
Nakatani et al.“Switching Convolutional Beamformer” Eusipco 2021 <URL: https://eusipco2021-virtual.org> Aug. 16, 2021. [cited by applicant]
Nakatani et al. “Computationally Efficient and Versatile Framework for Joint Optimization of Blind Speech Separation and Dereverberation” Interspeech 2020 <URL: http://www.interspeech2020.org/uploadfile/pdf/Mon-1-2-9.pd… [cited by applicant]
Nakatani et al. “Improved Switching Convolutional Beamformer.” Acoustical Science and Technology—Journal, Sep. 2021. [cited by applicant]
Ikeshita et al. “Blind Signal Dereverberation Based on Mixture of Weighted Prediction Error Models” IEEE Signal Processing Letters, vol. 28, Feb. 2, 2021 p. 399-403. [cited by applicant]
Ikeshita et al. “Independent Vector Extraction for Fast Joint Blind Source Separation and Dereverberation” IEEE Signal Processing Letters, vol. 28, Apr. 20, 2021 p. 972-976. [cited by applicant]
Yamaoka et al. “Time-Frequency-Bin-Wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined Situations,” Proc. IEEE ICASSP, pp. 7908-7912, 2019. [cited by applicant]
Nakatani et al. “Jointly optimal denoising, dereverberation, and source separation,” IEEE/ACM Trans. Audio, Speech, and Language Processing, vol. 28, pp. 2267-2282, 2020. [cited by applicant]