IP Library › Granted Patent US 12,254,250
Granted Patent B2
US 12,254,250 · App. 17/270,448 · Granted Mar 18, 2025

Mask estimation device, mask estimation method, and mask estimation program

Inventors: Tomohiro Nakatani (Musashino, JP); Marc Delcroix (Musashino, JP); Keisuke Kinoshita (Musashino, JP); Nobutaka Ito (Musashino, JP); Shoko Araki (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F30/27G06N3/08G10L21/0216
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,250
App. No.
17/270,448
Granted
Mar 18, 2025
Kind
B2
Abstract

A mask estimation apparatus includes processing circuitry configured to estimate, for a target segment to be processed among a plurality of segments of a continuous time, a first mask which is an occupancy ratio of a target signal to an observation signal of the target segment, based on a first feature obtained from a plurality of the observation signals of the target segment recorded at a plurality of locations, and estimate a parameter for modeling a second feature and a second mask which is an occupancy ratio of the target signal to the observation signal based on an estimation result of the first mask in the target segment and the second feature obtained from the plurality of the observation signals of the target segment.

Claims (36)

1. A mask estimation apparatus comprising:

processing circuitry configured to sequentially perform for a plurality of segments of continuous time:

online estimation, for a target segment to be processed among the plurality of segments of continuous time, of a first mask which is an occupancy ratio of a target signal to an observation signal of the target segment, based on a first feature obtained from a plurality of the observation signals of the target segment recorded at a plurality of locations, the occupancy ratio indicating a probability of a target sound occupying each time frequency point of the observation signal; and

online estimation of a parameter for modeling a second feature and a second mask which is an occupancy ratio of the target signal to the observation signal based on an estimation result of the first mask in the target segment and the second feature obtained from the plurality of the observation signals of the target segment, wherein the processing circuitry is further configured to:

update the parameter based on a cumulative sum of a plurality of first masks up to the target segment, and the second feature of the target segment and the second mask,

update the second mask based on the second feature of the target segment, the first mask, and the parameter, and

repeatedly perform the update of the parameter and the second mask until a predetermined convergence condition is satisfied,

wherein the cumulative sum of the plurality of the first masks indicates an amount of target signals included in the observation signal up to the target segment which is a segment for which mask estimation is to be performed,

the cumulative sum of the plurality of first masks indicates a cumulative sum of the probability that a signal of the target sound occupies the observation signal, and

the processing circuitry is further configured to determine whether the cumulative sum of the probability up to the time of the target segment exceeds a predetermined threshold value.

2. The mask estimation apparatus according to claim 1 , wherein the processing circuitry is further configured to:

estimate the first mask using a neural network, and

estimate the second mask based on the first mask and a distribution model of the second feature with the parameter as a condition.

3. The mask estimation apparatus according to claim 1 , wherein the processing circuitry is further configured to:

estimate the second mask based on the estimated parameter in a case where the amount of the plurality of the target signals exceeds the threshold value, and

substitute the estimation result of the first mask for an estimated value of the second mask in a case where the amount of the plurality of the target signals does not exceed the threshold value.

4. A mask estimation method comprising:

sequentially performing for a plurality of segments of continuous time:

online estimation, for a target segment to be processed among the plurality of segments of continuous time, of a first mask which is an occupancy ratio of a target signal to an observation signal of the target segment, based on a first feature obtained from a plurality of the observation signals of the target segment recorded at a plurality of locations, the occupancy ratio indicating a probability of a target sound occupying each time frequency point of the observation signal; and

online estimation of a parameter for modeling a second feature and a second mask which is an occupancy ratio of the target signal to the observation signal based on an estimation result of the first mask in the target segment and the second feature obtained from the plurality of the observation signals of the target segment, by processing circuitry, wherein the method further comprises:

updating the parameter based on a cumulative sum of a plurality of first masks up to the target segment, and the second feature of the target segment and the second mask,

updating the second mask based on the second feature of the target segment, the first mask, and the parameter, and

repeatedly performing the update of the parameter and the second mask until a predetermined convergence condition is satisfied,

wherein the cumulative sum of the plurality of the first masks indicates an amount of target signals included in the observation signal up to the target segment which is a segment for which mask estimation is to be performed,

the cumulative sum of the plurality of first masks indicates a cumulative sum of the probability that a signal of the target sound occupies the observation signal, and

the method further includes determining whether the cumulative sum of the probability up to the time of the target segment exceeds a predetermined threshold value.

5. A non-transitory computer-readable recording medium storing therein a mask estimation program that causes a computer to execute a process comprising:

sequentially performing for a plurality of segments of continuous time:

online estimation, for a target segment to be processed among the plurality of segments of continuous time, of a first mask which is an occupancy ratio of a target signal to an observation signal of the target segment, based on a first feature obtained from a plurality of the observation signals of the target segment recorded at a plurality of locations, the occupancy ratio indicating a probability of a target sound occupying each time frequency point of the observation signal; and

online estimation of a parameter for modeling a second feature and a second mask which is an occupancy ratio of the target signal to the observation signal based on an estimation result of the first mask in the target segment and the second feature obtained from the plurality of the observation signals of the target segment, wherein the process further comprises:

updating the parameter based on a cumulative sum of a plurality of first masks up to the target segment, and the second feature of the target segment and the second mask,

updating the second mask based on the second feature of the target segment, the first mask, and the parameter, and

repeatedly performing the update of the parameter and the second mask until a predetermined convergence condition is satisfied,

wherein the cumulative sum of the plurality of the first masks indicates an amount of target signals included in the observation signal up to the target segment which is a segment for which mask estimation is to be performed,

the cumulative sum of the plurality of first masks indicates a cumulative sum of the probability that a signal of the target sound occupies the observation signal, and

the process further includes determining whether the cumulative sum of the probability up to the time of the target segment exceeds a predetermined threshold value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: NAKATANI, TOMOHIRO; DELCROIX, MARC; KINOSHITA, KEISUKE; ITO, NOBUTAKA; ARAKI, SHOKO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 055364/0800 →
Priority Claims (1)
JP 2018-163856 · Aug 31, 2018 · national
Continuity (1)
Related Publication 20210216687A1 · Jul 15, 2021
References Cited (12)
US 10553236B1 · Ayrapetian · 2020 [cited by examiner]
US 10643633B2 · Nakatani · 2020 [cited by examiner]
US 20080215651A1 · Sawada · 2008 [cited by examiner]
US 20190267019A1 · Ito · 2019 [cited by examiner]
Nakatani T, Ito N, Higuchi T, Araki S, Kinoshita K. Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming. In2017 IEEE International Conference on Acoustics, Speech and Signal Pr… [cited by examiner]
Liu Y, Ganguly A, Kamath K, Kristjansson T. Neural network based time-frequency masking and steering vector estimation for two-channel MVDR beamforming. In2018 IEEE International Conference on Acoustics, Speech and Sign… [cited by examiner]
Xiao X, Zhao S, Jones DL, Chng ES, Li H. On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition. In2017 IEEE International Conference on Acoustics, Speech and Signal Process… [cited by examiner]
Yu, Y., Wang, W. & Han, P. Localization based stereo speech source separation using probabilistic time-frequency masking and deep neural networks. J Audio Speech Music Proc. 2016, 7 (2016). (Year: 2016). [cited by examiner]
Higuchi, Takuya, Nobutaka Ito, Takuya Yoshioka, and Tomohiro Nakatani. “Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise.” In 2016 IEEE International Conference on Acoustics, Speech and… [cited by examiner]
Higuchi T, Ito N, Yoshioka T, Nakatani T. Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise. In2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) Mar.… [cited by examiner]
Higuchi, T., Kinoshita, K., Ito, N., Karita, S. and Nakatani, T., Apr. 2018. Frame-by-frame closed-form update for mask-based adaptive MVDR beamforming. In 2018 IEEE International Conference on Acoustics, Speech and Sig… [cited by examiner]
Nakatani et al., “Integrating DNN-Based and Spatial Clustering-Based Mask Estimation for Robust MVDR Beamforming”, IEEE, ICASSP, 2017, pp. 286-290. [cited by applicant]