IP Library › Granted Patent US 12,431,158
Granted Patent B2
US 12,431,158 · App. 17/635,354 · Granted Sep 30, 2025

Speech signal processing device, speech signal processing method, speech signal processing program, training device, training method, and training program

Inventors: Hiroshi Sato (Tokyo, JP); Tsubasa Ochiai (Tokyo, JP); Keisuke Kinoshita (Tokyo, JP); Marc Delcroix (Tokyo, JP); Tomohiro Nakatani (Tokyo, JP); Atsunori Ogawa (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L25/30G10L19/008G10L21/0272
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,158
App. No.
17/635,354
Granted
Sep 30, 2025
Kind
B2
Abstract

An audio signal processing apparatus ( 10 ) includes a first auxiliary feature conversion unit ( 12 ) and a second auxiliary feature conversion unit ( 13 ) that convert a plurality of signals relating to processing of an audio signal of a target speaker into a plurality of auxiliary features for the plurality of signals using a plurality of auxiliary neural networks corresponding to the plurality of signals, and an audio signal processing unit ( 11 ) that estimates information regarding an audio signal of the target speaker included in a mixed audio signal using a main neural network based on an input feature of the mixed audio signal and the plurality of auxiliary features, wherein the plurality of signals relating to processing of the audio signal of the target speaker are two or more pieces of information of different modalities.

Claims (34)

1. A training apparatus comprising:

a selection unit configured to select a mixed audio signal for training and a plurality of signals relating to processing of an audio signal of a target speaker for training from training data;

a feature conversion unit configured to convert the plurality of signals relating to the processing of the audio signal of the target speaker for training into a plurality of auxiliary features for the plurality of signals using a plurality of auxiliary neural networks corresponding to the plurality of signals;

an audio signal processing unit configured to estimate information regarding processing of an audio signal of the target speaker included in the mixed audio signal for training using a main neural network based on a feature of the mixed audio signal for training and the plurality of auxiliary features; and

an update unit configured to update parameters of neural networks and cause the selection unit, the feature conversion unit, and the audio signal processing unit to repeatedly execute processing until a predetermined criterion is satisfied to set the parameters of the neural networks satisfying the predetermined criterion, wherein the plurality of signals relating to processing of the audio signal of the target speaker are two or more pieces of information of different modalities,

wherein the training apparatus further comprising:

an auxiliary information generation unit configured to generate a weighted sum of the plurality of auxiliary features multiplied by attentions corresponding to the plurality of auxiliary features using a neural network, wherein the audio signal processing unit is configured to receive as an input a second intermediate feature generated by integrating a first intermediate feature obtained by converting the mixed audio signal using a first main neural network included in the main neural network, and the weighted sum and estimate information regarding the audio signal of the target speaker included in the mixed audio signal for training using a second main neural network included in the main neural network, and the auxiliary information generation unit includes:

an attention calculation unit configured to calculate attentions corresponding to the plurality of auxiliary features based on the first intermediate feature and the plurality of auxiliary features; and

an aggregation unit configured to calculate the weighted sum of the plurality of auxiliary features multiplied by the attentions corresponding to the plurality of auxiliary features calculated by the attention calculation unit.

2. The training apparatus according to claim 1 , wherein

the selection unit is configured to select the mixed audio signal for training, the audio signal of the target speaker for training, and video information of speakers at a time of recording the mixed audio signal for training from the training data,

the feature conversion unit includes:

a first auxiliary feature conversion unit configured to convert the audio signal of the target speaker into a first auxiliary feature using a first auxiliary neural network; and

a second auxiliary feature conversion unit configured to convert the video information of the speakers at the time of recording the mixed audio signal for training into a second auxiliary feature using a second auxiliary neural network,

the audio signal processing unit is configured to estimate information regarding the audio signal of the target speaker included in the mixed audio signal for training using the main neural network based on the feature of the mixed audio signal for training, the first auxiliary feature, and the second auxiliary feature, and

the update unit is configured to update parameters of neural networks and cause the selection unit, the first auxiliary feature conversion unit, the second auxiliary feature conversion unit, and the audio signal processing unit to repeatedly execute processing until the predetermined criterion is satisfied to set the parameters of the neural networks satisfying the predetermined criterion.

3. The training apparatus according to claim 2 , wherein the update unit is configured to update parameters of neural networks to allow a weighted sum of a first loss, with respect to a teacher signal, of audio of the target speaker included in the mixed audio signal for training where the audio signal processing unit is estimated using the feature of the mixed audio signal for training, the first auxiliary feature, and the second auxiliary feature, a second loss, with respect to a teacher signal, of audio of the target speaker included in the mixed audio signal for training where the audio signal processing unit is estimated based on the feature of the mixed audio signal for training and the first auxiliary feature, and a third loss, with respect to a teacher signal, of audio of the target speaker included in the mixed audio signal for training that is estimated based on the feature of the mixed audio signal for training and the second auxiliary feature to become smaller.

4. The training apparatus according to claim 1 , wherein

the auxiliary information generation unit further includes:

a normalization unit configured to normalize norms of the plurality of auxiliary features; and

a scaling unit configured to output the weighted sum multiplied by a scale factor calculated based on magnitudes of the norms before normalization to the audio signal processing unit, and

the aggregation unit is configured to calculate a weighted sum of the plurality of normalized auxiliary features multiplied by the attentions corresponding to the plurality of auxiliary features calculated by the attention calculation unit.

5. The training apparatus according to claim 4 , wherein

the audio signal processing unit is configured to estimate the audio signal of the target speaker included in the mixed audio signal for training, and

the update unit is configured to update parameters of neural networks to optimize an objective function based on attentions corresponding to the plurality of auxiliary features calculated by the attention calculation unit, preset desired values of attentions corresponding to the plurality of auxiliary features, the audio signal of the target speaker included in the mixed audio signal for training estimated by the audio signal processing unit, and a teacher signal of audio of the target speaker included in the mixed audio signal for training.

6. The training apparatus according to claim 4 , further comprising a prediction unit configured to predict reliabilities of a plurality of signals relating to processing of the audio signal of the target speaker for training using a neural network based on the plurality of auxiliary features, wherein

the audio signal processing unit is configured to estimate the audio signal of the target speaker included in the mixed audio signal for training, and

the update unit is configured to update parameters of neural networks to optimize an objective function based on the reliabilities of the plurality of signals relating to processing of the audio signal of the target speaker for training predicted by the prediction unit, predetermined reliabilities of the plurality of signals relating to processing of the audio signal of the target speaker for training, the audio signal of the target speaker included in the mixed audio signal for training estimated by the audio signal processing unit, and a teacher signal of audio of the target speaker included in the mixed audio signal for training.

7. The training apparatus according to claim 1 , wherein

the audio signal processing unit is configured to estimate the audio signal of the target speaker included in the mixed audio signal for training, and

the update unit is configured to update parameters of neural networks to optimize an objective function based on attentions corresponding to the plurality of auxiliary features calculated by the attention calculation unit, preset desired values of attentions corresponding to the plurality of auxiliary features, the audio signal of the target speaker included in the mixed audio signal for training estimated by the audio signal processing unit, and a teacher signal of audio of the target speaker included in the mixed audio signal for training.

8. The training apparatus according to claim 1 , further comprising a prediction unit configured to predict reliabilities of a plurality of signals relating to processing of the audio signal of the target speaker for training using a neural network based on the plurality of auxiliary features, wherein

the audio signal processing unit is configured to estimate the audio signal of the target speaker included in the mixed audio signal for training, and

the update unit is configured to update parameters of neural networks to optimize an objective function based on the reliabilities of the plurality of signals relating to processing of the audio signal of the target speaker for training predicted by the prediction unit, predetermined reliabilities of the plurality of signals relating to processing of the audio signal of the target speaker for training, the audio signal of the target speaker included in the mixed audio signal for training estimated by the audio signal processing unit, and a teacher signal of audio of the target speaker included in the mixed audio signal for training.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2023
From: SATO, HIROSHI; OCHIAI, TSUBASA; KINOSHITA, KEISUKE; DELCROIX, MARC; NAKATANI, TOMOHIRO; OGAWA, ATSUNORI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 062722/0846 →
Priority Claims (1)
WO PCT/JP2019/032193 · Aug 16, 2019 · international
Continuity (1)
Related Publication 20220335965A1 · Oct 20, 2022
References Cited (7)
US 20170330586A1 · Roblek · 2017 [cited by examiner]
US 20190080689A1 · Kagoshima · 2019 [cited by examiner]
US 20190311711A1 · Fazeldehkordi · 2019 [cited by examiner]
US 20200335121A1 · Mosseri · 2020 [cited by examiner]
Delcroix et al. (2018) “Single Channel Target Speaker Extraction and Recognition with Speaker Beam” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, pp. 5554-5558. [cited by applicant]
Ephrat et al. (2018) “Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation” ACM Trans. on Graphics, vol. 37, No. 4. [cited by applicant]
Ochiai et al. (2019) “Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues” Interspeech, Sep. 15, 2019. [cited by applicant]