IP Library Granted Patent US 12,354,620
Granted Patent B2
US 12,354,620 · App. 18/020,084 · Granted Jul 8, 2025

Signal processing device, signal processing method, signal processing program, learning device, learning method, and learning program

Inventors: Tsubasa Ochiai (Musashino, JP); Marc Delcroix (Musashino, JP); Yuma Koizumi (Musashino, JP); Hiroaki Ito (Musashino, JP); Keisuke Kinoshita (Musashino, JP); Shoko Araki (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L21/028G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,620
App. No.
18/020,084
Granted
Jul 8, 2025
Kind
B2
Abstract

A signal processing device includes processing circuitry configured to receive an input of extraction target information indicating which audio class of an audio signal is to be extracted from a mixture audio signal constituted by a mixture of audio signals of a plurality of audio classes, and output a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal, with a neural network by using a feature value of the mixture audio signal and the extraction target information.

Claims (33)

1. A signal processing device comprising:

processing circuitry configured to:

receive an input of extraction target information indicating which audio class of an audio signal is to be extracted from a mixture audio signal constituted by a mixture of audio signals of a plurality of audio classes;

integrate information of the mixture audio signal and the extraction target information; and

output a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with a neural network by using a result of the integration of the information of the mixture audio signal and the extraction target information.

2. The signal processing device according to claim 1 , wherein

the extraction target information is a target class vector indicating, by a vector, which audio class of the audio signal is to be extracted from the mixture audio signal,

the processing circuitry is further configured to

perform processing of embedding the target class vector by using a neural network, and

output a result of extracting the audio signal of the audio class indicated by the target class vector from the mixture audio signal with the neural network by using a feature value obtained by integrating a feature value of the mixture audio signal and the target class vector after the embedding processing.

3. The signal processing device according to claim 1 ,

wherein the processing circuitry is further configured to

receive an input of a target class vector indicating, by a vector, which audio class of the audio signal is to be removed from the mixture audio signal, and

output a result of removing the audio signal of the audio class indicated by the target class vector from the mixture audio signal with the neural network by using a feature value obtained by integrating the target class vector after an embedding processing to a feature value of the mixture audio signal.

4. A signal processing method executed by a signal processing device, the signal processing method comprising:

receiving an input of extraction target information indicating which audio class of an audio signal is to be extracted from a mixture audio signal constituted by a mixture of audio signals of a plurality of audio classes;

integrate information of the mixture audio signal and the extraction target information; and

outputting a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with a neural network by using a result of the integration of the information of the mixture audio signal and the extraction target information.

5. A non-transitory computer-readable recording medium storing therein a signal processing program that causes a computer to execute a process comprising:

receiving an input of extraction target information indicating which audio class of an audio signal is to be extracted from a mixture audio signal constituted by a mixture of audio signals of a plurality of audio classes;

integrating information of the mixture audio signal and the extraction target information; and

outputting a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with a neural network by using a result of the integration of the information of the mixture audio signal and the extraction target information.

6. The signal processing device according to claim 1 , wherein

the information of the mixture audio signal includes a feature value of the mixture audio signal, and

the processing circuitry is configured to:

obtain a feature value by integrating the feature value of the mixture audio signal and the extraction target information; and

output a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with the neural network by using the feature value obtained by the integrating and the feature value of the mixture audio signal.

7. The signal processing device according to claim 1 , wherein the processing circuitry is configured to perform processing of embedding the extraction target information by using a neural network.

8. The signal processing device according to claim 1 , wherein

the extraction target information is a target class vector indicating, by a vector, which audio class of the audio signal is to be extracted from the mixture audio signal,

the processing circuitry is further configured to perform processing of embedding the target class vector by using a neural network.

9. The signal processing device according to claim 1 , wherein the processing circuitry is configured to output the result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with the neural network by using the result of the integration and an intermediate feature value of the mixture audio signal.

10. The signal processing device according to claim 1 , wherein the processing circuitry is configured to output the result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal with the neural network by using an intermediate feature value derived based on the result of the integration and the intermediate feature value of the mixture audio signal.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072556/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2023
From: OCHIAI, TSUBASA; DELCROIX, MARC; KOIZUMI, YUMA; ITO, HIROAKI; KINOSHITA, KEISUKE; ARAKI, SHOKO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 062614/0853 →
Continuity (1)
Related Publication 20240038254A1 · Feb 1, 2024
References Cited (12)
US 20080300702A1 · Gomez · 2008 [cited by examiner]
US 20210192220A1 · Qu · 2021 [cited by examiner]
US 20210281739A1 · Takahashi · 2021 [cited by examiner]
US 20220277040A1 · Xu · 2022 [cited by examiner]
AU 2009278263B2 · 2012 [cited by examiner]
CN 108615532A · 2018 [cited by examiner]
WO 2020022055A1 · 2020 [cited by applicant]
Žmolíková et al., “SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures” IEEE Journal of Selected Topics in Signal Processing, vol. 13, No. 4, Available Online at: https://www.fit.v… [cited by applicant]
Kavalerov et al., “Universal Sound Separation”, 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Available Online at: https://arxiv.org/pdf/1905.03330.pdf, analysisarXiv: 1905.03330v2 [cs.… [cited by applicant]
Luo et al., “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, No. 8, Aug. 2019, pp. 1256-1266. [cited by applicant]
Ochiai et al., “Listen to What You Want: Neural Network-based Universal Sound Selector”, Available Online at: https://arxiv.org/abs/2006.05712, arXiv:2006.05712v1, Jun. 10, 2020, 11 pages. [cited by applicant]
Listen to What You Want: Neural Network-based Universal Sound Selector by Tsubasa Ochiai, submitted on Jun. 10, 2020 to arxiv.org. [cited by applicant]