IP Library Granted Patent US 12,700,419
Granted Patent B2
US 12,700,419 · App. 18/275,950 · Granted Aug 4, 2026

Sound source separation method, sound source separation apparatus, and program

Inventors: Naoki Makishima (Tokyo, JP); Ryo Masumura (Tokyo, JP)
Assignee: NTT, Inc.
G10L21/028G10L15/063G10L15/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,419
App. No.
18/275,950
Filed
Aug 4, 2023
Granted
Aug 4, 2026
Kind
B2
Art Unit
2653
USPC
704/232
Abstract

A mixed acoustic signal including sound emitted from a plurality of sound sources and sound source video signals representing at least one video of the plurality of sound sources are received as inputs, and at least a separated signal including a signal representing a target sound emitted from one sound source represented by the video is acquired. However, at least the separated signal is acquired using properties of the sound source that affects sound emitted by the sound source acquired from the video and/or features of a structure used for the sound source to emit the sound.

Claims (20)

1 . A method for separating a sound source, comprising:

estimating a separated signal including a signal representing a target sound emitted from a sound source among a plurality of sound sources by applying a mixed acoustic signal and sound source video signals to a model, wherein the mixed acoustic signal represents a mixed sound of sound emitted from the plurality of sound sources, the sound source video signals represent videos of at least some of the plurality of sound sources to the model,

the model is obtained by learning using a neural network, wherein the model is learned based on differences between features of the separated signal and features of teaching data of sound source video signals, the features of the separated signal are obtained by applying teaching mixed acoustic signals as teaching data of the mixed acoustic signal and teaching sound source video signals as teaching data of the sound source video signals to the model, wherein

the learning is further based on:

a similarity between a first element representing the features of the teaching sound source video signals corresponding to a first sound source among the plurality of sound sources and a second element representing the features of the separated signal corresponding to a second sound source different from the first sound source decreases, and

a similarity between the first element representing the features of the teaching sound source video signals corresponding to the first sound source and a third element representing the features of the separated signal corresponding to the first sound source increases.

2 . The according to claim 1 , wherein

the model is further obtained by learning based on differences between the separated signal obtained by applying the teaching mixed acoustic signals and the teaching sound source video signals to the model and teaching sound source signals which are teaching data of the separated signal corresponding to the teaching mixed acoustic signals and the teaching sound source video signals.

3 . The method according to claim 1 , wherein

the sound source video signals represent a video of each of the plurality of sound sources.

4 . The method according to claim 1 , wherein

the plurality of sound sources includes a plurality of different speakers, the mixed acoustic signal includes a speech signal, and the sound source video signals represent videos of the speakers.

5 . The method according to claim 1 , wherein

the separated signal includes a signal representing a target sound emitted from a certain sound source among the plurality of sound sources and a signal representing a sound emitted from another sound source.

6 . The method according to claim 1 , wherein

the model is further obtained by learning based on differences between the separated signal obtained by applying the teaching mixed acoustic signals and the teaching sound source video signals to the model and teaching sound source signals which are teaching data of the separated signal corresponding to the teaching mixed acoustic signals and the teaching sound source video signals.

7 . The method according to claim 1 , wherein

the sound source video signals represent a video of each of the plurality of sound sources.

8 . The method according to claim 1 , wherein

the plurality of sound sources includes a plurality of different speakers, the mixed acoustic signal includes a speech signal, and the sound source video signals represent videos of the speakers.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0693 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2023
From: MAKISHIMA, NAOKI; MASUMURA, RYO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 064498/0833 →
Continuity (2)
Related Publication 20240135950A1 · Apr 25, 2024
Related Publication 20240233744A9 · Jul 11, 2024
References Cited (10)
US 20050195990A1 · Kondo · 2005 [cited by examiner]
US 20160180865A1 · Citerin · 2016 [cited by examiner]
US 20190206417A1 · Woodruff · 2019 [cited by examiner]
US 20210174817A1 · Grauman · 2021 [cited by examiner]
US 20220067386A1 · Rotman · 2022 [cited by examiner]
CN 110246512A · 2019 [cited by examiner]
JP 2020187346A · 2020 [cited by applicant]
Yu et al. (2017) “Permutation invariant training of deep models for speaker-independent multi-talker speech separation” ICASSP, Mar. 5, 2017, pp. 241-245. [cited by applicant]
Lu et al. (2019) “Audio-visual deep clustering for speech separation” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, No. 11, pp. 1697-1712. [cited by applicant]
Ephrat et al. (2018) “Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation” ACM Trans. on Graphics, vol. 37, No. 4, pp. 112:1-112:11. [cited by applicant]