IP Library › Granted Patent US 10,650,841
Granted Patent B2
US 10,650,841 · App. 15/558,259 · Granted May 12, 2020

Sound source separation apparatus and method

Inventor: Yuhki Mitsufuji (Tokyo, JP)
Assignee: SONY CORPORATION
G10L21/028G10L21/0272H04R1/406H04R3/005G10L2021/02166H04S2400/11H04S2420/07H04S2420/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,650,841
App. No.
15/558,259
Filed
Sep 14, 2017
Granted
May 12, 2020
Kind
B2
Art Unit
2653
USPC
381/92
Abstract

The present technology relates to a sound source separation apparatus and a method which make it possible to separate a sound source at lower calculation cost. A communication unit receives a spatial frequency spectrum of a sound collection signal which is obtained by a microphone array collecting a plane wave of sound from a sound source, and a spatial frequency mask generating unit generates a spatial frequency mask for masking a component of a predetermined region in a spatial frequency domain on the basis of the spatial frequency spectrum. A sound source separating unit extracts a component of a desired sound source from the spatial frequency spectrum as an estimated sound source spectrum on the basis of the spatial frequency mask. The present technology can be applied to a spatial frequency sound source separator.

Claims (33)

1. A sound source separation apparatus, comprising:

a central processing unit (CPU) configured to:

obtain a multichannel sound signal via a microphone array;

generate a spatial frequency spectrum based on the multichannel sound signal;

generate a spatial frequency mask to mask a component of a specific region in a spatial frequency domain, wherein the spatial frequency mask is generated based on:

a direction of arrival of the multichannel sound signal from a specific sound source, and

the spatial frequency spectrum; and

extract, as an estimated sound source spectrum, a component of the specific sound source based on a multiplication of the spatial frequency spectrum with the spatial frequency mask.

2. The sound source separation apparatus according to claim 1 , wherein the CPU is further configured to generate the spatial frequency mask through blind sound source separation.

3. The sound source separation apparatus according to claim 2 , wherein the CPU is further configured to generate the spatial frequency mask through the blind sound source separation by utilization of non-negative matrix factorization.

4. The sound source separation apparatus according to claim 1 , wherein the CPU is further configured to generate the spatial frequency mask through sound source separation based on information associated with the specific sound source.

5. The sound source separation apparatus according to claim 4 , wherein the information associated with the specific sound source indicates the direction of arrival.

6. The sound source separation apparatus according to claim 5 , wherein the CPU is further configured to generate the spatial frequency mask based on an adaptive beam former.

7. The sound source separation apparatus according to claim 1 , wherein the CPU is further configured to:

generate a drive signal in the spatial frequency domain based on the estimated sound source spectrum;

reproduce the multichannel sound signal based on the drive signal;

calculate a time-frequency spectrum based on spatial frequency synthesis on the drive signal;

generate a speaker drive signal based on time frequency synthesis on the time-frequency spectrum; and

reproduce, via a speaker array, the multichannel sound signal based on the speaker drive signal.

8. A sound source separation method, comprising:

obtaining a multichannel sound signal via a microphone array;

generating a spatial frequency spectrum based on the multichannel sound signal;

generating a spatial frequency mask for masking a component of a specific region in a spatial frequency domain, wherein the spatial frequency mask is generated based on:

a direction of arrival of the multichannel sound signal from a specific sound source, and

the spatial frequency spectrum; and

extracting, as an estimated sound source spectrum, a component of the specific sound source based on a multiplication of the spatial frequency spectrum with the spatial frequency mask.

9. A non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed by a processor, cause the processor to execute operations, the operations comprising:

obtaining a multichannel sound signal via a microphone array;

generating a spatial frequency spectrum based on the multichannel sound signal;

generating a spatial frequency mask for masking a component of a specific region in a spatial frequency domain, wherein the spatial frequency mask is generated based on:

a direction of arrival of the multichannel sound signal from a specific sound source, and

the spatial frequency spectrum; and

extracting, as an estimated sound source spectrum, a component of the specific sound source based on a multiplication of the spatial frequency spectrum with the spatial frequency mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2017
From: MITSUFUJI, YUHKI
To: SONY CORPORATION
Reel/Frame 043860/0404 →
Priority Claims (1)
JP 2015-059318 · Mar 23, 2015 · national
Continuity (1)
Related Publication 20180047407A1 · Feb 15, 2018