IP Library › Granted Patent US 12,148,432
Granted Patent B2
US 12,148,432 · App. 17/756,874 · Granted Nov 19, 2024

Signal processing device, signal processing method, and signal processing system

Inventor: Atsuo Hiroe (Tokyo, JP)
Assignee: SONY GROUP CORPORATION
G10L17/18G10L21/034G10L25/78H04R1/406H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,148,432
App. No.
17/756,874
Granted
Nov 19, 2024
Kind
B2
Abstract

Provided is a signal processing device including a main speech detection unit that detects, by using a neural network, whether or not a signal input to a sound collection device assigned to each of at least two speakers includes a main speech that is a voice of the corresponding speaker, and outputs frame information indicating presence or absence of the main speech.

Claims (70)

1. A signal processing device, comprising:

a main speech detection unit configured to:

receive a first signal from a first sound collection device, a second signal from a second sound collection device, and a third signal from a third sound collection device, wherein

the first sound collection device is associated with a first speaker,

the second sound collection device is associated with a second speaker, and

the third sound collection device is associated with a third speaker;

input, to a first neural network, the first signal and the second signal to obtain first information;

input, to a second neural network, the first signal and the third signal to obtain second information;

input, to a third neural network, the second signal and the third signal to obtain third information;

detect, by integration of the first information and the second information, a presence or an absence of a main speech of the first speaker in the first signal;

output first frame information indicating the presence or the absence of the main speech of the first speaker in the first signal;

detect, by integration of the first information and the third information, a presence or an absence of a main speech of the second speaker in the second signal;

output second frame information indicating the presence or the absence of the main speech of the second speaker in the second signal;

detect, by integration of the second information and the third information, a presence or an absence of a main speech of the third speaker in the third signal; and

output third frame information indicating the presence or the absence of the main speech of the third speaker in the third signal.

2. The signal processing device according to claim 1 , wherein the main speech detection unit is further configured to detect the presence or the absence of the main speech of the first speaker in the first signal, in a case where the main speech of the second speaker is included in the first signal.

3. The signal processing device according to claim 1 , wherein the main speech detection unit is further configured to output time information of a frame that includes the main speech of the first speaker.

4. The signal processing device according to claim 1 , further comprising a crosstalk reduction unit configured to reduce the main speech of the second speaker included in the first signal.

5. The signal processing device according to claim 4 , wherein the crosstalk reduction unit is further configured to reduce, based on the first frame information output from the main speech detection unit, the main speech of the second speaker included in the first signal.

6. The signal processing device according to claim 4 , wherein the crosstalk reduction unit is further configured to perform a process on a specific signal of a frame that includes the main speech of the first speaker.

7. The signal processing device according to claim 1 , further comprising a voice recognition unit configured to perform a voice recognition process on a specific signal, wherein the specific signal is based on the first frame information.

8. The signal processing device according to claim 7 , further comprising a text information generation unit configured to generate text information based on a result of the voice recognition process.

9. A signal processing method, comprising:

in a signal processing device:

receiving a first signal from a first sound collection device, a second signal from a second sound collection device, and a third signal from a third sound collection device, wherein

the first sound collection device is associated with a first speaker,

the second sound collection device is associated with a second speaker, and

the third sound collection device is associated with a third speaker;

inputting, to a first neural network, the first signal and the second signal to obtain first information;

inputting, to a second neural network, the first signal and the third signal to obtain second information;

inputting, to a third neural network, the second signal and the third signal to obtain third information;

detecting, by integrating the first information and the second information, a presence or an absence of a main speech of the first speaker in the first signal;

outputting first frame information indicating the presence or the absence of the main speech of the first speaker in the first signal;

detecting, by integrating the first information and the third information, a presence or an absence of a main speech of the second speaker in the second signal;

outputting second frame information indicating the presence or the absence of the main speech of the second speaker in the second signal;

detecting, by integrating the second information and the third information, a presence or an absence of a main speech of the third speaker in the third signal; and

outputting third frame information indicating the presence or the absence of the main speech of the third speaker in the third signal.

10. A non-transitory computer-readable medium having stored thereon computer-executable instructions which, when executed by a computer, cause the computer to execute operations, the operations comprising:

receiving a first signal from a first sound collection device, a second signal from a second sound collection device, and a third signal from a third sound collection device, wherein

the first sound collection device is associated with a first speaker,

the second sound collection device is associated with a second speaker, and

the third sound collection device is associated with a third speaker;

inputting, to a first neural network, the first signal and the second signal to obtain first information;

inputting, to a second neural network, the first signal and the third signal to obtain second information;

inputting, to a third neural network, the second signal and the third signal to obtain third information;

detecting, by integrating the first information and the second information, a presence or an absence of a main speech of the first speaker in the first signal;

outputting first frame information indicating the presence or the absence of the main speech of the first speaker in the first signal;

detecting, by integrating the first information and the third information, a presence or an absence of a main speech of the second speaker in the second signal;

outputting second frame information indicating the presence or the absence of the main speech of the second speaker in the second signal;

detecting, by integrating the second information and the third information, a presence or an absence of a main speech of the third speaker in the third signal; and

outputting third frame information indicating the presence or the absence of the main speech of the third speaker in the third signal.

11. A signal processing system, comprising:

a plurality of sound collection devices, wherein

a first sound collection device of the plurality of sound collection devices is associated with a first speaker,

a second sound collection device of the plurality of sound collection devices is associated with a second speaker, and

a third sound collection device of the plurality of sound collection devices is associated with a third speaker; and

a signal processing device including a main speech detection unit, wherein the main speech detection unit is configured to:

receive a first signal from the first sound collection device, a second signal from the second sound collection device, and a third signal from the third sound collection device;

input, to a first neural network, the first signal and the second signal to obtain first information;

input, to a second neural network, the first signal and the third signal to obtain second information;

input, to a third neural network, the second signal and the third signal to obtain third information;

detect, by integration of the first information and the second information, a presence or an absence of a main speech of the first speaker in the first;

output first frame information indicating the presence or the absence of the main speech of the first speaker in the first signal;

detect, by integration of the first information and the third information, a presence or an absence of a main speech of the second speaker in the second signal;

output second frame information indicating the presence or the absence of the main speech of the second speaker in the second signal;

detect, by integration of the second information and the third information, a presence or an absence of a main speech of the third speaker in the third signal; and

output third frame information indicating the presence or the absence of the main speech of the third speaker in the third signal.

12. The signal processing system according to claim 11 , wherein

each of the plurality of sound collection devices includes a microphone, and

the first sound collection device is one of wearable by the first speaker or has directivity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2022
From: HIROE, ATSUO
To: SONY GROUP CORPORATION
Reel/Frame 060099/0901 →
Priority Claims (1)
JP 2019-227192 · Dec 17, 2019 · national
Continuity (1)
Related Publication 20230005488A1 · Jan 5, 2023