IP Library Granted Patent US 10,332,545
Granted Patent B2
US 10,332,545 · App. 15/823,874 · Granted Jun 25, 2019

System and method for temporal and power based zone detection in speaker dependent microphone environments

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,332,545
App. No.
15/823,874
Granted
Jun 25, 2019
Kind
B2
Abstract

A method, computer program product, and computer system for receiving, by a computing device, a speech signal from a speaker via a plurality of microphone zones. A temporal cue based confidence may be determined for at least a portion of the plurality of microphone zones. A power cue based confidence may be determined for at least a portion of the plurality of microphone zones. A microphone zone of the plurality of microphone zones from which to use an output signal of the speaker may be identified based upon, at least in part, a combination of the temporal cue based confidence and the power cue based confidence.

Claims (33)

1. A computer-implemented method comprising:

receiving, by a computing device, a speech signal from a speaker via a plurality of microphone zones;

determining temporal cue based confidence for at least a portion of the plurality of microphone zones;

determining power cue based confidence for the at least the portion of the plurality of microphone zones, wherein the at least the portion includes at least two zones of the plurality of zones;

identifying a microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker, based upon, at least in part, a combination of the temporal cue based confidence determined for the at least the portion of the plurality of microphone zones and the power cue based confidence determined for the at least the portion of the plurality of microphone zones; and

using the speech signal from the identified microphone zone as an output signal in a speech system;

wherein the identifying includes comparing the temporal cue based confidence and the power cue based confidence, and wherein the identifying further includes selecting the temporal cue based confidence to identify the microphone zone when the temporal cue based confidence is higher than the power cue based confidence.

2. The computer-implemented method of claim 1 wherein the identifying the microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker further includes selecting the power cue based confidence to identify the microphone zone when the power cue based confidence is higher than the temporal cue based confidence.

3. The computer-implemented method of claim 1 wherein the temporal cue based confidence is based upon, at least in part, a signal-to-noise ratio.

4. The computer-implemented method of claim 1 wherein the temporal cue based confidence is based upon, at least in part, evaluation of a sign of an imaginary part of a delta phase in different frequency sub-bands in a sub-band domain.

5. The computer-implemented method of claim 1 wherein the temporal cue based confidence is based upon, at least in part, an observed delay being one of greater than and less than a pre-defined delay.

6. A computer program product residing on a non-transitory computer readable storage medium having a plurality of instructions stored thereon which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising:

receiving a speech signal from a speaker via a plurality of microphone zones;

determining temporal cue based confidence for at least a portion of the plurality of microphone zones;

determining power cue based confidence for the at least the portion of the plurality of microphone zones, wherein the at least the portion includes at least two zones of the plurality of zones;

identifying a microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker, based upon, at least in part, a combination of the temporal cue based confidence determined for the at least the portion of the plurality of microphone zones and the power cue based confidence determined for the at least the portion of the plurality of microphone zones; and

using the speech signal from the identified microphone zone as an output signal in a speech system;

wherein the identifying includes comparing the temporal cue based confidence and the power cue based confidence, and wherein the identifying further includes selecting the temporal cue based confidence to identify the microphone zone when the temporal cue based confidence is higher than the power cue based confidence.

7. The computer program product of claim 6 wherein the identifying the microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker further includes selecting the power cue based confidence to identify the microphone zone when the power cue based confidence is higher than the temporal cue based confidence.

8. The computer program product of claim 6 wherein the temporal cue based confidence is based upon, at least in part, a signal-to-noise ratio.

9. The computer program product of claim 6 wherein the temporal cue based confidence is based upon, at least in part, evaluation of a sign of an imaginary part of a delta phase in different frequency sub-bands in a sub-band domain.

10. The computer program product of claim 6 wherein the temporal cue based confidence is based upon, at least in part, an observed delay being one of greater than and less than a pre-defined delay.

11. A computing system including one or more processors and one or more memories configured to perform operations comprising:

receiving a speech signal from a speaker via a plurality of microphone zones;

determining temporal cue based confidence for at least a portion of the plurality of microphone zones;

determining power cue based confidence for the at least the portion of the plurality of microphone zones, wherein the at least the portion includes at least two zones of the plurality of zones;

identifying a microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker, based upon, at least in part, a combination of the temporal cue based confidence determined for the at least the portion of the plurality of microphone zones and the power cue based confidence determined for the at least the portion of the plurality of microphone zones; and

using the speech signal from the identified microphone zone as an output signal in a speech system;

wherein the identifying includes comparing the temporal cue based confidence and the power cue based confidence, and wherein the identifying further includes selecting the temporal cue based confidence to identify the microphone zone when the temporal cue based confidence is higher than the power cue based confidence.

12. The computing system of claim 11 wherein the identifying the microphone zone of the plurality of microphone zones from which the speech signal originates from the speaker further includes selecting the power cue based confidence to identify the microphone zone when the power cue based confidence is higher than the temporal cue based confidence.

13. The computing system of claim 11 wherein the temporal cue based confidence is based upon, at least in part, a signal-to-noise ratio.

14. The computing system of claim 11 wherein the temporal cue based confidence is based upon, at least in part, evaluation of a sign of an imaginary part of a delta phase in different frequency sub-bands in a sub-band domain.

15. The computing system of claim 11 wherein the temporal cue based confidence is based upon, at least in part, an observed delay being one of greater than and less than a pre-defined delay.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2017
From: MATHEJA, TIMO; BUCK, MARKUS; GRAF, SIMON
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 044233/0591 →