IP Library Granted Patent US 10,482,878
Granted Patent B2
US 10,482,878 · App. 15/831,808 · Granted Nov 19, 2019

System and method for speech enhancement in multisource environments

Inventors: Tobias Wolff (Neu-Ulm Burlafingen, DE); Jan Philip Janssen (Ulm, DE); Simon Graf (Ulm, DE); Tim Haulick (Blaubeuren, DE)
Assignee: Nuance Communications, Inc.
G10L15/20G10L15/183G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,482,878
App. No.
15/831,808
Granted
Nov 19, 2019
Kind
B2
Abstract

A method, computer program product, and computer system for receiving, by a computing device, a first signal emitted from one or more sources. A second signal may be received emitted from the one or more sources. A first confidence level that the wake-up-word is included in the first signal may be determined. A second confidence level that the wake-up-word is included in the second signal may be determined. It may be identified that the wake-up-word originated from a first source of the one or more sources based upon, at least in part, the first and second confidence levels. The first source may be enabled to participate in a dialog phase.

Claims (32)

1. A computer-implemented method comprising:

receiving, by a computing device, a first signal emitted from one or more sources;

receiving a second signal emitted from the one or more sources;

determining a first confidence level that a wake-up-word is included in the first signal;

determining a second confidence level that the wake-up-word is included in the second signal;

identifying that the wake-up-word originated from a first source of the one or more sources based upon, at least in part, the first and second confidence levels; and

enabling the first source to participate in a dialogue phase.

2. The computer-implemented method of claim 1 further comprising excluding at least a second source of the one or more sources from participating in the dialogue phase based upon, at least in part, historical information of the one or more sources.

3. The computer-implemented method of claim 1 wherein a Bayes decision process is used, at least in part, to identify the one or more sources.

4. The computer-implemented method of claim 1 wherein a wrapped Gaussian mixture model is used, at least in part, to identify the one or more sources.

5. The computer-implemented method of claim 1 further comprising tracking movement of the first source with at least one of one or more core localizers.

6. The computer-implemented method of claim 1 wherein the first signal and the second signal are received at a microphone group.

7. The computer-implemented method of claim 1 further comprising determining one or more angles at a given frequency expected to exhibit a maximum grating lobe based upon, at least in part, a latest model state.

8. The computer-implemented method of claim 1 further comprising extracting, by a beamformer, any source of the one or more sources except sources of the one or more sources deemed as interference.

9. A computer program product residing on a non-transitory computer readable storage medium having a plurality of instructions stored thereon which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising: receiving a first signal emitted from one or more sources; receiving a second signal emitted from the one or more sources; determining a first confidence level that a wake-up-word is included in the first signal; determining a second confidence level that the wake-up-word is included in the second signal; identifying that the wake-up-word originated from a first source of the one or more sources based upon, at least in part, the first and second confidence levels; and enabling the first source to participate in a dialogue phase.

10. The computer program product of claim 9 further comprising excluding at least a second source of the one or more sources from participating in the dialogue phase based upon, at least in part, historical information of the one or more sources.

11. The computer program product of claim 9 wherein a Bayes decision process is used, at least in part, to identify the one or more sources.

12. The computer program product of claim 9 wherein a wrapped Gaussian mixture model is used, at least in part, to identify the one or more sources.

13. The computer program product of claim 9 wherein the operations further comprise tracking movement of the first source with at least one of one or more core localizers.

14. The computer program product of claim 9 wherein the operations further comprise extracting, by a beamformer, any source of the one or more sources except sources of the one or more sources deemed as interference.

15. The computer program product of claim 9 wherein the operations further comprise determining one or more angles at a given frequency expected to exhibit a maximum grating lobe based upon, at least in part, a latest model state.

16. A computing system including one or more processors and one or more computer readable media having a plurality of instructions stored thereon which, when executed across the one or more processors, causes at least a portion of the one or more processors to perform operations comprising:

receiving a first signal emitted from one or more sources;

receiving a second signal emitted from the one or more sources;

determining a first confidence level that a wake-up-word is included in the first signal;

determining a second confidence level that the wake-up-word is included in the second signal;

identifying that the wake-up-word originated from a first source of the one or more sources based upon, at least in part, the first and second confidence levels; and

enabling the first source to participate in a dialogue phase.

17. The computing system of claim 16 wherein the operations further comprise excluding at least a second source of the one or more sources from participating in the dialogue phase based upon, at least in part, historical information of the one or more sources.

18. The computing system of claim 16 wherein a Bayes decision process is used, at least in part, to identify the one or more sources.

19. The computing system of claim 16 wherein a wrapped Gaussian mixture model is used, at least in part, to identify the one or more sources.

20. The computing system of claim 16 wherein the operations further comprise at least one of extracting, by a beamformer, any source of the one or more sources except sources of the one or more sources deemed as interference and determining one or more angles at a given frequency expected to exhibit a maximum grating lobe based upon, at least in part, a latest model state.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2017
From: WOLFF, TOBIAS; JANSSEN, JAN PHILIP; GRAF, SIMON; HAULICK, TIM
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 044299/0389 →
Continuity (2)
Continuation In Part 15825775 · Nov 29, 2017
Related Publication 20190164542A1 · May 30, 2019