IP Library Granted Patent US 9,454,959
Granted Patent B2
US 9,454,959 · App. 14/440,343 · Granted Sep 27, 2016

Method and apparatus for passive data acquisition in speech recognition and natural language understanding

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,454,959
App. No.
14/440,343
Granted
Sep 27, 2016
Kind
B2
Abstract

Speech recognition systems often process speech by employing models and analyzing audio data. An embodiment of the method and corresponding system described herein allow for passive monitoring of, for example, conversation between user(s) to determine context to use to prime model(s) for later speech recognition requests submitted to the speech recognition system. The embodiment improves the results of the speech recognition system by updating speech recognition model(s) with contextual information of the conversation. This increases the probability that the speech recognition system interprets the conversation to contextually relevant information.

Claims (31)

1. A method comprising:

extracting contextual information from a conversation by performing speech recognition passively in a background while a speech recognition system is inactive from executing speech recognition requests from a user;

applying the contextual information to a model of the speech recognition system to enhance analyzing speech directed to the speech recognition system by updating the model of the speech recognition system using the contextual information from the conversation; and

processing a speech recognition request from the user, while the speech recognition system is active to execute the speech recognition request, using the updated model of the speech recognition system;

wherein the model is at least one of an acoustic model, language model, contextual model, likely words model, vocabulary, domain-specific language model, and domain specific vocabulary.

2. The method of claim 1 , wherein the speech recognition system is included in a device, wherein the conversation is directed to the device via an audio channel and wherein the speech recognition system extracts contextual information by accessing the audio channel at times before the speech recognition system is activated to perform speech recognition.

3. The method of claim 1 , further comprising:

collecting contextual information from at least one of a user, user and a third party, one or more third parties, or device outputting audio that can be interpreted as the conversation.

4. The method of claim 1 , wherein extracting the contextual information and applying the contextual information are performed on at least one of (i) a client device configured to record the conversation and the speech, and (ii) a server, the method further comprising storing the extracted contextual information on the client device or a storage system accessible by the client device during speech recognition by the speech recognition system.

5. The method of claim 1 , further comprising filtering a selected speaker from the conversation by employing at least one of speaker identification and verification.

6. The method of claim 5 , wherein filtering includes employing at least one of speaker identification and verification by further employing speaker segmentation.

7. The method of claim 1 , further comprising weighting contextual information as a function of a time of the conversation.

8. The method of claim 1 , further comprising replacing or suppressing identifying information in the extracted contextual information.

9. A system comprising:

an extraction module configured to extract contextual information from a conversation by performing speech recognition passively in a background while a speech recognition system is inactive from executing speech recognition requests from a user;

an enhancement module configured to apply the contextual information to a model of the speech recognition system to enhance analyzing speech directed to the speech recognition system by updating the model of the speech recognition system using the contextual information from the conversation; and

a speech recognition module configured to process a speech recognition request from the user, while the speech recognition system is active to execute the speech recognition request, using the updated model of the speech recognition system;

wherein the model is at least one of an acoustic model, language model, contextual model, likely words model, vocabulary, domain-specific language model, and domain specific vocabulary.

10. The system of claim 9 , wherein the speech recognition system is included in a device, wherein the conversation is directed to the device via an audio channel and wherein the speech recognition system extracts contextual information by accessing the audio channel at times before the speech recognition system is activated to perform speech recognition.

11. The system of claim 9 , further comprising:

a recording module configured to collect contextual information from at least one of a user, user and a third party, one or more third parties, or device outputting audio that can be interpreted as the conversation.

12. The system of claim 9 , further comprising at least one of (i) a client device configured to record the conversation and the speech, and (ii) a server, wherein the extraction module is further configured to store the extracted contextual information on at least one of the client device and a storage system accessible by the client device during speech recognition by the speech recognition system.

13. The system of claim 9 , further comprising a filtering module configured to filter a selected speaker from the conversation by employing at least one of speaker identification and verification.

14. The system of claim 13 , wherein the filtering module is further configured to employ speaker segmentation.

15. The system of claim 9 , further comprising a weighting module configured to weigh contextual information as a function of a time of the conversation.

16. The system of claim 9 , further comprising a suppression module configured to replace or suppress identifying information in the extracted contextual information.

17. A computer program product comprising non-transitory computer readable medium storing instructions for performing a method, the method comprising:

extracting contextual information from a conversation by performing speech recognition passively in a background while a speech recognition system is inactive from executing speech recognition requests from a user;

applying the contextual information to a model of the speech recognition system to enhance analyzing speech directed to the speech recognition system by updating the model of the speech recognition system using the contextual information from the conversation; and

processing a speech recognition request from the user, while the speech recognition system is active to execute the speech recognition request, using the updated model of the speech recognition system;

wherein the model is at least one of an acoustic model, language model, contextual model, likely words model, vocabulary, domain-specific language model, and domain specific vocabulary.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2015
From: LENKE, NILS; GANONG, WILLIAM F., III
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 035646/0332 →