IP Library Granted Patent US 9,792,904
Granted Patent B2
US 9,792,904 · App. 14/338,602 · Granted Oct 17, 2017

Methods and systems for natural language understanding using human knowledge and collected data

Inventors: Srinivas Bangalore (Morristown, NJ); Mazin Gilbert (Warren, NJ); Narendra K. Gupta (Dayton, NJ)
Assignee: Nuance Communications, Inc.
G10L15/183G06F17/2818G10L15/14G10L15/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,792,904
App. No.
14/338,602
Granted
Oct 17, 2017
Kind
B2
Abstract

Disclosed herein are systems and methods to incorporate human knowledge when developing and using statistical models for natural language understanding. The disclosed systems and methods embrace a data-driven approach to natural language understanding which progresses seamlessly along the continuum of availability of annotated collected data, from when there is no available annotated collected data to when there is any amount of annotated collected data.

Claims (40)

1. A method comprising:

operating, via a processor of a computing device, a natural language understanding application with a statistical model that generates a sequence of tags assigned to a first sequence of words;

developing a first part without human knowledge, the first part being first data used to formulate a replacement statistical model;

developing a second part using the human knowledge and annotated words, the second part being a second data used to formulate the replacement statistical model;

assigning weighted versions of the first part and the second part to yield the replacement statistical model, wherein the weighted versions are determined based on an amount of annotated data that is available; and

during execution of the natural language understanding application in which a second sequences of words is received as speech for processing using the natural language understanding application, assigning, via the processor of the computer device executing the natural language understanding application, a new sequence of tags to the second sequence of words, the new sequence of tags being generated by the replacement statistical model.

2. The method of claim 1 , wherein developing of the second part further comprises using a language model executor.

3. The method of claim 2 , wherein the language model executor is configured in run time to output the sequence of tags for an inputted sequence of words by using a statistical classifier model and a language model.

4. The method of claim 1 , wherein the new sequence of tags comprises tags from a predefined set of tags relating to different types of named entities.

5. The method of claim 1 , wherein developing of the replacement statistical model further comprises:

enumerating a phrase based on human knowledge for each tag in a predetermined set of possible tags for the natural language understanding application, to yield an enumerated phrase; and

developing a language model for each tag in the predetermined set of tags based on the enumerated phrase.

6. The method of claim 1 , further comprising replacing the statistical model with the replacement statistical model.

7. The method of claim 1 , further comprising determining when the replacement statistical model is sufficiently different than the statistical model to replace the statistical model.

8. The method of claim 7 , wherein determining when the replacement statistical model is sufficiently different than the statistical model comprises evaluating if a sufficiently large body of annotated collected data has been collected.

9. A system comprising: a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

operating a natural language understanding application with a statistical model that generates a first sequence of tags assigned to a sequence of words;

developing a first part without human knowledge, the first part being first data used to formulate a replacement statistical model;

developing a second part using the human knowledge and annotated words, the second part being a second data used to formulate the replacement statistical model;

assigning weighted versions of the first part and the second part to yield the replacement statistical model, wherein the weighted version are determined based on an amount of annotated data that is available; and

during execution of the natural language understanding application in which a second sequences of words is received as speech for processing using the natural language understanding application, assigning a new sequence of tags to the second sequence of words, the new sequence of tags being generated by the replacement statistical model.

10. The system of claim 9 , wherein developing of the second part further comprises using a language model executor.

11. The system of claim 10 , wherein the language model executor is configured in run time to output the new sequence of tags for an inputted sequence of words by using a statistical classifier model and a language model.

12. The system of claim 9 , wherein the new sequence of tags comprises tags from a predefined set of tags relating to different types of named entities.

13. The system of claim 9 , wherein developing of the replacement statistical model further comprises:

enumerating a phrase based on human knowledge for each tag in a predetermined set of possible tags for the natural language understanding application, to yield an enumerated phrase; and

developing a language model for each tag in the predetermined set of tags based on the enumerated phrase.

14. The system of claim 9 , the computer-readable storage medium having additional instructions stored which result in operations comprising replacing the statistical model with the replacement statistical model.

15. The system of claim 9 , the computer-readable storage medium having additional instructions stored which result in operations comprising determining when the replacement statistical model is sufficiently different than the statistical model to replace the statistical model.

16. The system of claim 15 , wherein determining when the replacement statistical model is sufficiently different than the statistical model comprises evaluating if a sufficiently large body of annotated collected data has been collected.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

operating a natural language understanding application with a statistical model that generates a first sequence of tags assigned to a sequence of words;

developing a first part without human knowledge, the first part being first data used to formulate a replacement statistical model;

developing a second part using the human knowledge and annotated words, the second part being a second data used to formulate the replacement statistical model;

assigning weighted versions of the first part and the second part to yield the replacement statistical model, wherein the weighted version are determined based on an amount of annotated data that is available; and

during execution of the natural language understanding application in which a second sequences of words is received as speech for processing using the natural language understanding application, assigning a new sequence of tags to the second sequence of words, the new sequence of tags being generated by the replacement statistical model.

18. The computer-readable storage device of claim 17 , wherein developing of the second part further comprises using a language model executor.

19. The computer-readable storage device of claim 18 , wherein the language model executor is configured in run time to output the new sequence of tags for an inputted sequence of words by using a statistical classifier model and a language model.

20. The computer-readable storage device of claim 17 , wherein the new sequence of tags comprises tags from a predefined set of tags relating to different types of named entities.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2016
From: BANGALORE, SRINIVAS; GILBERT, MAZIN; GUPTA, NARENDRA K.
To: AT&T CORP.
Reel/Frame 038124/0094 →
Continuity (3)
Continuation 13873548 · Apr 30, 2013
Continuation 11188825 · Jul 25, 2005
Related Publication 20140330555A1 · Nov 6, 2014