IP Library › Granted Patent US 10,629,186
Granted Patent B1
US 10,629,186 · App. 13/793,856 · Granted Apr 21, 2020

Domain and intent name feature identification and processing

Inventor: Janet Louise Slifka (Cambridge, MA)
Assignee: Amazon Technologies, Inc.
G10L15/1815G06F17/278G10L15/1822G06F17/2785G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,629,186
App. No.
13/793,856
Granted
Apr 21, 2020
Kind
B1
Abstract

A system for improved natural language understanding (NLU) provides pre-feature input to a named entity recognition (NER) component. Pre-features may include no-textually derived information associated with the circumstances describing a user command (such as time, location, etc.). A domain and/or intent may also be determined prior to NER processing and may be passed to the NER component as an input. The pre-features and/or domain or intent information may assist the NER processing by providing context to a textual input, thereby improving NER processing such as semantic tagging, which in turn may improve overall NLU processing quality.

Claims (47)

1. A method of performing natural language processing, the method comprising:

receiving an audio signal comprising an utterance;

obtaining first text of the utterance using automatic speech recognition;

determining at least one of a category of commands or a potential intent corresponding to the utterance, wherein the at least one of the category or the potential intent are determined using data associated with the utterance and are not determined from the first text or based on a previous utterance that is associated with the utterance;

performing semantic tagging of the first text using a named entity recognition model based at least in part on at least one of the category or the potential intent, wherein the named entity recognition model was trained using a corpus of data comprising a first training text, and wherein the first training text was associated with at least one of a first category or a first intent;

determining, using the results of the semantic tagging, a selected intent for the utterance; and

performing an action using the selected intent.

2. The method of claim 1 , wherein the named entity recognition model comprises a conditional random field.

3. The method of claim 1 , wherein the category comprises at least one of a calendar category, a music category, a games category, or a communications category.

4. The method of claim 1 , wherein performing semantic tagging is further based at least in part on one or more of the following: a user identity, a user location, a user activity history, a time associated with the utterance, information stored on a user device, or a device type.

5. The method of claim 1 , wherein the selected intent comprises at least one of a play music intent, create calendar item intent, or a get directions to location intent.

6. The method of claim 1 , further comprising performing further natural language understanding processing based at least in part on the results of the semantic tagging.

7. The method of claim 1 , wherein the determining at least one of the category of commands or the potential intent is based at least in part on at least one of: a user identity, a user location, a time associated with the utterance, a date associated with the utterance, a volume associated with the utterance, a speech speed associated with the utterance, ambient noise associated with the utterance, information stored on a user device, or a device type.

8. A computing device, comprising:

at least one processor;

a memory device including instructions operable to be executed by the at least one processor to perform a set of actions, configuring the at least one processor:

to receive audio data corresponding to an input utterance of a user;

to perform automatic speech recognition (ASR) on the audio data to obtain text;

to determine one or more pre-features, wherein the one or more pre-features are data associated with the input utterance and are not determined from the text or based on a previous input utterance of the user;

to associate the one or more pre-features with the text;

to determine, based at least in part on the one or more pre-features, at least one of a category of commands corresponding to the input utterance or a potential intent corresponding to the input utterance, the potential intent corresponding to an intended command to be executed;

to generate a pre-feature vector including the one or more pre-features;

to associate an entity with at least one word of the text based at least in part on the pre-feature vector and the category or the intent; and

to determine, based at least in part on the entity, a selected intent for the input utterance.

9. The computing device of claim 7 , wherein the at least one processor is further configured:

to select a named entity recognition model based on the category or the potential intent; and

to associate the entity with the at least one word of the text based at least in part on the one or more pre-features and the selected named entity recognition model.

10. The computing device of claim 8 , wherein the one or more pre-features includes a user location.

11. The computing device of claim 8 , wherein the one or more pre-features includes a user activity history.

12. The computing device of claim 8 , wherein the one or more pre-features includes a user identity.

13. The computing device of claim 8 , wherein the at least one processor is further configured to determine the one or more pre-features based at least in part on at least one of: a user identity, a user location, a time associated with the input command, a date associated with the input command, a volume associated with the input command, a speech speed associated with the input command, ambient noise associated with the input command, information stored on a user device, or a device type.

14. A non-transitory computer-readable storage medium storing processor-executable instructions for controlling a computing device, comprising:

program code to receive audio data corresponding to an input utterance of a user;

program code to perform automatic speech recognition (ASR) on the audio data to obtain text;

program code to determine one or more pre-features, wherein the one or more pre-features are data associated with the input utterance and are not determined from the text or based on a previous input utterance of the user;

program code to associate the one or more pre-features with the text;

program code to determine, based at least in part on the one or more pre-features, at least one of a category of commands corresponding to the input utterance or a potential intent corresponding to the input utterance, the potential intent corresponding to an intended command to be executed;

program code to generate a pre-feature vector including the one or more pre-features;

program code to associate an entity with at least one word of the text based at least in part on the pre-feature vector; and

program code to determine, based at least in part on the entity, a selected intent for the input utterance.

15. The non-transitory computer-readable storage medium of claim 14 , further comprising:

program code to select a named entity recognition model based on the category or the potential intent based at least in part on the one or more pre-features, and

program code to associate the entity with the at least one word of the text based at least in part on the one or more pre-features and the selected named entity recognition model.

16. The non-transitory computer-readable storage medium of claim 14 , wherein the one or more pre-features includes a time associated with the text.

17. The non-transitory computer-readable storage medium of claim 14 , wherein the one or more pre-features includes information stored on a user device.

18. The non-transitory computer-readable storage medium of claim 14 , wherein the one or more pre-features includes a device type.

19. The non-transitory computer-readable storage medium of claim 14 , further comprising program code to determine the one or more pre-features based at least in part on at least one of: a user identity, a user location, a time associated with the input command, a date associated with the input command, a volume associated with the input command, a speech speed associated with the input command, ambient noise associated with the input command, information stored on a user device, or a device type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2013
From: SLIFKA, JANET LOUISE
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 030996/0865 →
Cited By (31)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,493 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,322,410 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,400,663 US 12,431,128 US 12,477,470 US 12,556,890 US 12,579,969 US 12,608,171 US 12,613,730 US 12,619,452 US 12,670,902 US 12,694,203