IP Library › Granted Patent US 10,769,385
Granted Patent B2
US 10,769,385 · App. 16/204,467 · Granted Sep 8, 2020

System and method for inferring user intent from speech inputs

Inventor: Gunnar Evermann (Lincoln, MA)
Assignee: APPLE INC.
G06F40/35G10L15/1822G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,385
App. No.
16/204,467
Granted
Sep 8, 2020
Kind
B2
Abstract

A text string with a first and a second portion is provided. A domain of the text string is determined by applying a first word-matching process to the first portion of the text string. It is then determined whether the second portion of the text string matches a word of a set of words associated with the domain by applying a second word-matching process to the second portion of the text string. Upon determining that the second portion of the text string matches the word of the set of words, it is determined whether a user intent from the text string based at least in part on the domain and the word of the set of words.

Claims (46)

1. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by an electronic device with one or more processors and memory, cause the device to:

receive audio input containing a user utterance;

perform speech-to-text processing on the audio input to determine a plurality of text representations of the user utterance and a plurality of speech recognition scores for the plurality of text representations;

perform natural language processing on each of the plurality of text representations to determine a plurality of candidate user intents and a plurality of intent deduction scores for the plurality of candidate user intents;

determine a plurality of composite scores for the plurality of candidate user intents based on a combination of the plurality of speech recognition scores and the plurality of intent deduction scores;

select a user intent from the plurality of candidate user intents based on the plurality of composite scores; and

perform a task corresponding to the selected user intent.

2. The computer readable storage medium of claim 1 , wherein each intent deduction score of the plurality of intent deduction scores is based on a number of words in a respective text representation of the plurality of text representations that correspond to a domain of a respective candidate user intent of the plurality of candidate user intents.

3. The computer readable storage medium of claim 1 , wherein each intent deduction score of the plurality of intent deduction scores is based on a quality of match between words in a respective text representation of the plurality of text representations and predefined words corresponding to a domain of a respective candidate user intent of the plurality of candidate user intents.

4. The computer readable storage medium of claim 1 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a property of a domain of a respective candidate user intent of the plurality of candidate user intents can be resolved from a respective text representation of the plurality of text representations.

5. The computer readable storage medium of claim 1 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a natural language processor is able to identify a specific task from a respective text representation of the plurality of text representations.

6. The computer readable storage medium of claim 1 , wherein the selected user intent is determined from a first text representation of the plurality of text representations, and wherein an intent deduction score for the selected user intent is a highest score among the plurality of intent deduction scores, and wherein a speech recognition score for the first text representation is not a highest score among the plurality of speech recognition scores.

7. The computer readable storage medium of claim 1 , wherein the instructions further cause the device to:

rank the plurality of candidate user intents according to the plurality of composite scores, wherein the selected user intent has a highest composite score of the plurality of composite scores.

8. An electronic device, comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving audio input containing a user utterance;

performing speech-to-text processing on the audio input to determine a plurality of text representations of the user utterance and a plurality of speech recognition scores for the plurality of text representations;

performing natural language processing on each of the plurality of text representations to determine a plurality of candidate user intents and a plurality of intent deduction scores for the plurality of candidate user intents;

determining a plurality of composite scores for the plurality of candidate user intents based on a combination of the plurality of speech recognition scores and the plurality of intent deduction scores;

selecting a user intent from the plurality of candidate user intents based on the plurality of composite scores; and

performing a task corresponding to the selected user intent.

9. The device of claim 8 , wherein each intent deduction score of the plurality of intent deduction scores is based on a number of words in a respective text representation of the plurality of text representations that correspond to a domain of a respective candidate user intent of the plurality of candidate user intents.

10. The device of claim 8 , wherein each intent deduction score of the plurality of intent deduction scores is based on a quality of match between words in a respective text representation of the plurality of text representations and predefined words corresponding to a domain of a respective candidate user intent of the plurality of candidate user intents.

11. The device of claim 8 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a property of a domain of a respective candidate user intent of the plurality of candidate user intents can be resolved from a respective text representation of the plurality of text representations.

12. The device of claim 8 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a natural language processor is able to identify a specific task from a respective text representation of the plurality of text representations.

13. The device of claim 8 , wherein the selected user intent is determined from a first text representation of the plurality of text representations, and wherein an intent deduction score for the selected user intent is a highest score among the plurality of intent deduction scores, and wherein a speech recognition score for the first text representation is not a highest score among the plurality of speech recognition scores.

14. The device of claim 8 , wherein the one or more programs further include instructions for:

ranking the plurality of candidate user intents according to the plurality of composite scores, wherein the selected user intent has a highest composite score of the plurality of composite scores.

15. A method for inferring user intent from speech input, comprising:

at an electronic device with one or more processors and memory storing one or more programs for execution by the one or more processors:

receiving audio input containing a user utterance;

performing speech-to-text processing on the audio input to determine a plurality of text representations of the user utterance and a plurality of speech recognition scores for the plurality text representations;

performing natural language processing on each of the plurality of text representations to determine a plurality of candidate user intents and a plurality of intent deduction scores for the plurality of candidate user intents;

determining a plurality of composite scores for the plurality of candidate user intents based on a combination of the plurality of speech recognition scores and the plurality of intent deduction scores;

selecting a user intent from the plurality of candidate user intents based on the plurality of composite scores; and

performing a task corresponding to the selected user intent.

16. The method of claim 15 , wherein each intent deduction score of the plurality of intent deduction scores is based on a number of words in a respective text representation of the plurality of text representations that correspond to a domain of a respective candidate user intent of the plurality of candidate user intents.

17. The method of claim 15 , wherein each intent deduction score of the plurality of intent deduction scores is based on a quality of match between words in a respective text representation of the plurality of text representations and predefined words corresponding to a domain of a respective candidate user intent of the plurality of candidate user intents.

18. The method of claim 15 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a property of a domain of a respective candidate user intent of the plurality of candidate user intents can be resolved from a respective text representation of the plurality of text representations.

19. The method of claim 15 , wherein each intent deduction score of the plurality of intent deduction scores is based on whether or not a natural language processor is able to identify a specific task from a respective text representation of the plurality of text representations.

20. The method of claim 15 , wherein the selected user intent is determined from a first text representation of the plurality of text representations, and wherein an intent deduction score for the selected user intent is a highest score among the plurality of intent deduction scores, and wherein a speech recognition score for the first text representation is not a highest score among the plurality of speech recognition scores.

21. The method of claim 15 , further comprising:

ranking the plurality of candidate user intents according to the plurality of composite scores, wherein the selected user intent has a highest composite score of the plurality of composite scores.

Continuity (3)
Continuation 14298725 · Jun 6, 2014
Provisional Application 61832896 · Jun 9, 2013
Related Publication 20190179890A1 · Jun 13, 2019
Cited By (26)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,431,128 US 12,477,470 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452 US 12,748,568