IP Library Granted Patent US 10,789,945
Granted Patent B2
US 10,789,945 · App. 15/679,595 · Granted Sep 29, 2020

Low-latency intelligent automated assistant

Inventors: Alejandro Acero (Monte Sereno, CA); Hepeng Zhang (Sunnyvale, CA)
Assignee: Apple Inc.
G10L15/1822G10L15/22G10L15/30G10L25/87G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,789,945
App. No.
15/679,595
Filed
Aug 17, 2017
Granted
Sep 29, 2020
Kind
B2
Art Unit
2677
USPC
704/275
Abstract

Systems and processes for operating a digital assistant are provided. In an example process, low-latency operation of a digital assistant is provided. In this example, natural language processing, task flow processing, dialogue flow processing, speech synthesis, or any combination thereof can be at least partially performed while awaiting detection of a speech end-point condition. Upon detection of a speech end-point condition, results obtained from performing the operations can be presented to the user. In another example, robust operation of a digital assistant is provided. In this example, task flow processing by the digital assistant can include selecting a candidate task flow from a plurality of candidate task flows based on determined task flow scores. The task flow scores can be based on speech recognition confidence scores, intent confidence scores, flow parameter scores, or any combination thereof. The selected candidate task flow is executed and corresponding results presented to the user.

Claims (112)

1. An electronic device, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a stream of audio, comprising:

receiving, from a first time to a second time, a first portion of the stream of audio containing at least a portion of a user utterance, wherein one or more candidate text representations are determined based on the at least a portion of the user utterance while receiving the first portion of the stream of audio; and

receiving, from the second time to a third time, a second portion of the stream of audio, wherein the electronic device stops receiving the stream of audio at the third time;

after determining the one or more candidate text representations, determining whether the first portion of the stream of audio satisfies a predetermined condition;

in response to determining that the first portion of the stream of audio satisfies the predetermined condition, performing, at least partially between the second time and the third time, operations comprising:

determining, based on the one or more candidate text representations of the at least a portion of the user utterance, a plurality of candidate user intents for the at least a portion of the user utterance, wherein each candidate user intent of the plurality of candidate user intents corresponds to a respective candidate task flow of a plurality of candidate task flows;

selecting a first candidate task flow of the plurality of candidate task flows; and

executing the first candidate task flow without providing an output to a user of the device;

determining whether a speech end-point condition is detected between the second time and the third time; and

in response to determining that a speech end-point condition is detected between the second time and the third time, presenting, to the user, results from executing the selected first candidate task flow.

2. The device of claim 1 , wherein the one or more programs further include instructions for:

in response to determining that a speech end-point condition is not detected between the second time and the third time, forgoing presentation of the results to the user.

3. The device of claim 1 , wherein the one or more programs further include instructions for:

determining whether the second portion of the stream of audio contains a continuation of the user utterance, wherein the results are presented to the user in response to:

determining that the second portion of the stream of audio does not contain a continuation of the user utterance; and

determining that a speech end-point condition is detected between the second time and the third time.

4. The device of claim 3 , wherein the one or more programs further include instructions for:

in response to determining that the second portion of the stream of audio contains a continuation of the user utterance, forgoing presentation of the results to the user.

5. The device of claim 3 , wherein the one or more programs further include instructions for:

in response to determining that the second portion of the stream of audio contains a continuation of the user utterance:

receiving, from the third time to a fourth time, a third portion of the stream of audio;

determining whether the second portion of the stream of audio satisfies a predetermined condition;

in response to determining that the second portion of the stream of audio satisfies a predetermined condition, performing, at least partially between the third time and a fourth time, operations comprising:

determining, based on a second plurality of candidate text representations of the user utterance in the first and second portions of the stream of audio, a second plurality of candidate user intents for the user utterance, wherein each second candidate user intent of the second plurality of candidate user intents corresponds to a respective second candidate task flow of a second plurality of candidate task flows;

selecting a second candidate task flow of the second plurality of candidate task flows; and

executing the selected second candidate task flow without providing an output to the user.

6. The device of claim 5 , wherein the one or more programs further include instructions for:

determining whether a speech end-point condition is detected between the third time and the fourth time; and

in response to determining that a speech end-point condition is detected between the third time and the fourth time, presenting, to the user, second results from executing the selected second candidate task flow.

7. The device of claim 1 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an absence of user speech for longer than a first predetermined duration after the at least a portion of the user utterance.

8. The device of claim 7 , wherein detecting the speech end-point condition comprises detecting, in the second portion of the stream of audio, an absence of user speech for greater than a second predetermined duration, and wherein the second predetermined duration is longer than the first predetermined duration.

9. The device of claim 1 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an audio energy level that is less than a predetermined threshold energy level for longer than a first predetermined duration after the at least a portion of the user utterance.

10. The device of claim 1 , wherein the predetermined condition comprises an end-of-sentence condition based on a sequence of words in the at least a portion of the user utterance.

11. The device of claim 1 , wherein detecting the speech end-point condition comprises detecting a predetermined type of non-speech input from the user.

12. The device of claim 1 , wherein executing the first candidate task flow comprises:

generating a text dialogue that is responsive to the at least a portion of the user utterance without outputting the text dialogue or a spoken representation of the text dialogue.

13. The device of claim 12 , wherein presenting the results to the user comprises:

determining whether the memory of the device stores an audio file having the spoken representation of the text dialogue; and

in response to determining that the memory of the device stores an audio file having the spoken representation of the text dialogue, outputting, to the user, the spoken representation of the text dialogue by playing the stored audio file.

14. The device of claim 12 , wherein the one or more programs further include instructions for performing, at least partially prior to detecting the speech end-point condition, operations comprising:

determining whether the memory of the device stores an audio file having the spoken representation of the text dialogue; and

in response to determining that the memory of the device does not store an audio file having the spoken representation of the text dialogue:

generating an audio file having the spoken representation of the text dialogue; and

storing the audio file in the memory.

15. The device of claim 1 , wherein the one or more programs further include instructions for:

determining a plurality of task flow scores for the plurality of candidate task flows, each task flow score of the plurality of task flow scores corresponding to a respective candidate task flow of the plurality of candidate task flows;

wherein selecting the first candidate task flow of the plurality of candidate task flows is based on the plurality of task flow scores.

16. The device of claim 15 , wherein the one or more programs further include instructions for:

resolving, for each candidate task flow of the plurality of candidate task flows, one or more flow parameters of the respective candidate task flow, wherein a respective task flow score for the respective candidate task flow is based on the resolving one or more flow parameters of the respective candidate task flow.

17. The device of claim 15 , wherein the one or more programs further include instructions for:

ranking the plurality of candidate task flows in accordance with the plurality of task flow scores, wherein selecting the first candidate task flow is based on the ranking of the plurality of candidate task flows.

18. The device of claim 15 , wherein the plurality of task flow scores are determined based on a plurality of intent confidence scores for the plurality of candidate user intents.

19. The device of claim 15 , wherein the plurality of task flow scores are determined based on one or more speech recognition confidence scores for the one or more candidate text representations.

20. The device of claim 1 , wherein the one or more programs further include instructions for:

in response to determining, while the operations are being performed, that the speech end-point condition is not detected between the second time and the third time and that the second portion of the stream of audio contains a continuation of the user utterance, ceasing to continue performing the operations.

21. The device of claim 1 , wherein the first candidate task flow is at least partially executed prior to determining that the speech end-point condition is detected.

22. A method for operating a digital assistant, the method comprising:

at an electronic device having one or more processors and memory:

receiving a stream of audio, comprising:

receiving, from a first time to a second time, a first portion of the stream of audio containing at least a portion of a user utterance, wherein one or more candidate text representations are determined based on the at least a portion of the user utterance while receiving the first portion of the stream of audio; and

receiving, from the second time to a third time, a second portion of the stream of audio, wherein the electronic device stops receiving the stream of audio at the third time;

after determining the one or more candidate text representations, determining whether the first portion of the stream of audio satisfies a predetermined condition;

in response to determining that the first portion of the stream of audio satisfies the predetermined condition, performing, at least partially between the second time and the third time, operations comprising:

determining, based on the one or more candidate text representations of the at least a portion of the user utterance, a plurality of candidate user intents for the at least a portion of the user utterance, wherein each candidate user intent of the plurality of candidate user intents corresponds to a respective candidate task flow of a plurality of candidate task flows ;

selecting a first candidate task flow of the plurality of candidate task flows; and

executing the first candidate task flow without providing an output to a user of the device;

determining whether a speech end-point condition is detected between the second time and the third time; and

in response to determining that a speech end-point condition is detected between the second time and the third time, presenting, to the user, results from executing the selected first candidate task flow.

23. The method of claim 22 , further comprising:

in response to determining that a speech end-point condition is not detected between the second time and the third time, forgoing presentation of the results to the user.

24. The method of claim 22 , further comprising:

determining whether the second portion of the stream of audio contains a continuation of the user utterance, wherein the results are presented to the user in response to:

determining that the second portion of the stream of audio does not contain a continuation of the user utterance; and

determining that a speech end-point condition is detected between the second time and the third time.

25. The method of claim 24 , further comprising:

in response to determining that the second portion of the stream of audio contains a continuation of the user utterance, forgoing presentation of the results to the user.

26. The method of claim 22 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an absence of user speech for longer than a first predetermined duration after the at least a portion of the user utterance.

27. The method of claim 22 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an audio energy level that is less than a predetermined threshold energy level for longer than a first predetermined duration after the at least a portion of the user utterance.

28. The method of claim 22 , further comprising:

determining a plurality of task flow scores for the plurality of candidate task flows, each task flow score of the plurality of task flow scores corresponding to a respective candidate task flow of the plurality of candidate task flows;

wherein selecting the first candidate task flow of the plurality of candidate task flows is based on the plurality of task flow scores.

29. The method of claim 28 , wherein the plurality of task flow scores are determined based on a plurality of intent confidence scores for the plurality of candidate user intents.

30. The method of claim 28 , wherein the plurality of task flow scores are determined based on one or more speech recognition confidence scores for the one or more candidate text representations.

31. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for:

receiving a stream of audio, comprising:

receiving, from a first time to a second time, a first portion of the stream of audio containing at least a portion of a user utterance, wherein one or more candidate text representations are determined based on the at least a portion of the user utterance while receiving the first portion of the stream of audio; and

receiving, from the second time to a third time, a second portion of the stream of audio, wherein the electronic device stops receiving the stream of audio at the third time;

after determining the one or more candidate text representations, determining whether the first portion of the stream of audio satisfies a predetermined condition;

in response to determining that the first portion of the stream of audio satisfies the predetermined condition, performing, at least partially between the second time and the third time, operations comprising:

determining, based on the one or more candidate text representations of the at least a portion of the user utterance, a plurality of candidate user intents for the at least a portion of the user utterance, wherein each candidate user intent of the plurality of candidate user intents corresponds to a respective candidate task flow of a plurality of candidate task flows ;

selecting a first candidate task flow of the plurality of candidate task flows; and

executing the first candidate task flow without providing an output to a user of the device;

determining whether a speech end-point condition is detected between the second time and the third time; and

in response to determining that a speech end-point condition is detected between the second time and the third time, presenting, to the user, results from executing the selected first candidate task flow.

32. The non-transitory computer-readable storage medium of claim 31 , wherein the one or more programs further include instructions for:

in response to determining that a speech end-point condition is not detected between the second time and the third time, forgoing presentation of the results to the user.

33. The non-transitory computer-readable storage medium of claim 31 , wherein the one or more programs further include instructions for:

determining whether the second portion of the stream of audio contains a continuation of the user utterance, wherein the results are presented to the user in response to:

determining that the second portion of the stream of audio does not contain a continuation of the user utterance; and

determining that a speech end-point condition is detected between the second time and the third time.

34. The non-transitory computer-readable storage medium of claim 33 , wherein the one or more programs further include instructions for:

in response to determining that the second portion of the stream of audio contains a continuation of the user utterance, forgoing presentation of the results to the user.

35. The non-transitory computer-readable storage medium of claim 31 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an absence of user speech for longer than a first predetermined duration after the at least a portion of the user utterance.

36. The non-transitory computer-readable storage medium of claim 31 , wherein the predetermined condition comprises a condition of detecting, in the first portion of the stream of audio, an audio energy level that is less than a predetermined threshold energy level for longer than a first predetermined duration after the at least a portion of the user utterance.

37. The non-transitory computer-readable storage medium of claim 31 , wherein the one or more programs further include instructions for:

determining a plurality of task flow scores for the plurality of candidate task flows, each task flow score of the plurality of task flow scores corresponding to a respective candidate task flow of the plurality of candidate task flows;

wherein selecting the first candidate task flow of the plurality of candidate task flows is based on the plurality of task flow scores.

38. The non-transitory computer-readable storage medium of claim 37 , wherein the plurality of task flow scores are determined based on a plurality of intent confidence scores for the plurality of candidate user intents.

39. The non-transitory computer-readable storage medium of claim 37 , wherein the plurality of task flow scores are determined based on one or more speech recognition confidence scores for the one or more candidate text representations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2017
From: ACERO, ALEJANDRO; ZHANG, HEPENG
To: APPLE INC.
Reel/Frame 043906/0950 →
Continuity (2)
Provisional Application 62505546 · May 12, 2017
Related Publication 20180330723A1 · Nov 15, 2018
Cited By (26)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,327,558 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,431,128 US 12,477,470 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452