IP Library Granted Patent US 11,954,405
Granted Patent B2
US 11,954,405 · App. 17/982,337 · Granted Apr 9, 2024

Zero latency digital assistant

Inventors: William F. Stasior (Los Altos, CA); David A. Carson (San Francisco, CA); Rohit Dasari (San Francisco, CA); Yoon Kim (Los Altos, CA)
Assignee: Apple Inc.
G06F3/167G06F3/038G06F3/0481G06F3/0604G06F3/0656G06F3/0673G10L15/22G10L15/32G10L2015/088G10L2015/223G10L15/285H04M2201/40H04M2250/74
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,954,405
App. No.
17/982,337
Filed
Nov 7, 2022
Granted
Apr 9, 2024
Kind
B2
Art Unit
2143
USPC
715/727
Abstract

An electronic device can implement a zero-latency digital assistant by capturing audio input from a microphone and using a first processor to write audio data representing the captured audio input to a memory buffer. In response to detecting a user input while capturing the audio input, the device can determine whether the user input meets a predetermined criteria. If the user input meets the criteria, the device can use a second processor to identify and execute a task based on at least a portion of the contents of the memory buffer.

Claims (92)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device with a microphone, cause the device to:

while a main processor is in a low-power mode:

capture audio input from the microphone;

write, using a low-power processor, the captured audio input to a memory buffer; and

determine, using the low-power processor, whether the buffered audio input satisfies a set of criteria, wherein determining whether the buffered audio input satisfies the set of criteria includes:

determining whether a first portion of the buffered audio input includes a trigger phrase, wherein a first criterion of the set of criteria is satisfied when the first portion of the buffered audio input includes the trigger phrase; and

determining, by comparing speech characteristics of the buffered audio input to a set of known characteristics for an authorized user, whether the trigger phrase was spoken by the authorized user, wherein a second criterion of the set of criteria is satisfied when the first portion of the buffered audio input is spoken by the authorized user; and

in accordance with a determination that the buffered audio input satisfies the set of criteria:

cause the main processor to exit the low-power mode; and

after causing the main processor to exit the low-power mode:

identify, using the main processor, a computing task based on a second portion of the buffered audio input; and

execute, using the main processor, the identified computing task; and

in accordance with a determination that the buffered audio input does not satisfy the set of criteria, forgo causing the main processor to exit the low-power mode, identifying the computing task, and executing the computing task.

2. The non-transitory computer-readable storage medium according to claim 1 , wherein the main processor executes the identified computing task immediately, without requiring a further input from a user.

3. The non-transitory computer-readable storage medium according to claim 1 , wherein the trigger phrase corresponds to a request to launch a digital assistant session.

4. The non-transitory computer-readable storage medium according to claim 1 , wherein identifying the computing task includes launching a digital assistant on the main processor, and wherein identifying and executing the identified computing task is performed by the digital assistant.

5. The non-transitory computer-readable storage medium according to claim 4 , further comprising instructions to cause the electronic device to:

further in accordance with the determination that the buffered audio input satisfies the set of criteria, provide the second portion of the buffered audio input to the digital assistant.

6. The non-transitory computer-readable storage medium according to claim 5 , wherein providing the second portion of the buffered audio input to the digital assistant includes providing the second portion of the buffered audio input to a remote server associated with the digital assistant.

7. The non-transitory computer-readable storage medium according to claim 4 , wherein launching the digital assistant includes displaying a user interface associated with the digital assistant.

8. The non-transitory computer-readable storage medium according to claim 7 , wherein the user interface associated with the digital assistant is displayed in a full-screen view.

9. The non-transitory computer-readable storage medium according to claim 4 , wherein launching the digital assistant includes activating one or more audio components on the device.

10. The non-transitory computer-readable storage medium of claim 1 , wherein the speech characteristics include a spectral pattern.

11. The non-transitory computer-readable storage medium of claim 1 , wherein the speech characteristics include speech patterns.

12. The non-transitory computer-readable storage medium of claim 1 , wherein the speech characteristics include a speech intonation.

13. The non-transitory computer-readable storage medium according to claim 4 , further comprising instructions to cause the electronic device to:

further in accordance with the determination that the buffered audio input satisfies the set of criteria:

activating a second microphone on the device, and

streaming audio detected by the second microphone to the digital assistant.

14. A method, comprising:

at an electronic device including a microphone, and one or more processors:

while a main processor is in a low-power mode:

capturing audio input from the microphone;

writing, using a low-power processor, the captured audio input to a memory buffer; and

determining, using the low-power processor, whether the buffered audio input satisfies a set of criteria, wherein determining whether the buffered audio input satisfies the set of criteria includes:

determining whether a first portion of the buffered audio input includes a trigger phrase, wherein a first criterion of the set of criteria is satisfied when the first portion of the buffered audio input includes the trigger phrase; and

determining, by comparing speech characteristics of the buffered audio input to a set of known characteristics for an authorized user, whether the trigger phrase was spoken by the authorized user, wherein a second criterion of the set of criteria is satisfied when the first portion of the buffered audio input is spoken by the authorized user; and

in accordance with a determination that the buffered audio input satisfies the set of criteria:

causing the main processor to exit the low-power mode; and

after causing the main processor to exit the low-power mode:

identifying, using the main processor, a computing task based on a second portion of the buffered audio input; and

executing, using the main processor, the identified computing task; and

in accordance with a determination that the buffered audio input does not satisfy the set of criteria, forgoing causing the main processor to exit the low-power mode, identifying the computing task, and executing the computing task.

15. The method of claim 14 , wherein the main processor executes the identified computing task immediately, without requiring a further input from a user.

16. The method of claim 14 , wherein the trigger phrase corresponds to a request to launch a digital assistant session.

17. The method of claim 14 , wherein identifying the computing task includes launching a digital assistant on the main processor, and wherein identifying and executing the identified computing task is performed by the digital assistant.

18. The method of claim 17 , further comprising:

further in accordance with the determination that the buffered audio input satisfies the set of criteria, providing the second portion of the buffered audio input to the digital assistant.

19. The method of claim 18 , wherein providing the second portion of the buffered audio input to the digital assistant includes providing the second portion of the buffered audio input to a remote server associated with the digital assistant.

20. The method of claim 17 , wherein launching the digital assistant includes displaying a user interface associated with the digital assistant.

21. The method of claim 20 , wherein the user interface associated with the digital assistant is displayed in a full-screen view.

22. The method of claim 17 , wherein launching the digital assistant includes activating one or more audio components on the device.

23. The method of claim 17 , further comprising:

further in accordance with the determination that the buffered audio input satisfies the set of criteria:

activating a second microphone on the device, and

streaming audio detected by the second microphone to the digital assistant.

24. The method of claim 14 , wherein the speech characteristics include a spectral pattern.

25. The method of claim 14 , wherein the speech characteristics include speech patterns.

26. The method of claim 14 , wherein the speech characteristics include a speech intonation.

27. An electronic device, comprising:

a microphone;

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

while a main processor is in a low-power mode:

capturing audio input from the microphone;

writing, using a low-power processor, the captured audio input to a memory buffer; and

determining, using the low-power processor, whether the buffered audio input satisfies a set of criteria, wherein determining whether the buffered audio input satisfies the set of criteria includes:

determining whether a first portion of the buffered audio input includes a trigger phrase, wherein a first criterion of the set of criteria is satisfied when the first portion of the buffered audio input includes the trigger phrase; and

determining, by comparing speech characteristics of the buffered audio input to a set of known characteristics for an authorized user, whether the trigger phrase was spoken by the authorized user, wherein a second criterion of the set of criteria is satisfied when the first portion of the buffered audio input is spoken by the authorized user; and

in accordance with a determination the buffered audio input satisfied the set of criteria:

causing the main processor to exit the low-power mode; and

after causing the main processor to exit the low-power mode:

identifying, using the main processor, a computing task based on a second portion of the buffered audio input; and

executing, using the main processor, the identified computing task; and

in accordance with a determination that the buffered audio input does not satisfy the set of criteria, forgoing causing the main processor to exit the low-power mode, identifying the computing task, and executing the computing task.

28. The electronic device of claim 27 , wherein the main processor executes the identified computing task immediately, without requiring a further input from a user.

29. The electronic device of claim 27 , wherein the trigger phrase corresponds to a request to launch a digital assistant session.

30. The electronic device of claim 27 , wherein identifying the computing task includes launching a digital assistant on the main processor, and wherein identifying and executing the identified computing task is performed by the digital assistant.

31. The electronic device of claim 30 , the one or more programs further including instructions for:

further in accordance with the determination that the buffered audio input satisfies the set of criteria, providing the second portion of the buffered audio input to the digital assistant.

32. The electronic device of claim 31 , wherein providing the second portion of the buffered audio input to the digital assistant includes providing the second portion of the buffered audio input to a remote server associated with the digital assistant.

33. The electronic device of claim 30 , wherein launching the digital assistant includes displaying a user interface associated with the digital assistant.

34. The electronic device of claim 33 , wherein the user interface associated with the digital assistant is displayed in a full-screen view.

35. The electronic device of claim 30 , wherein launching the digital assistant includes activating one or more audio components on the device.

36. The electronic device of claim 30 , the one or more programs further including instructions for:

further in accordance with the determination that the buffered audio input satisfies the set of criteria:

activating a second microphone on the device, and

streaming audio detected by the second microphone to the digital assistant.

37. The electronic device of claim 27 , wherein the speech characteristics include a spectral pattern.

38. The electronic device of claim 27 , wherein the speech characteristics include speech patterns.

39. The electronic device of claim 27 , wherein the speech characteristics include a speech intonation.

Continuity (5)
Continuation 17403674 · Aug 16, 2021
Continuation 16909852 · Jun 23, 2020
Continuation 15147726 · May 5, 2016
Provisional Application 62215608 · Sep 8, 2015
Related Publication 20230057442A1 · Feb 23, 2023
Cited By (1)
US 12,315,510