IP Library Granted Patent US 11,636,869
Granted Patent B2
US 11,636,869 · App. 17/150,513 · Granted Apr 25, 2023

Voice trigger for a digital assistant

Inventors: Justin Binder (Oakland, CA); Samuel D. Post (San Jose, CA); Onur Tackin (Sunnyvale, CA); Thomas R. Gruber (Santa Cruz, CA)
Assignee: Apple Inc.
G10L21/16G06F3/167G10L15/22G10L15/26G10L17/24G10L15/02G10L15/30G10L25/51G10L25/84G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,869
App. No.
17/150,513
Granted
Apr 25, 2023
Kind
B2
Abstract

A method for operating a voice trigger is provided. In some implementations, the method is performed at an electronic device including one or more processors and memory storing instructions for execution by the one or more processors. The method includes receiving a sound input. The sound input may correspond to a spoken word or phrase, or a portion thereof. The method includes determining whether at least a portion of the sound input corresponds to a predetermined type of sound, such as a human voice. The method includes, upon a determination that at least a portion of the sound input corresponds to the predetermined type, determining whether the sound input includes predetermined content, such as a predetermined trigger word or phrase. The method also includes, upon a determination that the sound input includes the predetermined content, initiating a speech-based service, such as a voice-based digital assistant.

Claims (88)

1. A non-transitory computer-readable storage medium, storing one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for:

determining, based on comparing an amount of light detected on at least a front surface of the electronic device to a threshold amount of light, whether to operate a voice trigger in a standby mode or in a listening mode;

in accordance with a determination to operate the voice trigger in the listening mode:

receiving a sound input;

determining whether the sound input includes predetermined content; and

upon a determination that the sound input includes the predetermined content, initiating a speech-based service; and

in accordance with a determination to operate the voice trigger in the standby mode, forgoing initiating the speech-based service based on received sound input.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the predetermined content is one or more words.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the predetermined content is one or more predetermined phonemes.

4. The non-transitory computer-readable storage medium of claim 3 , wherein the one or more predetermined phonemes constitute at least one word.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

prior to determining whether the sound input includes the predetermined content, determining whether the sound input satisfies a predetermined condition.

6. The non-transitory computer-readable storage medium of claim 5 , wherein the predetermined condition is an amplitude threshold.

7. The non-transitory computer-readable storage medium of claim 5 , wherein said determining whether the sound input satisfies the predetermined condition is performed by a first sound detector, wherein the first sound detector consumes less power while operating than a second sound detector configured to determine whether the sound input includes the predetermined content.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

storing at least a portion of the sound input in memory; and

providing the portion of the sound input to the speech-based service once the speech-based service is initiated.

9. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

determining whether the sound input corresponds to a voice of a particular user.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the speech-based service is initiated upon a determination that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the speech-based service is initiated in a limited access mode upon a determination that the sound input includes the predetermined content and that the sound input does not correspond to the voice of the particular user.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the one or more programs further include instructions for:

upon a determination that the sound input corresponds to the voice of the particular user, outputting a voice prompt including a name of the particular user.

13. The non-transitory computer-readable storage medium of claim 1 , wherein determining, based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether to operate the voice trigger in the standby mode or in the listening mode includes:

determining based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether the electronic device is face-up on a surface or face-down on the surface;

in accordance with determining that the electronic device is face-up on the surface, determining to operate the voice trigger in the listening mode; and

in accordance with determining that the electronic device is face-down on the surface, determining to operate the voice trigger in the standby mode.

14. The non-transitory computer-readable storage medium of claim 13 , wherein determining whether the electronic device is face-up on the surface or face-down on the surface includes comparing the amount of light detected on the front surface of the electronic device to an amount of light detected on a back surface of the electronic device.

15. A method for operating a voice trigger, comprising:

at an electronic device including one or more processors and memory storing instructions for execution by the one or more processors:

determining, based on comparing an amount of light detected on at least a front surface of the electronic device to a threshold amount of light, whether to operate the voice trigger in a standby mode or in a listening mode;

in accordance with a determination to operate the voice trigger in the listening mode:

receiving a sound input;

determining whether the sound input includes predetermined content; and

upon a determination that the sound input includes the predetermined content, initiating a speech-based service; and

in accordance with a determination to operate the voice trigger in the standby mode, forgoing initiating the speech-based service based on received sound input.

16. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

determining, based on comparing an amount of light detected on at least a front surface of the electronic device to a threshold amount of light, whether to operate a voice trigger in a standby mode or in a listening mode;

in accordance with a determination to operate the voice trigger in the listening mode:

receiving a sound input;

determining whether the sound input includes predetermined content; and

upon a determination that the sound input includes the predetermined content, initiating a speech-based service; and

in accordance with a determination to operate the voice trigger in the standby mode, forgoing initiating the speech-based service based on received sound input.

17. The method of claim 15 , wherein the predetermined content is one or more words.

18. The method of claim 15 , wherein the predetermined content is one or more predetermined phonemes.

19. The method of claim 18 , wherein the one or more predetermined phonemes constitute at least one word.

20. The method of claim 15 , further comprising:

prior to determining whether the sound input includes the predetermined content, determining whether the sound input satisfies a predetermined condition.

21. The method of claim 20 , wherein the predetermined condition is an amplitude threshold.

22. The method of claim 20 , wherein said determining whether the sound input satisfies the predetermined condition is performed by a first sound detector, wherein the first sound detector consumes less power while operating than a second sound detector configured to determine whether the sound input includes the predetermined content.

23. The method of claim 15 , further comprising:

storing at least a portion of the sound input in memory; and

providing the portion of the sound input to the speech-based service once the speech-based service is initiated.

24. The method of claim 15 , further comprising:

determining whether the sound input corresponds to a voice of a particular user.

25. The method of claim 24 , wherein the speech-based service is initiated upon a determination that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user.

26. The method of claim 25 , wherein the speech-based service is initiated in a limited access mode upon a determination that the sound input includes the predetermined content and that the sound input does not correspond to the voice of the particular user.

27. The method of claim 26 , further comprising:

upon a determination that the sound input corresponds to the voice of the particular user, outputting a voice prompt including a name of the particular user.

28. The method of claim 15 , wherein determining, based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether to operate the voice trigger in the standby mode or in the listening mode includes:

determining based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether the electronic device is face-up on a surface or face-down on the surface;

in accordance with determining that the electronic device is face-up on the surface, determining to operate the voice trigger in the listening mode; and

in accordance with determining that the electronic device is face-down on the surface, determining to operate the voice trigger in the standby mode.

29. The method of claim 28 , wherein determining whether the electronic device is face-up on the surface or face-down on the surface includes comparing the amount of light detected on the front surface of the electronic device to an amount of light detected on a back surface of the electronic device.

30. The electronic device of claim 16 , wherein the predetermined content is one or more words.

31. The electronic device of claim 16 , wherein the predetermined content is one or more predetermined phonemes.

32. The electronic device of claim 31 , wherein the one or more predetermined phonemes constitute at least one word.

33. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

prior to determining whether the sound input includes the predetermined content, determining whether the sound input satisfies a predetermined condition.

34. The electronic device of claim 33 , wherein the predetermined condition is an amplitude threshold.

35. The electronic device of claim 33 , wherein said determining whether the sound input satisfies the predetermined condition is performed by a first sound detector, wherein the first sound detector consumes less power while operating than a second sound detector configured to determine whether the sound input includes the predetermined content.

36. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

storing at least a portion of the sound input in memory; and

providing the portion of the sound input to the speech-based service once the speech-based service is initiated.

37. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

determining whether the sound input corresponds to a voice of a particular user.

38. The electronic device of claim 37 , wherein the speech-based service is initiated upon a determination that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user.

39. The electronic device of claim 38 , wherein the speech-based service is initiated in a limited access mode upon a determination that the sound input includes the predetermined content and that the sound input does not correspond to the voice of the particular user.

40. The electronic device of claim 39 , wherein the one or more programs further include instructions for:

upon a determination that the sound input corresponds to the voice of the particular user, outputting a voice prompt including a name of the particular user.

41. The electronic device of claim 16 , wherein determining, based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether to operate the voice trigger in the standby mode or in the listening mode includes:

determining based on comparing the amount of light detected on at least the front surface of the electronic device to the threshold amount of light, whether the electronic device is face-up on a surface or face-down on the surface;

in accordance with determining that the electronic device is face-up on the surface, determining to operate the voice trigger in the listening mode; and

in accordance with determining that the electronic device is face-down on the surface, determining to operate the voice trigger in the standby mode.

42. The electronic device of claim 41 , wherein determining whether the electronic device is face-up on the surface or face-down on the surface includes comparing the amount of light detected on the front surface of the electronic device to an amount of light detected on a back surface of the electronic device.

Continuity (5)
Continuation 16879348 · May 20, 2020
Continuation 16222249 · Dec 17, 2018
Continuation 14175864 · Feb 7, 2014
Provisional Application 61762260 · Feb 7, 2013
Related Publication 20210151070A1 · May 20, 2021