IP Library Granted Patent US 11,557,310
Granted Patent B2
US 11,557,310 · App. 17/713,741 · Granted Jan 17, 2023

Voice trigger for a digital assistant

Inventors: Justin Binder (Oakland, CA); Samuel D. Post (San Jose, CA); Onur Tackin (Sunnyvale, CA); Thomas R. Gruber (Santa Cruz, CA)
Assignee: Apple Inc.
G10L21/16G06F3/167G10L15/22G10L15/26G10L17/24G10L15/02G10L15/30G10L25/51G10L25/84G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,557,310
App. No.
17/713,741
Granted
Jan 17, 2023
Kind
B2
Abstract

A method for operating a voice trigger is provided. In some implementations, the method is performed at an electronic device including one or more processors and memory storing instructions for execution by the one or more processors. The method includes receiving a sound input. The sound input may correspond to a spoken word or phrase, or a portion thereof. The method includes determining whether at least a portion of the sound input corresponds to a predetermined type of sound, such as a human voice. The method includes, upon a determination that at least a portion of the sound input corresponds to the predetermined type, determining whether the sound input includes predetermined content, such as a predetermined trigger word or phrase. The method also includes, upon a determination that the sound input includes the predetermined content, initiating a speech-based service, such as a voice-based digital assistant.

Claims (111)

1. A non-transitory computer-readable storage medium, storing one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for:

receiving a sound input;

determining whether the sound input corresponds to predetermined content based on comparing an input representation of the sound input to one or more reference representations of one or more trigger words for a voice trigger;

in accordance with determining that the sound input does not correspond to the predetermined content, forgoing initiating a speech based service;

after determining that the sound input does not correspond to the predetermined content, detecting, within a predetermined duration of receiving the sound input, user input to initiate the speech based service, wherein the user input to initiate the speech based service corresponds to a user selection of a button of the electronic device or a user selection of an affordance displayed by the electronic device; and

in accordance with detecting the user input to initiate the speech based service within the predetermined duration, adjusting, based on a determination that the user input to initiate the speech based service is detected within the predetermined duration, the one or more reference representations using the sound input to increase a confidence level of a match between the one or more reference representations and the input representation of the sound input.

2. The non-transitory computer-readable storage medium of claim 1 , wherein adjusting the one or more reference representations includes updating a moving average of the one or more reference representations using the sound input.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

prior to receiving the sound input, receiving a plurality of additional sound inputs, wherein each sound input of the plurality of additional sound inputs includes the one or more trigger words; and

generating the one or more reference representations using the plurality of additional sound inputs.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

in accordance with determining that the sound input corresponds to the predetermined content, initiating the speech based service.

5. The non-transitory computer-readable storage medium of claim 4 , wherein comparing the input representation of the sound input to the one or more reference representations includes determining whether the input representation of the sound input matches the one or more reference representations with a predetermined confidence level, and wherein the one or more programs further include instructions for:

in accordance with initiating the speech based service:

in accordance with a determination that the input representation matches the one or more reference representations with the predetermined confidence level:

adjusting the one or more reference representations using the sound input; and

in accordance with a determination that the input representation of the sound input does not match the one or more reference representations with the predetermined confidence level:

forgoing adjusting the one or more reference representations using the sound input.

6. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

after determining that the sound input does not correspond to the predetermined content, receiving one or more sound inputs consecutive to the sound input, wherein each sound input of the one or more sound inputs includes the one or more trigger words; and

adjusting the one or more reference representations using the one or more sound inputs.

7. The non-transitory computer-readable storage medium of claim 6 , wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the one or more sound inputs include at least two successive sound inputs.

8. The non-transitory computer-readable storage medium of claim 6 , wherein the one or more sound inputs include at least two sound inputs, and wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the at least two sound inputs are received within a predetermined time period.

9. The non-transitory computer-readable storage medium of claim 1 , wherein comparing the input representation of the sound input to the one or more reference representations includes comparing the input representation to the one or more reference representations using an algorithm, and wherein the one or more programs further include instructions for:

in accordance with detecting the user input to initiate the speech based service, adjusting the algorithm.

10. The non-transitory computer-readable storage medium of claim 1 , wherein:

determining whether the sound input corresponds to predetermined content is performed using a first processor; and

adjusting the one or more reference representations using the sound input is performed using a second processor.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the second processor is in a standby mode when determining whether the sound input corresponds to predetermined content.

12. The non-transitory computer-readable storage medium of claim 10 , wherein the one or more processors further include instructions for:

storing the sound input in a memory of the electronic device;

in accordance with detecting the user input to initiate the speech based service, activating the second processor, wherein adjusting the one or more reference representations using the sound input includes accessing, by the activated second processor, the sound input.

13. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more processors further include instructions for:

prior to determining whether the sound input corresponds to predetermined content, determining whether the sound input corresponds to a predetermined type of sound.

14. The non-transitory computer-readable storage medium of claim 13 , wherein determining whether the sound input corresponds to predetermined content is performed in accordance with determining that the sound input corresponds to the predetermined type of sound.

15. The non-transitory computer-readable storage medium of claim 1 , wherein the speech based service includes a digital assistant.

16. An electronic device, comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a sound input;

determining whether the sound input corresponds to predetermined content based on comparing an input representation of the sound input to one or more reference representations of one or more trigger words for a voice trigger;

in accordance with determining that the sound input does not correspond to the predetermined content, forgoing initiating a speech based service;

after determining that the sound input does not correspond to the predetermined content, detecting, within a predetermined duration of receiving the sound input, user input to initiate the speech based service, wherein the user input to initiate the speech based service corresponds to a user selection of a button of the electronic device or a user selection of an affordance displayed by the electronic device; and

in accordance with detecting the user input to initiate the speech based service within the predetermined duration, adjusting, based on a determination that the user input to initiate the speech based service is detected within the predetermined duration, the one or more reference representations using the sound input to increase a confidence level of a match between the one or more reference representations and the input representation of the sound input.

17. A method for operating a voice trigger, performed at an electronic device including one or more processors and memory storing instructions for execution by the one or more processors, the method comprising:

receiving a sound input;

determining whether the sound input corresponds to predetermined content based on comparing an input representation of the sound input to one or more reference representations of one or more trigger words for the voice trigger;

in accordance with determining that the sound input does not correspond to the predetermined content, forgoing initiating a speech based service;

after determining that the sound input does not correspond to the predetermined content, detecting, within a predetermined duration of receiving the sound input, user input to initiate the speech based service, wherein the user input to initiate the speech based service corresponds to a user selection of a button of the electronic device or a user selection of an affordance displayed by the electronic device; and

in accordance with detecting the user input to initiate the speech based service within the predetermined duration, adjusting, based on a determination that the user input to initiate the speech based service is detected within the predetermined duration, the one or more reference representations using the sound input to increase a confidence level of a match between the one or more reference representations and the input representation of the sound input.

18. The electronic device of claim 16 , wherein adjusting the one or more reference representations includes updating a moving average of the one or more reference representations using the sound input.

19. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

prior to receiving the sound input, receiving a plurality of additional sound inputs, wherein each sound input of the plurality of additional sound inputs includes the one or more trigger words; and

generating the one or more reference representations using the plurality of additional sound inputs.

20. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

in accordance with determining that the sound input corresponds to the predetermined content, initiating the speech based service.

21. The electronic device of claim 20 , wherein comparing the input representation of the sound input to the one or more reference representations includes determining whether the input representation of the sound input matches the one or more reference representations with a predetermined confidence level, and wherein the one or more programs further include instructions for:

in accordance with initiating the speech based service:

in accordance with a determination that the input representation matches the one or more reference representations with the predetermined confidence level:

adjusting the one or more reference representations using the sound input; and

in accordance with a determination that the input representation of the sound input does not match the one or more reference representations with the predetermined confidence level:

forgoing adjusting the one or more reference representations using the sound input.

22. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

after determining that the sound input does not correspond to the predetermined content, receiving one or more sound inputs consecutive to the sound input, wherein each sound input of the one or more sound inputs includes the one or more trigger words; and

adjusting the one or more reference representations using the one or more sound inputs.

23. The electronic device of claim 22 , wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the one or more sound inputs include at least two successive sound inputs.

24. The electronic device of claim 22 , wherein the one or more sound inputs include at least two sound inputs, and wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the at least two sound inputs are received within a predetermined time period.

25. The electronic device of claim 16 , wherein comparing the input representation of the sound input to the one or more reference representations includes comparing the input representation to the one or more reference representations using an algorithm, and wherein the one or more programs further include instructions for:

in accordance with detecting the user input to initiate the speech based service, adjusting the algorithm.

26. The electronic device of claim 16 , wherein:

determining whether the sound input corresponds to predetermined content is performed using a first processor; and

adjusting the one or more reference representations using the sound input is performed using a second processor.

27. The electronic device of claim 26 , wherein the second processor is in a standby mode when determining whether the sound input corresponds to predetermined content.

28. The electronic device of claim 26 , wherein the one or more processors further include instructions for:

storing the sound input in a memory of the electronic device;

in accordance with detecting the user input to initiate the speech based service, activating the second processor, wherein adjusting the one or more reference representations using the sound input includes accessing, by the activated second processor, the sound input.

29. The electronic device of claim 16 , wherein the one or more processors further include instructions for:

prior to determining whether the sound input corresponds to predetermined content, determining whether the sound input corresponds to a predetermined type of sound.

30. The electronic device of claim 29 , wherein determining whether the sound input corresponds to predetermined content is performed in accordance with determining that the sound input corresponds to the predetermined type of sound.

31. The electronic device of claim 16 , wherein the speech based service includes a digital assistant.

32. The method of claim 17 , wherein adjusting the one or more reference representations includes updating a moving average of the one or more reference representations using the sound input.

33. The method of claim 17 , further comprising:

prior to receiving the sound input, receiving a plurality of additional sound inputs, wherein each sound input of the plurality of additional sound inputs includes the one or more trigger words; and

generating the one or more reference representations using the plurality of additional sound inputs.

34. The method of claim 17 , further comprising:

in accordance with determining that the sound input corresponds to the predetermined content, initiating the speech based service.

35. The method of claim 34 , wherein comparing the input representation of the sound input to the one or more reference representations includes determining whether the input representation of the sound input matches the one or more reference representations with a predetermined confidence level, the method further comprising:

in accordance with initiating the speech based service:

in accordance with a determination that the input representation matches the one or more reference representations with the predetermined confidence level:

adjusting the one or more reference representations using the sound input; and

in accordance with a determination that the input representation of the sound input does not match the one or more reference representations with the predetermined confidence level:

forgoing adjusting the one or more reference representations using the sound input.

36. The method of claim 17 , further comprising:

after determining that the sound input does not correspond to the predetermined content, receiving one or more sound inputs consecutive to the sound input, wherein each sound input of the one or more sound inputs includes the one or more trigger words; and

adjusting the one or more reference representations using the one or more sound inputs.

37. The method of claim 36 , wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the one or more sound inputs include at least two successive sound inputs.

38. The method of claim 36 , wherein the one or more sound inputs include at least two sound inputs, and wherein adjusting the one or more reference representations using the one or more sound inputs is performed in accordance with a determination that the at least two sound inputs are received within a predetermined time period.

39. The method of claim 17 , wherein comparing the input representation of the sound input to the one or more reference representations includes comparing the input representation to the one or more reference representations using an algorithm, the method further comprising:

in accordance with detecting the user input to initiate the speech based service, adjusting the algorithm.

40. The method of claim 17 , wherein:

determining whether the sound input corresponds to predetermined content is performed using a first processor; and

adjusting the one or more reference representations using the sound input is performed using a second processor.

41. The method of claim 40 , wherein the second processor is in a standby mode when determining whether the sound input corresponds to predetermined content.

42. The method of claim 40 , further comprising:

storing the sound input in a memory of the electronic device;

in accordance with detecting the user input to initiate the speech based service, activating the second processor, wherein adjusting the one or more reference representations using the sound input includes accessing, by the activated second processor, the sound input.

43. The method of claim 17 , further comprising:

prior to determining whether the sound input corresponds to predetermined content, determining whether the sound input corresponds to a predetermined type of sound.

44. The method of claim 43 , wherein determining whether the sound input corresponds to predetermined content is performed in accordance with determining that the sound input corresponds to the predetermined type of sound.

45. The method of claim 17 , wherein the speech based service includes a digital assistant.

Continuity (6)
Continuation 17150513 · Jan 15, 2021
Continuation 16879348 · May 20, 2020
Continuation 16222249 · Dec 17, 2018
Continuation 14175864 · Feb 7, 2014
Provisional Application 61762260 · Feb 7, 2013
Related Publication 20220230653A1 · Jul 21, 2022
Cited By (1)
US 12,288,165