IP Library Granted Patent US 11,348,582
Granted Patent B2
US 11,348,582 · App. 16/836,226 · Granted May 31, 2022

Electronic devices with voice command and contextual data processing capabilities

Inventor: Aram M. Lindahl (Menlo Park, CA)
Assignee: Apple Inc.
G10L15/22G06F3/167G06F16/43G10L15/1822G10L15/30G10L21/06G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,348,582
App. No.
16/836,226
Granted
May 31, 2022
Kind
B2
Abstract

An electronic device may capture a voice command from a user. The electronic device may store contextual information about the state of the electronic device when the voice command is received. The electronic device may transmit the voice command and the contextual information to computing equipment such as a desktop computer or a remote server. The computing equipment may perform a speech recognition operation on the voice command and may process the contextual information. The computing equipment may respond to the voice command. The computing equipment may also transmit information to the electronic device that allows the electronic device to respond to the voice command.

Claims (59)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by a first electronic device, cause the first electronic device to:

receive a plurality of predefined sample utterances from a first user;

cause speech recognition operations to be trained on the plurality of predefined sample utterances;

associate the trained speech recognition operations with the first user; and

share, with a second electronic device, information related to the trained speech recognition operations associated with the first user, wherein the information is used for processing user utterances received at the second electronic device.

2. The computer-readable storage medium of claim 1 , wherein the trained speech recognition operations are user recognition operations.

3. The computer-readable storage medium of claim 1 , wherein the trained speech recognition operations are stored as part of a user recognition profile.

4. The computer-readable storage medium of claim 1 , wherein a user utterance received at the second electronic device represents a user request.

5. The computer-readable storage medium of claim 4 , wherein the instructions, when executed by the first electronic device, further cause the first electronic device to:

process the user utterance using the speech recognition operations;

determine, based on one or more words in the user utterance, contextual information of the first electronic device that is associated with the user request;

based on the user utterance and the contextual information, cause determination of one or more actions responsive to the user request;

cause performance of the one or more actions to generate a result; and

present the result.

6. The computer-readable storage medium of claim 4 , wherein the result is presented at the second electronic device.

7. The computer-readable storage medium of claim 1 , wherein the instructions, when executed by the first electronic device, further cause the first electronic device to instruct the first user to speak the plurality of predefined sample utterances.

8. The computer-readable storage medium of claim 1 , wherein a third electronic device instructs the first user to speak the plurality of predefined sample utterances.

9. The computer-readable storage medium of claim 1 , wherein the second electronic device instructs the first user to speak the plurality of predefined sample utterances.

10. A method, comprising:

at a first electronic device with one or more processors and memory:

receiving a plurality of predefined sample utterances from a first user;

causing speech recognition operations to be trained on the plurality of predefined sample utterances;

associating the trained speech recognition operations with the first user; and

sharing, with a second electronic device; information related to the trained speech recognition operations associated with the first user, wherein the information is used for processing user utterances received at the second electronic device.

11. The method of claim 10 , wherein the trailed speech recognition operations are user recognition operations.

12. The method of claim 10 , wherein the trained speech recognition operations are stored as part of a user recognition profile.

13. The method of claim 10 , wherein a user utterance received at the second electronic device represents a user request.

14. The method of claim 13 , further comprising:

processing the user utterance using the speech recognition operations;

determining, based on one or more words in the user utterance, contextual information of the first electronic device that is associated with the user request;

based on the user utterance and the contextual information; causing determination of one or more actions responsive to the user request;

causing performance of the one or more actions to generate a result; and

presenting the result.

15. The method of claim 14 , wherein the result is presented at the second electronic device.

16. The method of claim 10 , further comprising:

instructing the first user to speak the plurality of predefined sample utterances.

17. The method of claim 10 , wherein a third electronic device instructs the first user to speak the plurality of predefined sample utterances.

18. The method of claim 10 , wherein the second electronic device instructs the first user to speak the plurality of predefined sample utterances.

19. A first electronic device, comprising:

a microphone;

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a plurality of predefined sample utterances from a first user;

causing speech recognition operations to be trained on the plurality of predefined sample utterances;

associating the trained speech recognition operations with the first user; and

sharing, with a second electronic device, information related to the trained speech recognition operations associated with the first user, wherein the information is used for processing user utterances received at the second electronic device.

20. The first electronic device of claim 19 , wherein the trained speech recognition operations are user recognition operations.

21. The first electronic device of claim 19 , wherein the trained speech recognition operations are stored as part of a user recognition profile.

22. The first electronic device of claim 19 , wherein a user utterance received at the second electronic device represents a user request.

23. The first electronic device of claim 22 , wherein the one or more programs further include instructions for:

processing the user utterance using the speech recognition operations;

determining, based on one or more words in the user utterance, contextual information of the first electronic device that is associated with the user request;

based on the user utterance and the contextual information, causing determination of one or more actions responsive to the user request;

causing performance of the one or more actions to generate a result; and

presenting the result.

24. The first electronic device of claim 23 , wherein the result is presented at the second electronic device.

25. The first electronic device of claim 19 , wherein the one or more programs further include instructions for instructing the first user to speak the plurality of predefined sample utterances.

26. The first electronic device of claim 19 , wherein a third electronic device instructs the first user to speak the plurality of predefined sample utterances.

27. The first electronic device of claim 18 , wherein the second electronic device instructs the first user to speak the plurality of predefined sample utterances.

Continuity (5)
Continuation 15938603 · Mar 28, 2018
Continuation 15207248 · Jul 11, 2016
Continuation 14165520 · Jan 27, 2014
Continuation 12244713 · Oct 2, 2008
Related Publication 20200227044A1 · Jul 16, 2020
Cited By (20)
US 12,197,817 US 12,200,297 US 12,236,938 US 12,236,952 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,437,747 US 12,477,470 US 12,505,748 US 12,562,172 US 12,567,415 US 12,608,171 US 12,613,621 US 12,620,179 US 12,626,702 US 12,694,791