IP Library Granted Patent US 10,083,688
Granted Patent B2
US 10,083,688 · App. 14/838,331 · Granted Sep 25, 2018

Device voice control for selecting a displayed affordance

Inventors: Philippe P. Piernot (Palo Alto, CA); Justin G. Binder (Oakland, CA)
Assignee: Apple Inc.
G10L15/22G06F3/167G10L15/1822G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,083,688
App. No.
14/838,331
Granted
Sep 25, 2018
Kind
B2
Abstract

Systems and processes for device voice control are provided. An example process includes, at an electronic device, receiving a spoken user input and interpreting the spoken user input to derive a representation of user intent. The process further includes determining whether a task may be identified based on the representation of user intent. In accordance with a determination that a task may be identified based on the representation of user intent, the task is performed, and in accordance with a determination that a task may not be identified based on the representation of user intent, the spoken user input is disambiguated.

Claims (118)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions for voice control of displayed content, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a first spoken user input;

obtain a first text string based on the first spoken user input;

derive a representation of a first user intent based on the first text string, wherein the first user intent is derived based on a degree of match between the first text string and one or more words associated with a first predefined domain;

determine whether a task associated with one or more displayed affordances may be identified based on the representation of the first user intent based on the first text string;

in accordance with a determination that a task may be identified based on the representation of the first user intent based on the first text string, perform the task associated with the one or more displayed affordances; and

in accordance with a determination that a task may not be identified based on the representation of the first user intent:

highlight one or more of the displayed affordances;

receive a second spoken user input corresponding to an affordance of the one or more affordances;

obtain a second text string based on the second spoken user input;

derive a representation of a second user intent based on the second text string, wherein the second user intent is derived based on a degree of match between the second text string and one or more words associated with a second predefined domain;

determine whether a task may be identified based on the representation of the second user intent based on the second text string; and

in accordance with a determination that a task may be identified based on the representation of the second user intent based on the second text string, select the affordance of the one or more affordances.

2. The non-transitory computer-readable storage medium of claim 1 , wherein performing the task includes selecting an affordance of the one or more displayed affordances.

3. The non-transitory computer-readable storage medium of claim 2 , wherein the first spoken user input includes an index of the affordance.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the first spoken user input includes a task and an argument.

5. The non-transitory computer-readable storage medium of claim 1 , wherein performing the task includes performing a task associated with a physical input device of the electronic device.

6. The non-transitory computer-readable storage medium of claim 1 , wherein deriving the representation of the first user intent includes:

identifying a name of each of the one or more displayed affordances; and

providing one or more of the identified names to a server.

7. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions, which when executed by the one or more processors of the electronic device, further cause the electronic device to:

in response to another spoken user input, cease to highlight the highlighted one or more displayed affordances; and

highlight an affordance other than the previously highlighted one or more displayed affordances.

8. The non-transitory computer-readable storage medium of claim 7 , wherein the another spoken user input includes at least one of a name or index of the affordance other than the previously highlighted one or more displayed affordances.

9. The non-transitory computer-readable storage medium of claim 1 , wherein performing the task includes performing a system function.

10. The non-transitory computer-readable storage medium of claim 1 , wherein performing the task includes:

initiating performance of the task, wherein the task is a continuous task;

receiving a third spoken user input; and

in response to receiving the third spoken user input, ceasing performance of the task.

11. The non-transitory computer-readable storage medium of claim 1 , wherein the first spoken user input includes context associated with the one or more displayed affordances and wherein determining whether a task associated with the one or more displayed affordances may be identified based on the representation of the first user intent includes:

identifying the one or more displayed affordances based on the context.

12. The non-transitory computer-readable storage medium of claim 1 , wherein performing the task includes requesting confirmation to perform the task.

13. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions, which when executed by one or more processors of the electronic device, further cause the electronic device to:

receive a third spoken user input; and

in response to receiving the third spoken user input, undo the task.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, which when executed by one or more processors of the electronic device, further cause the electronic device to:

receive a fourth spoken user input; and

in response to receiving the fourth spoken user input, redo the task.

15. A method for voice control of displayed content, comprising:

at an electronic device:

receiving a first spoken user input;

obtaining a first text string based on the first spoken user input;

deriving a representation of a first user intent based on the first text string, wherein the first user intent is derived based on a degree of match between the first text string and one or more words associated with a first predefined domain;

determining whether a task associated with one or more displayed affordances may be identified based on the representation of the first user intent based on the first text string;

in accordance with a determination that a task may be identified based on the representation of the first user intent based on the first text string, performing the task associated with the one or more displayed affordances; and

in accordance with a determination that a task may not be identified based on the representation of the first user intent:

highlighting one or more of the displayed affordances;

receiving a second spoken user input corresponding to an affordance of the one or more affordances;

obtaining a second text string based on the second spoken user input;

derive a representation of a second user intent based on the second text string, wherein the second user intent is derived based on a degree of match between the second text string and one or more words associated with a second predefined domain;

determining whether a task may be identified based on the representation of the second user intent based on the second text string; and

in accordance with a determination that a task may be identified based on the representation of the second user intent based on the second text string, selecting the affordance of the one or more affordances.

16. The method of claim 15 , wherein performing the task includes selecting an affordance of the one or more displayed affordances.

17. The method of claim 16 , wherein the first spoken user input includes an index of the affordance.

18. The method of claim 15 , wherein the first spoken user input includes a task and an argument.

19. The method of claim 15 , wherein performing the task includes performing a task associated with a physical input device of the electronic device.

20. The method of claim 15 , wherein deriving the representation of the first user intent includes:

identifying a name of each of the one or more displayed affordances; and

providing one or more of the identified names to a server.

21. The method of claim 15 , further comprising:

in response to another spoken user input, ceasing to highlight the highlighted one or more displayed affordances; and

highlighting an affordance other than the previously highlighted one or more displayed affordances.

22. The method of claim 21 , wherein the another spoken user input includes at least one of a name or index of the affordance other than the previously highlighted one or more displayed affordances.

23. The method of claim 15 , wherein performing the task includes performing a system function.

24. The method of claim 15 , wherein performing the task includes:

initiating performance of the task, wherein the task is a continuous task;

receiving a third spoken user input; and

in response to receiving the third spoken user input, ceasing performance of the task.

25. The method of claim 15 , wherein the first spoken user input includes context associated with the one or more displayed affordances and wherein determining whether a task associated with the one or more displayed affordances may be identified based on the representation of the first user intent includes:

identifying the one or more displayed affordances based on the context.

26. The method of claim 15 , wherein performing the task includes requesting confirmation to perform the task.

27. The method of claim 15 , further comprising:

receive a third spoken user input; and

in response to receiving the third spoken user input, undo the task.

28. The method of claim 27 , further comprising:

receive a fourth spoken user input; and

in response to receiving the fourth spoken user input, redo the task.

29. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for voice control of displayed content, the instructions for:

receiving a first spoken user input;

obtaining a first text string based on the first spoken user input;

deriving a representation of a first user intent based on the first text string, wherein the first user intent is derived based on a degree of match between the first text string and one or more words associated with a first predefined domain;

determining whether a task associated with one or more displayed affordances may be identified based on the representation of the first user intent based on the first text string;

in accordance with a determination that a task may be identified based on the representation of the first user intent based on the first text string, performing the task associated with the one or more displayed affordances; and

in accordance with a determination that a task may not be identified based on the representation of the first user intent:

highlighting one or more of the displayed affordances;

receiving a second spoken user input corresponding to an affordance of the one or more affordances;

obtaining a second text string based on the second spoken user input;

derive a representation of a second user intent based on the second text string, wherein the second user intent is derived based on a degree of match between the second text string and one or more words associated with a second predefined domain;

determining whether a task may be identified based on the representation of the second user intent based on the second text string; and

in accordance with a determination that a task may be identified based on the representation of the second user intent based on the second text string, selecting the affordance of the one or more affordances.

30. The electronic device of claim 29 , wherein performing the task includes selecting an affordance of the one or more displayed affordances.

31. The electronic device of claim 30 , wherein the first spoken user input includes an index of the affordance.

32. The electronic device of claim 29 , wherein the first spoken user input includes a task and an argument.

33. The electronic device of claim 29 , wherein performing the task includes performing a task associated with a physical input device of the electronic device.

34. The electronic device of claim 29 , wherein deriving the representation of the first user intent includes:

identifying a name of each of the one or more displayed affordances; and

providing one or more of the identified names to a server.

35. The electronic device of claim 29 , the one or more programs further including instructions for:

in response to another spoken user input, ceasing to highlight the highlighted one or more displayed affordances; and

highlighting an affordance other than the previously highlighted one or more displayed affordances.

36. The electronic device of claim 35 , wherein the another spoken user input includes at least one of a name or index of the affordance other than the previously highlighted one or more displayed affordances.

37. The electronic device of claim 29 , wherein performing the task includes performing a system function.

38. The electronic device of claim 29 , wherein performing the task includes:

initiating performance of the task, wherein the task is a continuous task;

receiving a third spoken user input; and

in response to receiving the third spoken user input, ceasing performance of the task.

39. The electronic device of claim 29 , wherein the first spoken user input includes context associated with the one or more displayed affordances and wherein determining whether a task associated with the one or more displayed affordances may be identified based on the representation of the first user intent includes:

identifying the one or more displayed affordances based on the context.

40. The electronic device of claim 29 , wherein performing the task includes requesting confirmation to perform the task.

41. The electronic device of claim 29 , the one or more programs further including instructions for:

receive a third spoken user input; and

in response to receiving the third spoken user input, undo the task.

42. The electronic device of claim 41 , the one or more programs further including instructions for:

receive a fourth spoken user input; and

in response to receiving the fourth spoken user input, redo the task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2015
From: PIERNOT, PHILIPPE P.; BINDER, JUSTIN G.
To: APPLE INC.
Reel/Frame 036741/0137 →
Continuity (2)
Provisional Application 62167190 · May 27, 2015
Related Publication 20160351190A1 · Dec 1, 2016
Cited By (29)
US 12,190,873 US 12,197,817 US 12,200,297 US 12,211,502 US 12,216,894 US 12,219,314 US 12,236,952 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,307,505 US 12,327,558 US 12,333,404 US 12,361,943 US 12,367,879 US 12,380,876 US 12,386,434 US 12,386,491 US 12,423,917 US 12,477,470 US 12,567,415 US 12,608,171 US 12,613,621 US 12,619,452 US 12,620,179 US 12,640,151 US 12,646,510 US 12,675,839