IP Library › Granted Patent US 12,014,118
Granted Patent B2
US 12,014,118 · App. 17/555,030 · Granted Jun 18, 2024

Multi-modal interfaces having selection disambiguation and text modification capability

Inventors: Thomas R. Gruber (Santa Cruz, CA); Mohammed A. Tayyeb (Cupertino, CA); Ron C. Santos (San Jose, CA); Madhusudan Chinthakunta (Saratoga, CA)
Assignee: Apple Inc.
G06F3/167G06F3/04817G06F3/0482G06F3/0488G06F2203/0381G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,014,118
App. No.
17/555,030
Filed
Dec 17, 2021
Granted
Jun 18, 2024
Kind
B2
Art Unit
3662
USPC
715/727
Abstract

Systems and processes for operating an intelligent automated assistant to perform intelligent list reading are provided. In accordance with one example, a method includes, at an electronic device having one or more processors, receiving a first user input of a first input type, the first user input including a plurality of words; displaying, on the touch-sensitive display, the plurality of words; receiving a second user input of a second input type indicating a selection of a word of the plurality of words, the second input type different than the first input type; receiving a third user input; modifying the selected word based on the third user input to provide a modified one or more words; and displaying, on the touch-sensitive display, the modified one or more words.

Claims (121)

1. A method, comprising:

at an electronic device having one or more processors and a touch-sensitive display:

receiving a first user input of a first input type, the first user input corresponding to a plurality of words;

displaying, on a view of the touch-sensitive display, the plurality of words of the first user input;

receiving a second user input of a second input type indicating an attempted selection of a word of the plurality of words of the first user input, the second input type different than the first input type;

determining that more than one word of the plurality of words of the first user input potentially corresponds to the attempted selection;

simultaneously displaying, on the view of the touch-sensitive display, both the plurality of words and an affordance for each word of the plurality of words of the first user input that potentially corresponds to the attempted selection, wherein each affordance includes a copy of the word of the plurality of words associated with the affordance;

receiving a selection of one of the displayed words of the first user input that potentially corresponds to the attempted selection;

designating the selected one of the displayed words as a selected word of the plurality of words;

receiving a third user input;

modifying the selected word based on the third user input to provide a modified one or more words; and

displaying, on the view of the touch-sensitive display, the modified one or more words.

2. The method of claim 1 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

determining a confidence score for each of the plurality of words; and

displaying the plurality of words based on the confidence score of each word of the plurality of words.

3. The method of claim 2 , wherein determining a confidence score for each of the plurality of words comprises:

for each word of the plurality of words:

determining a speech-to-text score;

determining a natural-language score; and

determining a confidence score based on the speech-to-text score and the natural-language score.

4. The method of claim 1 , wherein the first user input of the first input type is a typed input and wherein the third user input is a speech input.

5. The method of claim 4 , wherein the speech input includes a first command and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises modifying the one or more words based on the first command.

6. The method of claim 5 , wherein the first command is a command to modify capitalization of the selected word.

7. The method of claim 5 , wherein the first command is a command to modify spelling of the selected word.

8. The method of claim 5 , wherein the first command is a command to modify textual effects of the selected word.

9. The method of claim 1 , further comprising:

determining a user intent based on the third user input; and

determining a task associated with the user intent,

wherein modifying the selected word based on the third user input to provide a modified one or more words comprises performing the task.

10. The method of claim 9 , wherein the third user input includes a replacement word and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the replacement word.

11. The method of claim 9 , wherein the third user input includes a reference to a word of the plurality of words.

12. The method of claim 1 , wherein the second user input corresponds to a direction of a user's gaze.

13. The method of claim 1 , further comprising:

determining whether the selected word requires disambiguation;

in accordance with a determination that the selected word requires disambiguation, determining a plurality of candidate words; and

providing the plurality of candidate words,

wherein receiving the third user input comprises receiving a selection of a candidate word of the plurality of candidate words,

wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the selected candidate word.

14. The method of claim 13 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

displaying the plurality of words using a word processing application, a messaging application, an email application, or any combination thereof.

15. The method of claim 1 , wherein the first input includes a second command.

16. The method of claim 1 , wherein the second user input of the second input type is a speech input.

17. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a first user input of a first input type, the first user input corresponding to a plurality of words;

display, on a view of a touch-sensitive display, the plurality of words of the first user input;

receive a second user input of a second input type indicating an attempted selection of a word of the plurality of words of the first user input, the second input type different than the first input type;

determine that more than one word of the plurality of words of the first user input potentially corresponds to the attempted selection;

simultaneously display, on the view of the touch-sensitive display, both the plurality of words and an affordance for each word of the plurality of words of the first user input that potentially corresponds to the attempted selection, wherein each affordance includes a copy of the word of the plurality of words associated with the affordance;

receive a selection of one of the displayed words of the first user input that potentially corresponds to the attempted selection;

designate the selected one of the displayed words as a selected word of the plurality of words;

receive a third user input;

modify the selected word based on the third user input to provide a modified one or more words; and

display, on the view of the touch-sensitive display, the modified one or more words.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the second user input of the second input type is a speech input.

19. The non-transitory computer-readable storage medium of claim 17 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

determining a confidence score for each of the plurality of words; and

displaying the plurality of words based on the confidence score of each word of the plurality of words.

20. The non-transitory computer readable storage medium of claim 19 , wherein determining a confidence score for each of the plurality of words comprises:

for each word of the plurality of words:

determining a speech-to-text score;

determining a natural-language score; and

determining a confidence score based on the speech-to-text score and the natural-language score.

21. The non-transitory computer readable storage medium of claim 17 , wherein the first user input of the first input type is a typed input and wherein the third user input is a speech input.

22. The non-transitory computer readable storage medium of claim 21 , wherein the speech input includes a first command and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises modifying the one or more words based on the first command.

23. The non-transitory computer readable storage medium of claim 21 , wherein the first command is a command to modify capitalization of the selected word.

24. The non-transitory computer readable storage medium of claim 21 , wherein the first command is a command to modify spelling of the selected word.

25. The non-transitory computer readable storage medium of claim 21 , wherein the first command is a command to modify textual effects of the selected word.

26. The non-transitory computer readable storage medium of claim 17 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors of the electronic device, cause the electronic device to:

determine a user intent based on the third user input; and

determine a task associated with the user intent, wherein modifying the selected word based on the third user input to provide a modified one or more words comprises performing the task.

27. The non-transitory computer readable storage medium of claim 26 , wherein the third user input includes a replacement word and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the replacement word.

28. The non-transitory computer readable storage medium of claim 26 , wherein the third user input includes a reference to a word of the plurality of words.

29. The non-transitory computer readable storage medium of claim 17 , wherein the second user input corresponds to a direction of a user's gaze.

30. The non-transitory computer readable storage medium of claim 17 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine whether the selected word requires disambiguation;

in accordance with a determination that the selected word requires disambiguation, determine a plurality of candidate words; and

provide the plurality of candidate words, wherein receiving the third user input comprises receiving a selection of a candidate word of the plurality of candidate words, and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the selected candidate word.

31. The non-transitory computer readable storage medium of claim 30 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

displaying the plurality of words using a word processing application, a messaging application, an email application, or any combination thereof.

32. The non-transitory computer readable storage medium of claim 17 , wherein the first input includes a second command.

33. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a first user input of a first input type, the first user input corresponding to a plurality of words;

displaying, on a view of a touch-sensitive display, the plurality of words of the first user input;

receiving a second user input of a second input type indicating an attempted selection of a word of the plurality of words of the first user input, the second input type different than the first input type;

determining that more than one word of the plurality of words of the first user input potentially corresponds to the attempted selection;

simultaneously displaying, on the view of the touch-sensitive display, both the plurality of words and an affordance for each word of the plurality of words of the first user input that potentially corresponds to the attempted selection, wherein each affordance includes a copy of the word of the plurality of words associated with the affordance;

receiving a selection of one of the displayed words of the first user input that potentially corresponds to the attempted selection;

designating the selected one of the displayed words as a selected word of the plurality of words;

receiving a third user input;

modifying the selected word based on the third user input to provide a modified one or more words; and

displaying, on the view of the touch-sensitive display, the modified one or more words.

34. The electronic device of claim 33 , wherein the second user input of the second input type is a speech input.

35. The electronic device of claim 33 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

determining a confidence score for each of the plurality of words; and

displaying the plurality of words based on the confidence score of each word of the plurality of words.

36. The electronic device of claim 35 , wherein determining a confidence score for each of the plurality of words comprises:

for each word of the plurality of words:

determining a speech-to-text score;

determining a natural-language score; and

determining a confidence score based on the speech-to-text score and the natural-language score.

37. The electronic device of claim 33 , wherein the first user input of the first input type is a typed input and wherein the third user input is a speech input.

38. The electronic device of claim 37 , wherein the speech input includes a first command and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises modifying the one or more words based on the first command.

39. The electronic device of claim 38 , wherein the first command is a command to modify capitalization of the selected word.

40. The electronic device of claim 38 , wherein the first command is a command to modify spelling of the selected word.

41. The electronic device of claim 38 , wherein the first command is a command to modify textual effects of the selected word.

42. The electronic device of claim 33 , wherein the one or more programs further include instructions for:

determining a user intent based on the third user input; and

determining a task associated with the user intent, wherein modifying the selected word based on the third user input to provide a modified one or more words comprises performing the task.

43. The electronic device of claim 42 , wherein the third user input includes a replacement word and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the replacement word.

44. The electronic device of claim 42 , wherein the third user input includes a reference to a word of the plurality of words.

45. The electronic device of claim 33 , wherein the second user input corresponds to a direction of a user's gaze.

46. The electronic device of claim 33 , wherein the one or more programs further include instructions for:

determining whether the selected word requires disambiguation;

in accordance with a determination that the selected word requires disambiguation, determining a plurality of candidate words; and

providing the plurality of candidate words, wherein receiving the third user input comprises receiving a selection of a candidate word of the plurality of candidate words, and wherein modifying the selected word based on the third user input to provide a modified one or more words comprises replacing the selected word with the selected candidate word.

47. The electronic device of claim 46 , wherein displaying, on the view of the touch-sensitive display, the plurality of words comprises:

displaying the plurality of words using a word processing application, a messaging application, an email application, or any combination thereof.

48. The electronic device of claim 33 , wherein the first input includes a second command.

Continuity (3)
Continuation 15917410 · Mar 9, 2018
Provisional Application 62506571 · May 15, 2017
Related Publication 20220107780A1 · Apr 7, 2022
Cited By (31)
US 12,197,699 US 12,223,228 US 12,224,972 US 12,238,058 US 12,242,702 US 12,244,755 US 12,260,059 US 12,262,089 US 12,265,364 US 12,265,696 US 12,267,622 US 12,288,165 US 12,301,979 US 12,302,035 US 12,348,663 US 12,368,946 US 12,379,827 US 12,381,924 US 12,422,976 US 12,432,169 US 12,449,961 US 12,452,389 US 12,504,944 US 12,526,361 US 12,541,338 US 12,563,299 US 12,578,837 US 12,615,491 US 12,620,155 US 12,732,665 US 12,737,049