IP Library Granted Patent US 12,236,938
Granted Patent B2
US 12,236,938 · App. 18/414,321 · Granted Feb 25, 2025

Digital assistant for providing and modifying an output of an electronic document

Inventors: Daniel A. Castellani (San Jose, CA); Didier Guzzoni (Mont-sur-Rolle, CH); Pierre-Louis Jallerat (Saint-Cyr-en-Val, FR)
Assignee: Apple Inc.
G10L13/10G10L13/0335G10L2013/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,938
App. No.
18/414,321
Granted
Feb 25, 2025
Kind
B2
Abstract

Systems and processes for providing and modifying an output of an electronic document using a digital assistant of an electronic device are provided. An example method includes, receiving, at a first electronic device, a user input requesting an audible output of an electronic document including text; and in accordance with a determination to provide the audible output of the electronic document: generating a media item based on the text of the electronic document; after generating the media item, outputting, based on a semantic structure of the electronic document, the media item; while outputting the media item, receiving a second user input; and in accordance with a determination that the second user input is associated with an intent to modify the output: modifying, based on the second user input and the semantic structure of the electronic document, the output of the media item.

Claims (116)

1. A first electronic device, comprising:

a display;

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors,

the one or more programs including instructions for:

receiving, at the first electronic device, a user input requesting an audible output of an electronic document including text, wherein the user input is the first audio input;

invoking a digital assistant;

determining, based on the digital assistant, to provide the audible output of the electronic document;

in accordance with a determination to provide the audible output of the electronic document:

generating a media item based on the text of the electronic document; and

after generating the media item, audibly outputting the media item, wherein the audible output of the media item is based on a semantic structure of the electronic document;

while audibly outputting the media item, displaying the electronic document and a graphical representation of the media item concurrently;

receiving a second user input, wherein the second user input is an audio input;

in accordance with a determination that the second user input is associated with an intent to modify the audible output of the media item:

modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item;

while audibly outputting the media item, receiving a user input corresponding to a request to cease display of the electronic document; and

in response to receiving the user input corresponding to the request to cease display of the electronic document:

ceasing to display the electronic document while continuing to audibly output the media item; and

continuing to display the graphical representation of the media item.

2. The first electronic device of claim 1 , wherein the semantic structure of the electronic document includes a length of at least one currently outputting semantic object in the electronic document.

3. The first electronic device of claim 1 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item further includes outputting based on a type of semantic object currently being outputted.

4. The first electronic device of claim 1 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes determining that a length of a currently outputting semantic object is greater than a threshold.

5. The first electronic device of claim 4 , wherein the one or more programs further include instructions for:

in accordance with the determination that the length of the currently outputting semantic object is greater than the threshold:

extending a pause after the currently outputting semantic object.

6. The first electronic device of claim 1 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes visually indicating each semantic object that is outputted.

7. The first electronic device of claim 6 , wherein visually indicating each semantic object comprises highlighting all outputted semantic objects.

8. The first electronic device of claim 1 , wherein the electronic document is an article saved on a user's reading list.

9. The first electronic device of claim 1 , wherein the electronic document is displayed on a browser.

10. The first electronic device of claim 1 , wherein the one or more programs further include instructions for:

determining whether the electronic document is readable;

in accordance with a determination that the electronic document is not readable, providing a response indicating the electronic document is not readable; and

in accordance with a determination that the electronic document is readable, determining to provide an output of the electronic document.

11. The first electronic device of claim 1 , wherein the one or more programs further include instructions for:

while audibly outputting the media item, receiving a fourth input associated with an intent to open an application; and

in accordance with a determination that the fourth input is associated with the intent to open the application: opening, while continuing to audibly output the media item, the application.

12. The first electronic device of claim 1 , wherein the audible output of the media item is performed via streaming the electronic document to the media item.

13. The first electronic device of claim 12 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a pitch of the output.

14. The first electronic device of claim 12 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a volume of the output.

15. The first electronic device of claim 12 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a speed of the output.

16. The first electronic device of claim 12 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes recommencing the audible output at a location in the electronic document.

17. The first electronic device of claim 10 , wherein determining whether the electronic document is readable is based on whether the electronic document is compatible with a reader mode of an application.

18. The first electronic device of claim 1 , wherein the user input corresponding to the request to cease display of the electronic document includes a third audio input.

19. The first electronic device of claim 1 , wherein the user input corresponding to the request to cease display of the electronic document includes a gesture input.

20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first electronic device with a display, cause the first electronic device to:

receive, at the first electronic device, a user input requesting an audible output of an electronic document including text, wherein the user input is a first audio input; invoke a digital assistant;

determine, based on the digital assistant, to provide the audible output of the electronic document;

in accordance with a determination to provide the audible output of the electronic document: generate a media item based on the text of the electronic document;

after generating the media item, audibly output the media item, wherein the audible output of the media item is based on a semantic structure of the electronic document;

while audibly outputting the media item, display the electronic document and a graphical representation of the media item concurrently;

receive a second user input, wherein the second user input is a second audio input;

in accordance with a determination that the second user input is associated with an intent to modify the audible output of the media item:

modify, based on the second user input and the semantic structure of the electronic document, the output of the media item;

while audibly outputting the media item, receive a user input corresponding to a request to cease display of the electronic document; and

in response to receiving the user input corresponding to the request to cease display of the electronic document:

cease to display the electronic document while continuing to audibly output the media item; and

continue to display the graphical representation of the media item.

21. The non-transitory computer-readable storage medium of claim 20 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item further includes outputting based on a type of semantic object currently being outputted.

22. The non-transitory computer-readable storage medium of claim 20 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes determining that a length of a currently outputting semantic object is greater than a threshold.

23. The non-transitory computer-readable storage medium of claim 22 , wherein the one or more programs further comprising instructions, which when executed by the one or more processors of the first electronic device, cause the first electronic device to:

in accordance with the determination that the length of the currently outputting semantic object is greater than the threshold:

extend a pause after the currently outputting semantic object.

24. The non-transitory computer-readable storage medium of claim 20 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes visually indicating each semantic object that is outputted.

25. The non-transitory computer-readable storage medium of claim 24 , wherein visually indicating each semantic object comprises highlighting all outputted semantic objects.

26. The non-transitory computer-readable storage medium of claim 20 , wherein the electronic document is an article saved on a user's reading list.

27. The non-transitory computer-readable storage medium of claim 20 , wherein the one or more programs further comprise instructions, which when executed by the one or more programs, cause the first electronic device to:

determine whether the electronic document is readable;

in accordance with a determination that the electronic document is not readable, provide a response indicating the electronic document is not readable; and

in accordance with a determination that the electronic document is readable, determine to provide an output of the electronic document.

28. The non-transitory computer-readable storage medium of claim 27 , wherein determining whether the electronic document is readable is based on whether the electronic document is compatible with a reader mode of an application.

29. The non-transitory computer-readable storage medium of claim 20 , wherein the one or more programs further comprise instructions, which when executed by the one or more programs, cause the first electronic device to:

while audibly outputting the media item, receive a fourth input associated with an intent to open an application; and

in accordance with a determination that the fourth input is associated with the intent to open the application:

open, while continuing to audibly output the media item, the application.

30. The non-transitory computer-readable storage medium of claim 20 , wherein the audible output of the media item is performed via streaming the electronic document to the media item.

31. The non-transitory computer-readable storage medium of claim 30 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a pitch of the output.

32. The non-transitory computer-readable storage medium of claim 30 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a volume of the output.

33. The non-transitory computer-readable storage medium of claim 30 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a speed of the output.

34. The non-transitory computer-readable storage medium of claim 30 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes recommencing the audible output at a location in the electronic document.

35. A method, comprising:

at a first electronic device with a display, one or more processors, and memory:

receiving, at the first electronic device, a user input requesting an audible output of an electronic document including text, wherein the user input is a first audio input;

invoking a digital assistant;

determining, based on the digital assistant, to provide the audible output of the electronic document;

in accordance with a determination to provide the audible output of the electronic document: generating a media item based on the text of the electronic document; and

after generating the media item, audibly outputting the media item, wherein the audible output of the media item is based on a semantic structure of the electronic document;

while audibly outputting the media item, displaying the electronic document and a graphical representation of the media item concurrently;

receiving a second user input, wherein the second user input is a second audio input;

in accordance with a determination that the second user input is associated with an intent to modify the audible output of the media item: modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item;

while audibly outputting the media item, receiving a user input corresponding to a request to cease display of the electronic document; and

in response to receiving the user input corresponding to the request to cease display of the electronic document:

ceasing to display the electronic document while continuing to audibly output the media item; and

continuing to display the graphical representation of the media item.

36. The method of claim 35 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item further includes outputting based on a type of semantic object currently being outputted.

37. The method of claim 35 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes determining that a length of a currently outputting semantic object is greater than a threshold.

38. The method of claim 37 , further comprising:

in accordance with the determination that the length of the currently outputting semantic object is greater than the threshold:

extending a pause after the currently outputting semantic object.

39. The method of claim 35 , wherein audibly outputting, based on the semantic structure of the electronic document, the media item includes visually indicating each semantic object that is outputted.

40. The method of claim 39 , wherein visually indicating each semantic object comprises highlighting all outputted semantic objects.

41. The method of claim 35 , wherein the electronic document is an article saved on a user's reading list.

42. The method of claim 35 , further comprising:

determining whether the electronic document is readable;

in accordance with a determination that the electronic document is not readable, providing a response indicating the electronic document is not readable; and

in accordance with a determination that the electronic document is readable, determining to provide an output of the electronic document.

43. The method of claim 42 , wherein determining whether the electronic document is readable is based on whether the electronic document is compatible with a reader mode of an application.

44. The method of claim 35 , further comprising:

while audibly outputting the media item, receiving a fourth input associated with an intent to open an application; and

in accordance with a determination that the fourth input is associated with the intent to open the application:

opening, while continuing to audibly output the media item, the application.

45. The method of claim 35 , wherein the audible output of the media item is performed via streaming the electronic document to the media item.

46. The method of claim 45 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a pitch of the output.

47. The method of claim 45 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a volume of the output.

48. The method of claim 45 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes modifying a speed of the output.

49. The method of claim 45 , wherein modifying, based on the second user input and the semantic structure of the electronic document, the audible output of the media item includes recommencing the audible output at a location in the electronic document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: JALLERAT, PIERRE-LOUIS; CASTELLANI, DANIEL A.; GUZZONI, DIDIER
To: APPLE INC.
Reel/Frame 066503/0577 →
Continuity (2)
Provisional Application 63459580 · Apr 14, 2023
Related Publication 20240347041A1 · Oct 17, 2024
References Cited (63)
US 6033224A · Kurzweil · 2000 [cited by examiner]
US 6250928B1 · Poggio · 2001 [cited by examiner]
US 7096183B2 · Junqua · 2006 [cited by applicant]
US 7684991B2 · Stohr et al. · 2010 [cited by applicant]
US 8073695B1 · Hendricks · 2011 [cited by examiner]
US 8832584B1 · Killalea et al. · 2014 [cited by applicant]
US 9002703B1 · Crosley · 2015 [cited by applicant]
US 9087508B1 · Dzik · 2015 [cited by examiner]
US 9286287B1 · Tierney · 2016 [cited by examiner]
US 9606986B2 · Bellegarda · 2017 [cited by applicant]
US 10147421B2 · Liddell et al. · 2018 [cited by applicant]
US 10303715B2 · Graham et al. · 2019 [cited by applicant]
US 10317992B2 · Prokofieva et al. · 2019 [cited by applicant]
US 10643611B2 · Lindahl · 2020 [cited by applicant]
US 10691473B2 · Karashchuk et al. · 2020 [cited by applicant]
US 10747498B2 · Stasior et al. · 2020 [cited by applicant]
US 10904488B1 · Weisz et al. · 2021 [cited by applicant]
US 10909171B2 · Graham et al. · 2021 [cited by applicant]
US 11087759B2 · Lemay et al. · 2021 [cited by applicant]
US 11152002B2 · Walker et al. · 2021 [cited by applicant]
US 11348582B2 · Lindahl · 2022 [cited by applicant]
US 11388291B2 · Van Os et al. · 2022 [cited by applicant]
US 11704552B2 · Sim et al. · 2023 [cited by applicant]
US 20030212559A1 · Xie · 2003 [cited by examiner]
US 20060008122A1 · Kurzweil · 2006 [cited by examiner]
US 20090259475A1 · Yamagami et al. · 2009 [cited by applicant]
US 20110119590A1 · Seshadri · 2011 [cited by examiner]
US 20110167287A1 · Walsh · 2011 [cited by examiner]
US 20110177481A1 · Haff · 2011 [cited by examiner]
US 20120278082A1 · Borodin · 2012 [cited by examiner]
US 20130307855A1 · Lamb et al. · 2013 [cited by applicant]
US 20140108010A1 · Maltseff · 2014 [cited by examiner]
US 20140283111A1 · Dolph et al. · 2014 [cited by applicant]
US 20150066506A1 · Romano · 2015 [cited by examiner]
US 20150163610A1 · Sampat · 2015 [cited by examiner]
US 20160027431A1 · Kurzweil · 2016 [cited by examiner]
US 20160071510A1 · Li et al. · 2016 [cited by applicant]
US 20160171980A1 · Liddell et al. · 2016 [cited by applicant]
US 20180035155A1 · Garner · 2018 [cited by examiner]
US 20180108356A1 · Mizumoto · 2018 [cited by examiner]
US 20180285063A1 · Huynh · 2018 [cited by examiner]
US 20200143330A1 · Perumalla et al. · 2020 [cited by applicant]
US 20200159838A1 · Kikin-Gil · 2020 [cited by examiner]
US 20200365134A1 · Tu · 2020 [cited by examiner]
US 20210097134A1 · Livshits · 2021 [cited by examiner]
US 20220093191A1 · Yang · 2022 [cited by applicant]
US 20220291792A1 · Alston · 2022 [cited by examiner]
US 20230352014A1 · Tennant et al. · 2023 [cited by applicant]
EP 1291847A2 · 2003 [cited by applicant]
JP 2018523102A · 2018 [cited by applicant]
KR 1020130075783A · 2013 [cited by applicant]
WO 2012051052A1 · 2012 [cited by applicant]
WO 2016191737A2 · 2016 [cited by applicant]
“Nuance Dragon Naturally Speaking”, Version 13 End-User Workbook, Nuance Communications, Inc. Online Available at: https://www.nuance.com/content/dam/nuance/en_us/collateral/dragon/guide/gd-dragon-naturally-speaking-13-… [cited by applicant]
Guo et al., “VizLens: A Robust and Interactive Screen Reader for Interfaces in the Real World”, In Proceedings of the 29th Annual Symposium on User Interface Software and Technology (UIST '16), Tokyo, Japan, Online avai… [cited by applicant]
Majerus, Wesley, “Cell phone accessibility for your blind child”, Retrieved from the Internet: <URL:https://web.archive.org/web/20100210001100/https://nfb.org/images/nfb/publications/fr/fr28/3/fr280314.htm>, 2010, pp. 1… [cited by applicant]
Zhang et al., “Interaction Proxies for Runtime Repair and Enhancement of Mobile Application Accessibility”, In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI '17). ACM, Denver, CO, USA… [cited by applicant]
Zhao et al., “SeeingVR: A Set of Tools to Make Virtual Reality More Accessible to People with Low Vision”, In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI '19). ACM, Article 111, Gla… [cited by applicant]
Matias, Yossi, “Easier access to web pages: Ask Google Assistant to read it aloud”, Available online at: https://blog.google/products/assistant/easier-access-web-pages-let-assistant-read-it-aloud/, Mar. 4, 2020, 3 pages. [cited by applicant]
Invitation to Pay Additional Fees and Partial International Search Report received for PCT Patent Application No. PCT/US2024/024028, mailed on Jun. 18, 2024, 15 pages. [cited by applicant]
Ermolina et al., “Voice-Controlled Intelligent Personal Assistants in Health Care: International Delphi Study”, Journal of Medical Internet Research, 2021. Online Available at: doi: 10.2196/25312., Apr. 9, 2021, 20 page… [cited by applicant]
Yeh et al., “Dialog modeling in audiobook synthesis”. Retrieved on Sep. 27, 2023. 6 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2024/024028, mailed on Aug. 8, 2024, 24 pages. [cited by applicant]