IP Library Granted Patent US 12,436,619
Granted Patent B1
US 12,436,619 · App. 17/376,072 · Granted Oct 7, 2025

Devices and methods for digital assistant

Inventors: Pravalika Avvaru (Milpitas, CA); Novaira Masood (San Jose, CA); Jeremy Ray Bernstein (San Francisco, CA); Eric Geusz (San Francisco, CA)
Assignee: Apple Inc.
G06F3/017G06F3/013G06F3/167G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,619
App. No.
17/376,072
Filed
Jul 14, 2021
Granted
Oct 7, 2025
Kind
B1
Art Unit
2624
USPC
345/156
Abstract

A digital assistant may be used with a computer-generated environment. In some embodiments, a digital assistant is invoked and/or terminated via user input. In some embodiments, the digital assistant is displaying in a computer-generated environment. In some embodiments, the digital assistant may be used to interact with objects in a computer-generated environment. Devices, methods, and graphical user interfaces for a digital assistant provide an improved user experience for computer-generated environments.

Claims (42)

1. A method comprising:

at an electronic device in communication with a display and one or more input devices comprising at least one hand-tracking sensor:

displaying, via the display, a three-dimensional environment;

while displaying the three-dimensional environment, detecting, via the at least one hand-tracking sensor, a first input including a gesture by a hand; and

in response to the first input:

in accordance with a determination that the first input satisfies first criteria, the first criteria including a first criterion that is satisfied when the hand is oriented in a specified direction relative to the electronic device or within a threshold of the specified direction relative to the electronic device, when a dorsal aspect of the hand is facing the electronic device and including a second criterion that is satisfied when a gaze is detected at or within a threshold distance of the hand or a representation of the hand in the three-dimensional environment, activating a digital assistant and displaying a visual representation of the digital assistant in the three-dimensional environment, wherein;

the digital assistant interprets natural language input and performs one or more actions based on the natural language input; and

in accordance with a determination that the digital assistant is activated:

 obtaining audio data; and

 changing the visual representation of the digital assistant based on the audio data while obtaining the audio data; and

in accordance with a determination that the first input fails to satisfy the first criteria, forgoing activating the digital assistant.

2. The method of claim 1 , further comprising: while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, a second input; and in response to the second input: in accordance with a determination that the second input satisfies one or more second criteria, ceasing displaying of the visual representation of the digital assistant; and in accordance with a determination that the second input fails to satisfy the one or more second criteria, continuing displaying the visual representation of the digital assistant.

3. The method of claim 1 , wherein the visual representation of the digital assistant in the three-dimensional environment is anchored to a representation of the hand in the three-dimensional environment.

4. The method of claim 1 , wherein the first criteria include one or more of: a criterion that is satisfied when the hand corresponds to a predetermined hand; a criterion that is satisfied when the hand is in a predetermined pose; a criterion that is satisfied when the hand is orientated in a specified direction or within a threshold of the specified direction; or a criterion that is satisfied when the hand is within a field of view of a sensor of the electronic device.

5. The method of claim 1 , wherein the first criteria include a criterion that is satisfied when a gaze is detected at or within a threshold distance of the hand or a representation of the hand in the three-dimensional environment.

6. The method of claim 1 , wherein the first criteria include a criterion that is satisfied when a subset of the first criteria are satisfied for a threshold period of time.

7. The method of claim 1 , further comprising: while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, audio data; and in response to the audio data: displaying a representation of the audio data in the three-dimensional environment.

8. An electronic device comprising: a display; one or more input devices comprising at least one hand-tracking sensor; one or more processors; memory; and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display, a three-dimensional environment; while displaying the three-dimensional environment, detecting, via the at least one hand-tracking sensor a first input including a gesture by a hand; and in response to the first input: in accordance with a determination that the first input satisfies first criteria, the first criteria including a criterion that is satisfied when the hand is oriented in a specified direction relative to the electronic device or within a threshold of the specified direction relative to the electronic device, when a dorsal aspect of the hand is facing the electronic device and a criterion that is satisfied when a gaze is detected at or within a threshold distance of the hand or a representation of the hand in the three-dimensional environment, activating a digital assistant and displaying a visual representation of the digital assistant in the three-dimensional environment, wherein:

the digital assistant interprets natural language input and performs one or more actions based on the natural language input; and

in accordance with a determination that a digital assistant is activated:

obtaining audio data; and

changing the visual representation of the digital assistant based on the audio data while obtaining the audio data; and

in accordance with a determination that the first input fails to satisfy the first criteria, forgoing activating the digital assistant.

9. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform a method comprising:

displaying, via a display, a three-dimensional environment; while displaying the three-dimensional environment, detect, via one or more input devices comprising at least one hand-tracking sensor, a first input including a gesture by a hand; and in response to the first input: in accordance with a determination that the first input satisfies first criteria, the first criteria including a criterion that is satisfied when the hand is oriented in a specified direction relative to the electronic device or within a threshold of the specified direction relative to the electronic device, when a dorsal aspect of the hand is facing the electronic device and a criterion that is satisfied when a gaze is detected at or within a threshold distance of the hand or a representation of the hand in the three-dimensional environment, activating a digital assistant and displaying a visual representation of the digital assistant in the three-dimensional environment, wherein:

the digital assistant interprets natural language input and performs one or more actions based on the natural language input; and

in accordance with a determination that a digital assistant is activated:

obtaining audio data; and

changing the visual representation of the digital assistant based on the audio data while obtaining the audio data; and

in accordance with a determination that the first input fails to satisfy the first criteria, forgoing activating the digital assistant.

10. The electronic device of claim 8 , the one or more programs further including instructions for: while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, a second input; and in response to the second input: in accordance with a determination that the second input satisfies one or more second criteria, ceasing displaying of the visual representation of the digital assistant; and in accordance with a determination that the second input fails to satisfy the one or more second criteria, continuing displaying the visual representation of the digital assistant.

11. The electronic device of claim 8 , wherein the visual representation of the digital assistant in the three-dimensional environment is anchored to a representation of the hand in the three-dimensional environment.

12. The electronic device of claim 8 , wherein the digital assistant interprets natural language input and performs one or more actions based on the natural language input.

13. The electronic device of claim 8 , wherein the first criteria include one or more of: a criterion that is satisfied when the hand corresponds to a predetermined hand; a criterion that is satisfied when the hand is in a predetermined pose; a criterion that is satisfied when a gaze is detected at or within a threshold distance of the hand or a representation of the hand in the three-dimensional environment; or a criterion that is satisfied when the hand is within a field of view of a sensor of the electronic device.

14. The electronic device of claim 8 , the one or more programs further including instructions for: in response to the first input: in accordance with the determination that the first input satisfies the first criteria, obtaining audio data via an audio sensor; while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, a second input; and in response to the second input: in accordance with a determination that the second input satisfies one or more second criteria, ceasing obtaining additional audio data; and in accordance with a determination that the second input fails to satisfy the one or more second criteria, continuing obtaining the audio data.

15. The electronic device of claim 14 , wherein the one or more second criteria include one or more of: a criterion that is satisfied when the hand is in a predetermined pose; a criterion that is satisfied when the hand is not oriented in a specified direction or within a threshold of the specified direction; a criterion that is satisfied when the audio data indicate a user ceases to speak for a threshold period of time; or a criterion that is satisfied when the audio data indicate a predefined audio command.

16. The electronic device of claim 14 , the one or more programs further including instructions for: in response to the first input: in accordance with the determination that the first input satisfies the first criteria, obtaining audio data via an audio sensor while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, a second input; and in response to the second input: in accordance with a determination that the second input satisfies one or more second criteria, executing a command in accordance with the audio data; and in accordance with a determination that the second input fails to satisfy the one or more second criteria, continuing obtaining the audio data.

17. The electronic device of claim 16 , wherein the command comprises adding a new object or manipulating an existing object in the three-dimensional environment, the one or more programs further including instructions for: adding the new object or manipulating the existing object in accordance with executing the command.

18. The non-transitory computer readable storage medium of claim 9 , the method further comprising: while the digital assistant is activated and while displaying the three-dimensional environment, detecting, via the one or more input devices, a second input; and in response to the second input: in accordance with a determination that the second input satisfies one or more second criteria, deactivating the digital assistant; and in accordance with a determination that the second input fails to satisfy the one or more second criteria, continuing obtaining audio data.

19. The non-transitory computer readable storage medium of claim 9 , wherein the first criteria include one or more of: a criterion that is satisfied when the hand corresponds to a predetermined hand; a criterion that is satisfied when the hand is in a predetermined pose; a criterion that is satisfied when the hand is orientated in a specified direction or within a threshold of the specified direction; a criterion that is satisfied when the hand is within a field of view of a sensor of the electronic device; or a criterion that is satisfied when a subset of the first criteria are satisfied for a threshold period of time.

20. The non-transitory computer readable storage medium of claim 9 , wherein the digital assistant interprets natural language input and performs one or more actions based on the natural language input.

21. The method of claim 1 , wherein the first criteria include a criterion that is satisfied when the gesture by the hand includes a predetermined pose, wherein the predetermined pose corresponds to the hand in a fist.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2021
From: AVVARU, PRAVALIKA; MASOOD, NOVAIRA; BERNSTEIN, JEREMY RAY; GEUSZ, ERIC
To: APPLE INC.
Reel/Frame 056868/0757 →
Continuity (1)
Provisional Application 63052412 · Jul 15, 2020
References Cited (57)
US 5758122A · Corda et al. · 1998 [cited by applicant]
US 8489984B1 · Violet et al. · 2013 [cited by applicant]
US 9143770B2 · Bennett et al. · 2015 [cited by applicant]
US 9318108B2 · Gruber et al. · 2016 [cited by applicant]
US 10048748B2 · Sridharan et al. · 2018 [cited by applicant]
US 10452360B1 · Burman et al. · 2019 [cited by applicant]
US 10521195B1 · Swope et al. · 2019 [cited by applicant]
US 11119735B1 · Baafi et al. · 2021 [cited by applicant]
US 11269889B1 · Aversano et al. · 2022 [cited by applicant]
US 20030007005A1 · Kandogan · 2003 [cited by applicant]
US 20030135533A1 · Cook · 2003 [cited by applicant]
US 20070150864A1 · Goh · 2007 [cited by applicant]
US 20070238520A1 · Kacmarcik · 2007 [cited by applicant]
US 20070244847A1 · Au · 2007 [cited by applicant]
US 20080092111A1 · Kinnucan et al. · 2008 [cited by applicant]
US 20120259762A1 · Tarighat et al. · 2012 [cited by applicant]
US 20140072115A1 · Makagon et al. · 2014 [cited by applicant]
US 20140129961A1 · Zubarev et al. · 2014 [cited by applicant]
US 20140287397A1 · Chong et al. · 2014 [cited by applicant]
US 20150077325A1 · Ferens et al. · 2015 [cited by applicant]
US 20150095882A1 · Jaeger et al. · 2015 [cited by applicant]
US 20150130716A1 · Sridharan · 2015 [cited by examiner]
US 20150149912A1 · Moore · 2015 [cited by applicant]
US 20160026253A1 · Bradski et al. · 2016 [cited by applicant]
US 20160269508A1 · Sharma · 2016 [cited by examiner]
US 20160342318A1 · Melchner et al. · 2016 [cited by applicant]
US 20160379418A1 · Osborn · 2016 [cited by examiner]
US 20170052767A1 · Bennett et al. · 2017 [cited by applicant]
US 20170085445A1 · Layman et al. · 2017 [cited by applicant]
US 20170147296A1 · Kumar et al. · 2017 [cited by applicant]
US 20170161123A1 · Zhao et al. · 2017 [cited by applicant]
US 20170255375A1 · Kim · 2017 [cited by applicant]
US 20170255450A1 · Mullins et al. · 2017 [cited by applicant]
US 20170277516A1 · Grebnov et al. · 2017 [cited by applicant]
US 20170315789A1 · Lam et al. · 2017 [cited by applicant]
US 20170316355A1 · Shrestha et al. · 2017 [cited by applicant]
US 20170316363A1 · Siciliano et al. · 2017 [cited by applicant]
US 20180095542A1 · Mallinson · 2018 [cited by examiner]
US 20180107461A1 · Balasubramanian et al. · 2018 [cited by applicant]
US 20180213048A1 · Messner et al. · 2018 [cited by applicant]
US 20180285084A1 · Mimlitch et al. · 2018 [cited by applicant]
US 20180285476A1 · Siciliano et al. · 2018 [cited by applicant]
US 20190005228A1 · Singh et al. · 2019 [cited by applicant]
US 20190065026A1 · Kiemele · 2019 [cited by examiner]
US 20190086997A1 · Kang et al. · 2019 [cited by applicant]
US 20190220863A1 · Novick et al. · 2019 [cited by applicant]
US 20200225758A1 · Tang · 2020 [cited by examiner]
US 20200301678A1 · Burman et al. · 2020 [cited by applicant]
US 20200356350A1 · Penland et al. · 2020 [cited by applicant]
US 20210152966A1 · Mathur · 2021 [cited by examiner]
US 20220101040A1 · Zhang et al. · 2022 [cited by applicant]
WO 2019026357A1 · 2019 [cited by applicant]
Grzyb et al., “Beyond Robotic Speech: Mutual Benefits to Cognitive Psychology and Artificial Intelligence From the Study of Multimodal Communication”, PsyArXiv Preprints, Available online at: <https://psyarxiv.com/h5dxy… [cited by applicant]
Hou et al., “Visual Feedback Design and Research of Handheld Mobile Voice Assistant Interface”, Atlantis Press, Advances in Intelligent Systems Research, vol. 146, Available online at: <https://download.atlantis-press.c… [cited by applicant]
Ravenet et al., “Automating the Production of Communicative Gestures in Embodied Characters”, Frontiers in Psychology, vol. 9, Article 1144, Available online at: <https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6046454/>, … [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/376,062, mailed on Apr. 21, 2023, 27 pages. [cited by applicant]
Apaza-Yllachura, et al., “SimpleAR: Augmented Reality High-Level Content Design Framework Using Visual Programming”, https://ieeexplore.ieee.org/document/8966427, IEEE, Apr. 9, 2019, 7 pages. [cited by applicant]
Cited By (1)
US 12,718,492