IP Library Granted Patent US 12,373,027
Granted Patent B2
US 12,373,027 · App. 18/344,921 · Granted Jul 29, 2025

Gaze initiated actions

Inventors: Rohit Prasad (Lexington, MA); Arushan Rajasekaram (North Bend, WA); Furqan Muhammad Khan (Culver City, CA); Rajiv M Reddy (Bellevue, WA)
Assignee: Amazon Technologies, Inc.
G06F3/013G06F3/0481G06F3/04842G06F3/165G06F3/167G06F9/451G06T7/70G10L15/22G06T2207/30201G10L15/193G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,027
App. No.
18/344,921
Granted
Jul 29, 2025
Kind
B2
Abstract

A system that detects the location of a user gaze at a display and in response to the duration of the gaze exceeding a threshold, auto-playing content on the display. The system may also determine gaze event data associating the gaze event with the source of the content the user is gazing at. Other information may also be associated with the gaze event such as user ID, time/duration data, or the like. Various actions can be taken in response to the gaze event such as auto-playing of content, outputting a visual indication of the detected gaze, interpreting detected speech using the gaze event data, data aggregation, etc.

Claims (103)

1. A computer-implemented method comprising:

presenting, by a display of a first device, a first graphical user interface (GUI) element;

capturing, by a camera of the first device, first image data representing a user;

processing the first image data to determine a gaze of the user is directed, at a first time, at the display;

determining the first GUI element corresponds to video content corresponding to a first application;

based at least in part on the first GUI element corresponding to the video content, selecting a first time threshold from a plurality of time thresholds, the first time threshold corresponding to playing of video content in response to gaze detection;

capturing, by the camera, second image data;

processing the second image data to determine the gaze of the user is directed, at a second time after the first time, at the display;

based at least in part on the first time and the second time, determining a first duration corresponding to the gaze of the user being directed at the display;

determining that the first duration satisfies the first time threshold;

receiving video data corresponding to the video content; and

in response to the first duration satisfying the first time threshold, causing the first device to present the video content using the video data.

2. The computer-implemented method of claim 1 , further comprising:

capturing, by the camera, third image data;

processing the third image data to determine the gaze of the user is directed, at a third time between the first time and the second time, at the display;

determining a second duration between the first time and the third time; and

based at least in part on the second duration, presenting, by the display, a visual indicator representing detection of the gaze of the user.

3. The computer-implemented method of claim 1 , further comprising, prior to processing the second image data:

sending a request for the video data to a second device associated with the first application;

receiving a first portion of the video data; and

storing, by the first device, the first portion of the video data,

wherein causing the first device to present the video content comprises processing the first portion of the video data.

4. A computer-implemented method comprising:

receiving first content to display on a device;

determining that the first content corresponds to a gaze initiated playback function;

based on the first content, determining a first time threshold to be used to initiate gaze related playback of the first content;

presenting, by a display of a first device, a first graphical user interface (GUI) element corresponding to first content;

capturing, by a camera, first image data representing a user;

processing the first image data to determine a first gaze of the user is directed at the display;

determining a first duration corresponding to the first gaze;

determining that the first duration satisfies the first time threshold; and

in response to the first duration satisfying the first time threshold, causing playback of the first content.

5. The computer-implemented method of claim 4 , wherein the first content corresponds to video content comprising image content and audio content and wherein causing playback of the first content comprises:

in response to the first duration satisfying the first time threshold, causing playback of the image content.

6. The computer-implemented method of claim 5 , further comprising:

receiving a user input corresponding to playback of the audio content; and

in response to the user input, causing playback of the audio content.

7. The computer-implemented method of claim 4 , further comprising:

determining user profile data corresponding to the user;

wherein determining the first time threshold is further based at least in part on the user profile data.

8. The computer-implemented method of claim 4 , further comprising:

determining the first GUI element corresponds to a first application;

determining first data associating the first gaze with the first application;

capturing, by a microphone, audio representing an utterance;

determining audio data representing the utterance; and

causing speech processing to be performed using the audio data and the first data.

9. The computer-implemented method of claim 4 , further comprising:

presenting, by the display, a visual indicator representing detection of the first gaze of the user.

10. The computer-implemented method of claim 4 , further comprising:

determining an identifier associated with the user,

wherein presenting the first GUI element is based at least in part on the identifier.

11. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

present, by a display of a first device, a first graphical user interface (GUI) element corresponding to first content;

capture, by a camera, first image data representing a user;

process the first image data to determine a gaze of the user is directed to at the display;

determine a duration corresponding to the gaze;

determine user profile data corresponding to the user;

determine that the user profile data includes information identifying a first time threshold associated with the user;

based at least in part on the user profile data including the information, select the first time threshold for use in evaluating the duration;

determine the duration satisfies the first time threshold; and

in response to the duration satisfying the first time threshold, cause playback of the first content.

12. The system of claim 11 , wherein the first content corresponds to video content comprising image content and audio content and wherein the instructions that cause the system to cause playback of the first content comprise instructions that, when executed by the at least one processor, further cause the system to:

in response to the duration satisfying the first time threshold, causing playback of the image content.

13. The system of claim 12 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive a user input corresponding to playback of the audio content; and

in response to the user input, cause playback of the audio content.

14. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

select the first time threshold further based at least in part on the first content.

15. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine the first GUI element corresponds to a first application;

determine first data associating the gaze with the first application;

capture, by a microphone, audio representing an utterance;

determine audio data representing the utterance; and

cause speech processing to be performed using the audio data and the first data.

16. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

present, by the display, a visual indicator representing detection of the gaze of the user.

17. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine an identifier associated with the user,

wherein presentation of the first GUI element is based at least in part on the identifier.

18. The computer-implemented method of claim 4 , further comprising:

receiving second content to display on a device;

determining that the second content does not correspond to a gaze initiated playback function;

presenting, by the display, a second GUI element corresponding to second content, wherein the first image data was captured while the first GUI element and the second GUI element were presented simultaneously on the display; and

determining to cause playback of the first content rather than the second content based at least in part on the first content corresponding to a gaze initiated playback function and the second content not corresponding to a gaze initiated playback function.

19. The computer-implemented method of claim 4 , further comprising:

receiving second content to display on a device;

determining that the second content corresponds to a gaze initiated playback function;

based on the second content, determining a second time threshold, different than the first time threshold, to be used to initiate gaze related playback of the second content;

presenting, by the display, a second GUI element corresponding to the second content;

capturing, by the camera, second image data representing the user;

processing the second image data to determine a second gaze of the user is directed at the display;

determining a second duration corresponding to the second gaze;

determining that the second duration satisfies the second time threshold; and

in response to the second duration satisfying the second time threshold, causing playback of the second content.

20. The computer-implemented method of claim 19 , wherein:

determining the first time threshold further comprises:

determining that the first content corresponds to a first type of content, and

determining that the first type of content is associated with the first time threshold; and

determining the second time threshold further comprises:

determining that the second content corresponds to a second type of content different than the first type of content, and

determining that the second type of content is associated with the second time threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: PRASAD, ROHIT; RAJASEKARAM, ARUSHAN; KHAN, FURQAN MUHAMMAD; REDDY, RAJIV M
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064122/0470 →
Continuity (1)
Related Publication 20250004544A1 · Jan 2, 2025
References Cited (129)
US 5818439A · Nagasaka et al. · 1998 [cited by applicant]
US 6484189B1 · Gerlach et al. · 2002 [cited by applicant]
US 6724864B1 · Denenberg et al. · 2004 [cited by applicant]
US 6934756B2 · Maes · 2005 [cited by applicant]
US 7117442B1 · Kemble et al. · 2006 [cited by applicant]
US 7194409B2 · Balentine et al. · 2007 [cited by applicant]
US 7324636B2 · Sauvage et al. · 2008 [cited by applicant]
US 7447636B1 · Schwartz et al. · 2008 [cited by applicant]
US 8229750B2 · Creamer et al. · 2012 [cited by applicant]
US 8406457B2 · Matsuoka et al. · 2013 [cited by applicant]
US 8432368B2 · Momeyer et al. · 2013 [cited by applicant]
US 8788977B2 · Bezos · 2014 [cited by applicant]
US 8970625B2 · Chavez et al. · 2015 [cited by applicant]
US 9026438B2 · Buck et al. · 2015 [cited by applicant]
US 9041734B2 · Look et al. · 2015 [cited by applicant]
US 9092435B2 · Douthitt et al. · 2015 [cited by applicant]
US 9094576B1 · Karakotsios · 2015 [cited by applicant]
US 9213163B2 · Lewis et al. · 2015 [cited by applicant]
US 9294670B2 · Jafarzadeh et al. · 2016 [cited by applicant]
US 9317175B1 · Lockhart · 2016 [cited by applicant]
US 9318129B2 · Vasilieff et al. · 2016 [cited by applicant]
US 9319221B1 · Awad et al. · 2016 [cited by applicant]
US 9354709B1 · Heller et al. · 2016 [cited by applicant]
US 9368105B1 · Freed et al. · 2016 [cited by applicant]
US 9384594B2 · Maciocci et al. · 2016 [cited by applicant]
US 9431021B1 · Scalise et al. · 2016 [cited by applicant]
US 9269012B2 · Fotland · 2016 [cited by applicant]
US 9542941B1 · Weksler et al. · 2017 [cited by applicant]
US 9552093B2 · Zawacki et al. · 2017 [cited by applicant]
US 9591295B2 · Lockhart et al. · 2017 [cited by applicant]
US 9626152B2 · Kim et al. · 2017 [cited by applicant]
US 9652031B1 · Savastinuk et al. · 2017 [cited by applicant]
US 9880640B2 · Gray et al. · 2018 [cited by applicant]
US 9940929B2 · VanBlon et al. · 2018 [cited by applicant]
US 10027888B1 · Mackraz · 2018 [cited by applicant]
US 10055013B2 · Ramaswamy et al. · 2018 [cited by applicant]
US 10067634B2 · Ames et al. · 2018 [cited by applicant]
US 10146303B2 · Cerriteno · 2018 [cited by examiner]
US 10320353B1 · Wahlberg et al. · 2019 [cited by applicant]
US 10325409B2 · Costa · 2019 [cited by applicant]
US 10332506B2 · Pappu et al. · 2019 [cited by applicant]
US 10349120B2 · Graham et al. · 2019 [cited by applicant]
US 10366692B1 · Adams et al. · 2019 [cited by applicant]
US 10373617B2 · Piernot et al. · 2019 [cited by applicant]
US 10430024B2 · Sachidanandam et al. · 2019 [cited by applicant]
US 10503770B2 · Lee et al. · 2019 [cited by applicant]
US 10573312B1 · Thomson et al. · 2020 [cited by applicant]
US 10579912B2 · Holtmann · 2020 [cited by applicant]
US 10592064B2 · Ames et al. · 2020 [cited by applicant]
US 10594757B1 · Shevchenko et al. · 2020 [cited by applicant]
US 10705794B2 · Gruber et al. · 2020 [cited by applicant]
US 10923101B2 · Guo et al. · 2021 [cited by applicant]
US 10979331B2 · Alsina et al. · 2021 [cited by applicant]
US 11140459B2 · Ingel et al. · 2021 [cited by applicant]
US 11170761B2 · Thomson et al. · 2021 [cited by applicant]
US 11258841B2 · Vincent et al. · 2022 [cited by applicant]
US 11315572B2 · Nishikawa et al. · 2022 [cited by applicant]
US 11403466B2 · Peng et al. · 2022 [cited by applicant]
US 11468892B2 · Lim et al. · 2022 [cited by applicant]
US 11620997B2 · Kurasawa et al. · 2023 [cited by applicant]
US 12105874B2 · Kelly · 2024 [cited by examiner]
US 20040152054A1 · Gleissner et al. · 2004 [cited by applicant]
US 20040218768A1 · Zhurin et al. · 2004 [cited by applicant]
US 20060093998A1 · Vertegaal · 2006 [cited by applicant]
US 20090018828A1 · Nakadai et al. · 2009 [cited by applicant]
US 20100226487A1 · Harder et al. · 2010 [cited by applicant]
US 20110267374A1 · Sakata · 2011 [cited by examiner]
US 20130268954A1 · Hulten · 2013 [cited by examiner]
US 20140078039A1 · Woods · 2014 [cited by examiner]
US 20150082145A1 · Ames et al. · 2015 [cited by applicant]
US 20150106386A1 · Lee · 2015 [cited by examiner]
US 20150213784A1 · Jafarzadeh et al. · 2015 [cited by applicant]
US 20150215532A1 · Jafarzadeh et al. · 2015 [cited by applicant]
US 20160162020A1 · Lehman · 2016 [cited by examiner]
US 20160170710A1 · Kim · 2016 [cited by examiner]
US 20160378747A1 · Orr et al. · 2016 [cited by applicant]
US 20170212583A1 · Krasadakis · 2017 [cited by applicant]
US 20180040046A1 · Gotoh et al. · 2018 [cited by applicant]
US 20180133900A1 · Breazeal et al. · 2018 [cited by applicant]
US 20180227630A1 · Schmidt · 2018 [cited by examiner]
US 20180276706A1 · Hoffman · 2018 [cited by examiner]
US 20190199993A1 · Babu J D · 2019 [cited by examiner]
US 20200128177A1 · Gotou · 2020 [cited by examiner]
US 20200285314A1 · Cieplinski · 2020 [cited by examiner]
US 20200333875A1 · Bansal et al. · 2020 [cited by applicant]
US 20200349966A1 · Konzelmann · 2020 [cited by examiner]
US 20210104242A1 · Hashimoto et al. · 2021 [cited by applicant]
US 20210166686A1 · Mok et al. · 2021 [cited by applicant]
US 20210312938A1 · Yun et al. · 2021 [cited by applicant]
US 20220021914A1 · Lintz et al. · 2022 [cited by applicant]
US 20220093093A1 · Krishnan et al. · 2022 [cited by applicant]
US 20220093094A1 · Krishnan et al. · 2022 [cited by applicant]
US 20220093101A1 · Krishnan et al. · 2022 [cited by applicant]
US 20230036042A1 · Niioka · 2023 [cited by examiner]
US 20230094522A1 · Stauber · 2023 [cited by examiner]
JP 2011227236A · 2011 [cited by applicant]
WO 2019077012A1 · 2019 [cited by applicant]
WO 2019138651A1 · 2019 [cited by applicant]
WO 2021188266A1 · 2021 [cited by applicant]
WO 2023049418A2 · 2023 [cited by applicant]
Oleg Akhtiamov, et al. 2017. Speech and Text Analysis for Multimodal Addressee Detection in Human-Human-Computer Interaction. In Interspeech 2017, 5 pages, https://www.isca-speech.org/archive_v0/Interspeech_2017/pdfs/05… [cited by applicant]
Oleg Akhtiamov, et al. 2017. Are you Addressing Me: Multimodal Addressee Detection in Human-Human-Computer Conversations. In Speech and Computer, SPECOM 2017, Lecture Notes in Computer Science, vol. 10458, pp. 152-161, … [cited by applicant]
Francois Chollet. 2017. Xception: Deep Learning with Depthwise Separable Convolutions. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1800-1807, Retrieved from https://arxiv.org/abs/1610.… [cited by applicant]
Gourav Datta, et al. 2022. ASD-transformer: Efficient active speaker detection using self and multimodal transformers. In ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
Jiankang Deng, et al. 2020. RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5202-5211, Retrieved from https://openacc… [cited by applicant]
Mihail Eric, et al. 2020. MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 4… [cited by applicant]
Lifeng Fan, et al. 2019. Understanding Human Gaze Communication by Spatio-Temporal Graph Reasoning. In IEEE International Conference on Computer Vision (ICCV), pp. 5724-5733, Retrieved from https://arxiv.org/abs/1909.02… [cited by applicant]
Vineet Garg, et al. 2022. Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models. In Interspeech 2022, pp. 1258-1262, https://www.isca-speech.org/archive/pdfs/interspeech_2022/gar… [cited by applicant]
Kellen Gillespie, et al. 2020. Improving Device Directedness Classification Of Utterances With Semantic Lexical Features. In IEEE ICASSP 2020 Virtual Conference May 2020, pp. 7859-7863, Retrieved from https://arxiv.org/… [cited by applicant]
Kristen Grauman, et al. 2022. Ego4D: Around the World in 3,000 Hours of Egocentric Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 18995-19012, Retrieved fro… [cited by applicant]
Jianzhu Guo, et al. 2020. Towards Fast, Accurate and Stable 3D Dense Face Alignment. In Computer Vision—ECCV 2020. ECCV 2020. Lecture Notes in Computer Science, vol. 12364, pp. 152-168 pages, Retrieved from https://www.… [cited by applicant]
Che-Wei Huang, et al. 2019. A Study for Improving Device-Directed Speech Detection toward Frictionless Human-Machine Interaction. In Interspeech 2019, pp. 3342-3346, Retrieved from https://www.isca-speech.org/archive/pd… [cited by applicant]
Kazunori Komatani, et al. 2021. Multimodal Dialogue Data Collection and Analysis of Annotation Disagreement. In Increasing Naturalness and Flexibility in Spoken Dialogue Interaction, Lecture Notes in Electrical Engineer… [cited by applicant]
Alina Kuznetsova, et al. 2020. The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection as scale. In International Journal of Computer Vision, vol. 128, pp. 1956-1981… [cited by applicant]
Thao Le Minh, et al. 2018. Deep Learning Based Multi-modal Addressee Recognition in Visual Scenes with Utterances. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI-18… [cited by applicant]
Sri Harish Mallidi, et al. 2018. Device-directed Utterance Detection. In Interspeech 2018, pp. 1225-1228, Retrieved from https://arxiv.org/abs/1808.02504. [cited by applicant]
Joseph Roth, et al. 2020. Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection. In ICASSP 2020- 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4492-4496, … [cited by applicant]
Mark Sandler, et al. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4510-4520. Retrieved from https://arxiv.org/abs/1801.… [cited by applicant]
Kalin Stefanov, et al. 2016. A Multi-party Multi-modal Dataset for Focus of Visual Attention in Human-human and Human-robot Interaction. In Proceedings of the Tenth International Conference on Language Resources and Eva… [cited by applicant]
Timothy Tsai, et al. 2015. Multimodal addressee detection in multiparty dialogue systems. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2314-2318, Retrieved from https://… [cited by applicant]
Xiangyu Zhu, et al. 2016. Face Alignment Across Large Poses: A 3D Solution. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 146-155, Retrieved from https://arxiv.org/abs/1511.07212v1. [cited by applicant]
International Preliminary Report on Patentability mailed on Mar. 30, 2023 for International Patent Application No. PCT/US2021/050645. [cited by applicant]
International Search Report and Written Opinion mailed on Mar. 9, 2022 for International Patent Application No. PCT/US2021/050645. [cited by applicant]
Invitation to Pay Additional Fees and, Where Applicable, Protest Fee mailed on Dec. 10, 2021 for International Patent Application No. PCT/US2021/050645. [cited by applicant]
U.S. Final Office Action mailed on Feb. 16, 2023 for U.S. Appl. No. 17/112,520. [cited by applicant]
U.S. Non-Final Office Action mailed on Mar. 30, 2023 for U.S. Appl. No. 17/112,512. [cited by applicant]
U.S. Non-Final Office Action mailed on Nov. 3, 2022 for U.S. Appl. No. 17/112,520. [cited by applicant]
U.S. Non-Final Office Action mailed on Oct. 26, 2022 for U.S. Appl. No. 17/112,227. [cited by applicant]
International Search Report and Written Opinion mailed Sep. 12, 2024 for International Patent Application No. PCT/US2024/033204, filed Jun. 10, 2024. [cited by applicant]