IP Library › Granted Patent US 12,541,940
Granted Patent B2
US 12,541,940 · App. 17/952,810 · Granted Feb 3, 2026

Visual attention tracking using gaze and visual content analysis

Inventors: Yuhu Chang (Shanghai, CN); Yingying Zhao (Shanghai, CN); Qin Lv (Boulder, CO); Robert P. Dick (Chelsea, MI); Li Shang (Shanghai, CN)
Assignee: The Regents of the University of Michigan
G06V10/25G02B27/0172G06F3/013G06V10/44G06V10/806G06V10/82G06V40/19G02B2027/0178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,940
App. No.
17/952,810
Filed
Sep 26, 2022
Granted
Feb 3, 2026
Kind
B2
Art Unit
2677
USPC
382/103
Abstract

A method for detecting content of interest to a user includes obtaining a first data stream indicative of eye movement and/or gaze direction of the user as the user is viewing a scene in a field of view of the user, obtaining a second data stream indicative of visual content in the field of view of the user, determining, based on the first data stream and the second data stream, that content of interest to the user is present in the scene in the field of view of the user, and, in response to determining that content of interest to the user is present in the scene in the field of view of the user, triggering, with the processor, an operation to be performed with respect to the scene in the field of view of the user.

Claims (61)

1 . A method for detecting content of interest to a user, the method comprising:

obtaining, by a processor, a first data stream indicative of one or both of i) eye movement and ii) gaze direction of the user as the user is viewing a scene in a field of view of the user;

obtaining, by the processor, a second data stream indicative of visual content in the field of view of the user;

determining, by the processor, based on the first data stream and the second data stream, i) whether the user is paying attention to the scene in the field of view of the user, including generating a decision indicating whether or not the user is paying attention to the scene in the field of view of the user and ii) a region to which the user is paying attention within the scene, wherein determining whether the user is paying attention to the scene includes:

extracting a set of gaze features from the first data stream,

concurrently with extracting the set of gaze features from the first data stream, extracting a set of scene features from the second data stream,

fusing the set of gaze features with the set of scene features to generate a fused set of features, and

detecting, based on the fused set of features, that the user is paying attention to the scene in the field of view of the user; and

in response to determining that the user is paying attention to the scene, triggering, with the processor, an operation to be performed with respect to the region to which the user is paying attention within the scene.

2 . The method of claim 1 , further comprising:

analyzing, by the processor, the first data stream prior to obtaining the second data stream;

based on analyzing the first data stream, detecting, by the processor, that visual attention of the user is focused on the region within the scene; and

in response to detecting that visual attention of the user is focused on the region within the scene, triggering capture of the second data stream to capture visual content in the field of view of the user.

3 . The method of claim 2 , wherein analyzing the first data stream includes detecting a change in eye movement of the user as the user is viewing the scene, wherein the change in eye movement indicates that the visual attention of the user is focused on the region within the scene.

4 . The method of claim 3 , wherein detecting the change in the eye movement of the user as the user is viewing the scene comprises detecting saccade to smooth pursuit transitions in eye gaze of the user.

5 . The method of claim 1 , wherein determining whether the user is paying attention to the scene in the field of view of the user includes performing a unified analysis of the first data stream and the second data stream by a temporal visual analysis network.

6 . The method of claim 1 , further comprising:

extracting, by the processor, from the first data stream, a likelihood of historical eye movement types, and

detecting a saccade-smooth pursuit transition in eye gaze of the user based at least in part on the likelihood of eye movement types.

7 . The method of claim 6 , further comprising:

extracting, by the processor, historical gaze positions from the set of gaze features, and

detecting the saccade-smooth pursuit transition in the eye gaze of the user further based on determining that a majority of the historical gaze positions fall within a region in the scene.

8 . The method of claim 6 , further comprising, in response to detecting the saccade-smooth pursuit transition, triggering, with the processor, capture of the second data stream to capture visual content in the field of view of the user.

9 . The method of claim 8 , wherein detecting the saccade-smooth pursuit transition in the eye gaze of the user comprising detecting the saccade-smooth pursuit transition for a pre-determined period of time prior to triggering capture of the second data stream to capture visual content in the field of view of the user.

10 . The method of claim 1 , wherein determining whether the user is paying attention to content in the scene includes determining that visual attention of the user is directed to an object in the scene.

11 . The method of claim 1 , wherein triggering the operation to be performed with respect to the region within the scene that has captured attention the user comprises triggering recording of a video snippet capturing the scene in the field of view of the user.

12 . The method of claim 11 , wherein triggering recording of the video snippet capturing the scene in the field of view of the user comprises triggering the recording to be performed for a predetermined duration of time.

13 . A method for tracking visual attention of a user, the method comprising:

obtaining, by a processor, a first data stream indicative of one or both of i) eye movement and ii) gaze direction of the user as the user is viewing a scene in a field of view of the user;

obtaining, by the processor, a second data stream indicative of visual content in the field of view of the user;

using a unified neural network to detect, by the processor, based on the first data stream and the second data stream, i) whether the user is paying attention to the scene in the field of view of the user, including generating a decision indicating whether or not the user is paying attention to the scene and ii) a region to which the user is paying attention within the scene, wherein using the unified neural network to detect whether the user is paying attention to the scene includes:

extracting a set of gaze features from the first data stream,

concurrently with extracting the set of gaze features from the first data stream, extracting a set of scene features from the second data stream,

fusing the set of gaze features with the set of scene features to generate a fused set of features, and

detecting, based on the fused set of features, that the user is paying attention to the scene in the field of view of the user; and

in response to determining that the user is paying attention to the scene, triggering recording of a video snippet capturing the region to which the user is paying attention within the scene.

14 . The method of claim 13 , further comprising:

analyzing, by the processor, the first data stream prior to obtaining the second data stream;

based on analyzing the first data stream, detecting, by the processor, that visual attention of the user is focused on the region within the scene; and

in response to detecting that visual attention of the user is focused on the region within the scene, triggering capture of the second data stream to capture visual content in the field of view of the user.

15 . A system comprising:

a first sensor configured to generate data indicative of one or both of i) eye movement and ii) gaze direction of a user as the user is viewing a scene in a field of view of the user;

a second sensor configured to generate data indicative of visual content in the field of view of the user; and

a processor configured to:

obtain, from the first sensor, a first data stream indicative of the one or both of i) eye movement and ii) gaze direction of the user as the user is viewing the scene,

obtain, from the second sensor, a second data stream indicative of visual content in the field of view of the user,

based on analyzing the first data stream and the second data stream, i) determine whether the user is paying attention to the scene in the field of view of the user, wherein the processor is configured to generate a decision indicating whether or not the user is paying attention to the scene, and ii) determine a region to which the user is paying attention within the scene, the processor being configured to:

extract a set of gaze features from the first data stream,

concurrently with extracting the set of gaze features from the first data stream, extract a set of scene features from the second data stream,

fuse the set of gaze features with the set of scene features to generate a fused set of features, and

detect, based on the fused set of features, that content of interest to the user is present in the scene in the field of view of the user; and

in response to determining that the user is paying attention to the scene, triggering, with the processor, an operation to be performed with respect to the region to which the user is paying attention within the scene.

16 . The system of claim 15 , wherein the first sensor and second sensor are mounted to a frame of smart eyewear to be worn by the user.

17 . The system of claim 16 , wherein:

the first sensor is configured as an inward-facing camera facing eyes of the user when the user wears the smart eyewear, and

the second sensor is configured as a forwarding-facing camera with respect to the field of view of the user when the user wears the smart eyewear.

18 . The system of claim 17 , a first resolution of the inward-facing camera is lower than a second resolution of the forwarding-facing camera.

19 . The system of claim 15 , wherein the processor is further configured to:

prior to obtaining the second data stream, analyze the first data stream, obtained from the first sensor,

based on analyzing the first data stream, detect that visual attention of the user is focused on the region within the scene, and

in response to detecting that visual attention of the user is focused on the region within the scene, trigger capture of the second data stream to capture visual content in the field of view of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2022
From: DICK, ROBERT P.
To: THE REGENTS OF THE UNIVERSITY OF MICHIGAN
Reel/Frame 061999/0878 →
Continuity (2)
Provisional Application 63247893 · Sep 24, 2021
Related Publication 20230133579A1 · May 4, 2023
References Cited (77)
US 10685488B1 · Kumar · 2020 [cited by examiner]
US 20150070386A1 · Ferens · 2015 [cited by examiner]
US 20160217327A1 · Osterhout · 2016 [cited by examiner]
US 20170293356A1 · Khaderi · 2017 [cited by examiner]
US 20170347039A1 · Baumert · 2017 [cited by examiner]
US 20180224658A1 · Teller · 2018 [cited by examiner]
US 20180232575A1 · Strombom · 2018 [cited by examiner]
US 20180261253A1 · Gilson · 2018 [cited by examiner]
US 20190179423A1 · Rose · 2019 [cited by examiner]
US 20190286227A1 · Samadani · 2019 [cited by examiner]
US 20190371075A1 · Stafford · 2019 [cited by examiner]
US 20190391638A1 · Khaderi · 2019 [cited by examiner]
US 20200034617A1 · Croxford · 2020 [cited by examiner]
US 20200081524A1 · Schmidt · 2020 [cited by examiner]
US 20200103967A1 · Bar-Zeev · 2020 [cited by examiner]
US 20200128232A1 · Hwang · 2020 [cited by examiner]
US 20200348515A1 · Peuhkurinen · 2020 [cited by examiner]
US 20200380767A1 · Xu · 2020 [cited by examiner]
US 20210258554A1 · Bruls · 2021 [cited by examiner]
US 20210373656A1 · Watola · 2021 [cited by examiner]
CN 112507799A · 2021 [cited by examiner]
WO WO2017209978A1 · 2017 [cited by examiner]
Krafka et al., Eye Tracking for Everyone, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1-9, AAPA furnished via IDS. [cited by examiner]
Silva et al., Leveraging Eye-gaze and Time-series Features to Predict User Interests and Build a Recommendation Model for Visual Analysis, ETRA '18: Proceedings of the 2018 ACM Symposium on Eye Tracking Research & Appli… [cited by examiner]
Alam et al., Analyzing Eye-Tracking Information in Visualization and Data Space: From Where on the Screen to What on the Screen , in IEEE Transactions on Visualization and Computer Graphics, vol. 23, No. 5, pp. 1492-150… [cited by examiner]
Feng et al., Low-Cost Eye Gaze Prediction System for Interactive Networked Video Streaming, in IEEE Transactions on Multimedia, vol. 15, No. 8, pp. 1865-1879, Dec. 2013, doi: 10.1109/TMM.2013.2272918. [cited by examiner]
Katti et al., Online Estimation of Evolving Human Visual Interest, ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 11, Issue 1, Article No. 8, pp. 1-21, Sep. 2014, doi.org/10.1145… [cited by examiner]
E. S. Lubana and R. P. Dick; Digital Foveation: an energy-aware machine vision framework; IEEE Trans. on Computer-Aided Design, pp. 2371-2380, Nov. 2018. [cited by applicant]
A. Torralba and P. Sinha; Statistical context priming for object detection; Proceedings Eighth IEEE International Conference on Computer Vision. ICCV 2001, pp. 763-770. [cited by applicant]
Ambarella introduces CV28M SoC with CVflow to enable new categories of intelligent sensing devices; Ambarella; 2020; https://www.ambarella.com; 5 pp. [cited by applicant]
Andrew Howard et al.; Searching for MobileNetV3; Proceedings of the IEEE International Conference on Computer Vision; 2019; pp. 1314-1324. [cited by applicant]
Benjamin Tag et al., Facial Thermography for Attention Tracking on Smart Eyewear: An Initial Study. Proceedings of the 2017 CHI Conference Extended Abstracts on Human Factors in Computing Systems; May 2017; pp. 2959-296… [cited by applicant]
Broadbent, D.E.; A mechanical model for human attention and immediate memory; Psychological Review vol. 64, 3; 1957; pp. 205-215. [cited by applicant]
Cherlynn Low; Google Clips review: A smart, but unpredictable camera; https://www.engadget.com/2018-02-27-google-clips-ai-camera-review.html; 2018; 13 pp. [cited by applicant]
Dan Witzner Hansen and Qiang Ji; In the eye of the beholder: A survey of models for eyes and gaze; IEEE transactions on pattern analysis and machine intelligence, vol. 32, 3; 2009, pp. 478-500. [cited by applicant]
Dario Cazzato, Marco Leo, Cosimo Distante, and Holger Voos; When I Look into Your Eyes: A Survey on Computer Vision Contributions for Human Gaze Estimation and Tracking; Sensors vol. 20, 13; 2020, pp. 1-42. [cited by applicant]
Dario D Salvucci and Joseph H Goldberg; Identifying fixations and saccades in eye-tracking protocols; Proceedings of the 2000 symposium on Eye tracking research & applications; ACM, 2020; pp. 71-78. [cited by applicant]
E. Chong, Y. Wang, N. Ruiz, and J. M. Rehg; Detecting Attended Visual Targets in Video; In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020; pp. 5395-5405. [cited by applicant]
Fakta Jain, Yaser Sheikh, Ariel Shamir, and Jessica Hodgins; Gaze-driven video re-editing; ACM Transactions on Graphics vol. 34, 2; 2015, pp. 1-12. [cited by applicant]
En Teng Wong et al.; Gaze Estimation Using Residual Neural Network; 2019 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops); IEEE, 2019; pp. 411-414. [cited by applicant]
Enkelejda Tafaj, Gjergji Kasneci, Wolfgang Rosenstiel, and Martin Bogdan; Bayesian Online Clustering of Eye Movement Data; Proceedings of the Symposium on Eye Tracking Research and Applications (ETRA '12); 2012, pp. 285… [cited by applicant]
Enkelejda Tafaj, Thomas C. Kübler, Gjergji Kasneci, Wolfgang Rosenstiel, and Martin Bogdan; Online Classification of Eye Tracking Data for Automated Analysis of Traffic Hazard Perception; Proceedings of the 23rd Interna… [cited by applicant]
Erroll Wood and Andreas Bulling; Eyetab: Model-based gaze estimation on unmodified tablet computers; Proceedings of the Symposium on Eye Tracking Research and Applications; 2014; pp. 207-210. [cited by applicant]
G. Huang, Z. Liu, K. Q. Weinberger, and L. van derMaaten; Densely connected convolutional network; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 4700-4708. [cited by applicant]
Grace W Lindsay; Attention in Psychology, Neuroscience, and Machine Learning; Frontiers in Computational Neuroscience vol. 14, 29; 2020; pp. 1-21. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, The IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Jeremiah W Johnson; Adapting mask-rcnn for automatic nucleus segmentation; 2018; arXiv preprint arXiv:1805.00500; pp 1-7. [cited by applicant]
Kang Wang, Hui Su, and Qiang Ji; Neuro-inspired Eye Tracking with Eye Movement Dynamics; Proceedings of the IEEE conference on computer vision and pattern recognition; 2019; pp. 9831-9840. [cited by applicant]
Kang Wang, Rui Zhao, Hui Su, and Qiang Ji; Generalizing eye tracking with bayesian adversarial learning; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2019; pp. 11907-11916. [cited by applicant]
Katharina Anton-Erxleben and Marisa Carrasco; Attentional enhancement of spatial resolution: linking behavioural and heurophysiological evidence; Nature Reviews Neuroscience vol. 14, 3; 2013, pp. 188-200. [cited by applicant]
Kingma, Diederik P., and Jimmy Ba; “Adam: A Method for Stochastic Optimization.” arXiv:1412.6980 [Cs], Dec. 22, 2014. http://arxiv.org/abs/1412.6980; 9 pp. [cited by applicant]
Kyle Krafka et al., Eye tracking for everyone; Proceedings of the IEEE conference on computer vision and pattern recognition; 2016; pp. 2176-2184. [cited by applicant]
Laxmidhar Behera, Indrani Kar, and Avshalom C Elitzur; A recurrent quantum neural network model to describe eye tracking of moving targets; Foundations of Physics Letters, vol. 18, 4; 2005; pp. 357-370. [cited by applicant]
Linjie Yang, Yuchen Fan, and Ning Xu; Video instance segmentation; Proceedings of the IEEE International Conference on Computer Vision; 2019; pp. 5188-5197. [cited by applicant]
Marisa Carrasco; Visual attention: The past 25 years; Vision Research vol. 51, 13; 2011, pp. 1484-1525. [cited by applicant]
Mark Sandler et al., MobileNetV2: Inverted residuals and linear bottlenecks; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018; pp. 4510-4520. [cited by applicant]
Mikhail Startsev, Ioannis Agtzidis, and Michael Dorr; 1D CNN with BLSTM for automated classification of fixations, saccades, and smooth pursuits; Behavior Research Methods vol. 51, 2; 2019, pp. 556-572. [cited by applicant]
Moritz Kassner, William Patera, and Andreas Bulling; Pupil: an open source platform for pervasive eye tracking and mobile gaze-based interaction; Proceedings of the 2014 ACM international joint conference on pervasive a… [cited by applicant]
Nelson Silva et al., Leveraging Eye-Gaze and Time-Series Features to Predict User Interests and Build a Recommendation Model for Visual Analysis; Proceedings of the 2018 ACM Symposium on Eye Tracking Research & Applicat… [cited by applicant]
Oleg V Komogortsev et al., Qualitative and quantitative scoring and evaluation of the eye movement classification algorithms; Proceedings of the 2010 Symposium on Eye-Tracking Research & Applications; ACM, 2010; pp. 65-… [cited by applicant]
Rachavarapu, Kranthi Kumar et al., Watch to Edit: Video Retargeting using Gaze; Computer Graphics Forum vol. 37, 2, 2018; pp. 205-215. [cited by applicant]
Raimondas Zemblys, Diederick C. Niehorster, Kenneth Holmqvist; gazeNet: End-to-end eye-movement event detection with deep neural networks; Behavior Research Methods vol. 51; 2019; pp. 840-864. [cited by applicant]
Shagen Djanian; Eye Movement Classification Using Deep Learning; Master Thesis. Department of Vision, Graphics and Interactive Systems, Aalborg University, 2019; 82 pp. [cited by applicant]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun; Faster r-cnn: Towards real-time object detection with region proposal networks; Advances in Neural Information Processing Systems 28 (NIPS), 2015; pp. 91-99. [cited by applicant]
Shuo Wang et al., Atypical visual saliency in autism spectrum disorder quantified through model-based eye tracking; Neuron vol. 88, 3; 2015; pp. 604-616. [cited by applicant]
Sophie Marat et al., Modelling spatio-temporal saliency to predict gaze direction for short videos; International Journal of Computer Vision vol. 82, 3; 2009, pp. 231-243. [cited by applicant]
Thiago Santini, Wolfgang Fuhl, Thomas Kubler, and Enkelejda Kasneci; Bayesian identification of fixations, saccades, and smooth pursuits; Proceedings of the Ninth Biennial ACM Symposium on Eye Tracking Research & Applic… [cited by applicant]
Trixie A Katz et al., Visual attention on a respiratory function monitor during simulated neonatal resuscitation: an eye-tracking study; Archives of Disease in Childhood-Fetal and Neonatal Edition vol. 104, 3, 2019, pp.… [cited by applicant]
Vannevar Bush; As we may think; The Atlantic Monthly, vol. 176, 1; 1945, pp. 101-108. [cited by applicant]
Wenguan Wang et al., Learning unsupervised video object segmentation through visual attention; Proceedings of the IEEE conference on computer vision and pattern recognition; 2019; pp. 3064-3074. [cited by applicant]
Xucong Zhang, Yusuke Sugano, and Andreas Bulling; Evaluation of Appearance-Based Methods and Implications for Gaze-Based Applications; Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems; ACM, 2… [cited by applicant]
Xucong Zhang, Yusuke Sugano, and Andreas Bulling; Revisiting Data Normalization for Appearance-Based Gaze Estimation; Proceedings of the International Symposium on Eye Tracking Research and Applications (ETRA); 2018; pp… [cited by applicant]
Xucong Zhang, Yusuke Sugano, Mario Fritz, and Andreas Bulling; MPIIGaze: Real-world dataset and deep appearance-based gaze estimation; IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 41, 1; 2019, pp.… [cited by applicant]
Yingying Zhao et al., A Reinforcement-Learning-Based Energy-Efficient Framework for Multi-Task Video Analytics Pipeline; IEEE Transactions on Multimedia vol. 24; 2021; 14 pp; arXiv:2104.04443. [cited by applicant]
Yomna Abdelrahman et al.; Classifying attention types with thermal imaging and eye tracking; Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies; vol. 3, 3; 2019; pp. 1-27. [cited by applicant]
Yoon Min Hwang and Kun Chang Lee; Using an eye-tracking approach to explore gender differences in visual attention and shopping attitudes in an online shopping environment; International Journal of Human-Computer Intera… [cited by applicant]
Yujiang Wang et al., Dynamic Face Video Segmentation via Reinforcement Learning; Proceedings of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition; 2020; pp. 6959-6969. [cited by applicant]