IP Library Granted Patent US 12,314,468
Granted Patent B2
US 12,314,468 · App. 18/545,172 · Granted May 27, 2025

Eye tracking in near-eye displays

Inventors: Eric J. Seibel (Seattle, WA); Steven L. Brunton (Seattle, WA); Chen Gong (Seattle, WA); Brian T. Schowengerdt (Plantation, FL)
Assignee: Magic Leap, Inc.
G06F3/013G02B27/0172G02B27/0179G06N3/04G06N3/08G06T7/248G06T7/277G06T7/33G06T7/74G02B2027/0138G02B2027/014G02B2027/0187G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,468
App. No.
18/545,172
Granted
May 27, 2025
Kind
B2
Abstract

Techniques for tracking eye movement in an augmented reality system identify a plurality of base images of an object or a portion thereof. A search image may be generated based at least in part upon at least some of the plurality of base images. A deep learning result may be generated at least by performing a deep learning process on a base image using a neural network in a deep learning mode. A captured image may be localized at least by performing an image registration process on the captured image and the search image using a Kalman filter model and the deep learning result.

Claims (60)

1. A method, comprising:

determining a search image that is constructed based at least in part upon a plurality of template image frames;

capturing, using a spatial computing headset, a plurality of captured image frames of an object or a portion thereof;

performing an image registration process that registers two or more captured image frames of the plurality of captured image frames in the search image using a deep network;

tracking, using the spatial computing headset, the movement of the object based at least in part upon respective results of registering the plurality of captured image frames in the search image; and

training the deep network using at least some captured image frames of the plurality of captured image frames, training the deep network comprising

identifying a captured image frame of the at least some captured image frames captured by the spatial computing headset,

extracting, using the deep network, a capture image frame feature from a first region in the captured image frame and a search image feature from a second region in the search image,

converting the captured image frame feature into a plurality of features that comprises a pair of a first feature and a second feature for the deep network,

providing the first feature converted from the captured image frame feature to a classification subnetwork of the deep network, and

providing the second feature converted from the captured image frame feature to a regression subnetwork of the deep network.

2. The method of claim 1 , wherein tracking the movement of the object is performed by the spatial computing headset, without capturing glint reflected from the object in response to an input light pattern.

3. The method of claim 1 , wherein training the deep network is accomplished without using regularization techniques.

4. The method of claim 1 , training the deep network further comprising:

producing, by the classification subnetwork, a first output data structure at least by convolving, at the classification subnetwork, the first feature converted from the captured image frame feature and search image feature from the search image; and

producing, by the regression subnetwork, a second output data structure at least by convolving, at the regression subnetwork, the second feature converted from the captured image frame feature and search image feature from the search image.

5. The method of claim 3 , training the deep network further comprising:

determining, by the classification subnetwork, whether the first region belongs to a target region or a non-target region based at least in part upon the first output data structure; and

predicting, by the regression subnetwork, a position refinement for the region that is determined to belong to the target region or the non-target region based at least in part upon the second output data structure.

6. The method of claim 1 , wherein the spatial computing headset uses a position of the object or a feature or a portion of the object, instead of geometric attributes of bounding boxes for the object or the feature or the portion of the object in tracking the movement of the object.

7. A system, comprising:

a spatial computing headset comprising:

a processor,

a scanning fiber assembly, and

memory storing thereupon a sequence of instructions, which, when executed by the processor, causes the processor to perform a set of acts, the set of acts comprising:

determining a search image that is constructed based at least in part upon a plurality of template image frames;

capturing, using a spatial computing headset, a plurality of captured image frames of an object or a portion thereof;

performing an image registration process that registers two or more captured image frames of the plurality of captured image frames in the search image using a deep network;

tracking, using the spatial computing headset, the movement of the object based at least in part upon respective results of registering the plurality of captured image frames in the search image; and

training the deep network using at least some captured image frames of the plurality of captured image frames, training the deep network comprising

identifying a captured image frame of the at least some captured image frames captured by the spatial computing headset,

extracting, using the deep network, a capture image frame feature from a first region in the captured image frame and a search image feature from a second region in the search image,

converting the captured image frame feature into a plurality of features that comprises a pair of a first feature and a second feature for the deep network,

providing the first feature converted from the captured image frame feature to a classification subnetwork of the deep network, and

providing the second feature converted from the captured image frame feature to a regression subnetwork of the deep network.

8. The system of claim 7 , wherein tracking the movement of the object is performed by the spatial computing headset, without capturing glint reflected from the object in response to an input light pattern.

9. The system of claim 7 , wherein the set of acts comprises training the deep network, training the deep network comprising:

producing, by the classification subnetwork, a first output data structure at least by convolving, at the classification subnetwork, the first feature converted from the captured image frame feature and search image feature from the search image; and

producing, by the regression subnetwork, a second output data structure at least by convolving, at the regression subnetwork, the second feature converted from the captured image frame feature and search image feature from the search image.

10. The system of claim 9 , wherein the set of acts comprises training the deep network, training the deep network comprising:

determining, by the classification subnetwork, whether the first region belongs to a target region or a non-target region based at least in part upon the first output data structure; and

predicting, by the regression subnetwork, a position refinement for the region that is determined to belong to the target region or the non-target region based at least in part upon the second output data structure.

11. A non-transitory computer-readable medium storing thereupon instructions which, when executed by a microprocessor, causes the microprocessor to perform a set of acts, the set of acts comprising:

determining a search image that is constructed based at least in part upon a plurality of template image frames;

capturing, using a spatial computing headset, a plurality of captured image frames of an object or a portion thereof;

performing an image registration process that registers two or more captured image frames of the plurality of captured image frames in the search image using a deep network;

tracking, using the spatial computing headset, the movement of the object based at least in part upon respective results of registering the plurality of captured image frames in the search image; and

training the deep network using at least some captured image frames of the plurality of captured image frames, training the deep network comprising

identifying a captured image frame of the at least some captured image frames captured by the spatial computing headset,

extracting, using the deep network, a capture image frame feature from a first region in the captured image frame and a search image feature from a second region in the search image,

converting the captured image frame feature into a plurality of features that comprises a pair of a first feature and a second feature for the deep network,

providing the first feature converted from the captured image frame feature to a classification subnetwork of the deep network, and

providing the second feature converted from the captured image frame feature to a regression subnetwork of the deep network,

wherein tracking the movement of the object is performed by the spatial computing headset, without capturing glint reflected from the object in response to an input light pattern.

12. The non-transitory computer-readable medium of claim 11 , wherein the set of acts comprises training the deep network, training the deep network comprising:

producing, by the classification subnetwork, a first output data structure at least by convolving, at the classification subnetwork, the first feature converted from the captured image frame feature and search image feature from the search image; and

producing, by the regression subnetwork, a second output data structure at least by convolving, at the regression subnetwork, the second feature converted from the captured image frame feature and search image feature from the search image.

13. The system of claim 12 , wherein the set of acts comprises training the deep network, training the deep network comprising:

determining, by the classification subnetwork, whether the first region belongs to a target region or a non-target region based at least in part upon the first output data structure; and

predicting, by the regression subnetwork, a position refinement for the region that is determined to belong to the target region or the non-target region based at least in part upon the second output data structure.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: SEIBEL, ERIC J.; BRUNTON, STEVEN L.; GONG, CHEN
To: UNIVERSITY OF WASHINGTON
Reel/Frame 068106/0312 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: SCHOWENGERDT, BRIAN T.
To: MAGIC LEAP, INC.
Reel/Frame 068107/0007 →
Continuity (3)
Continuation 17345305 · Jun 11, 2021
Provisional Application 63038414 · Jun 12, 2020
Related Publication 20240184359A1 · Jun 6, 2024
References Cited (84)
US 6317103B1 · Furness, III et al. · 2001 [cited by applicant]
US 10567641B1 · Rueckner · 2020 [cited by applicant]
US 10852551B1 · Sharma · 2020 [cited by examiner]
US 10852817B1 · Ouderkirk · 2020 [cited by examiner]
US 20040005083A1 · Fujimura et al. · 2004 [cited by applicant]
US 20060146046A1 · Longhurst et al. · 2006 [cited by applicant]
US 20080166025A1 · Thorne · 2008 [cited by applicant]
US 20130293530A1 · Perez · 2013 [cited by examiner]
US 20160292921A1 · Evans · 2016 [cited by examiner]
US 20180068450A1 · Yuan · 2018 [cited by examiner]
US 20180124300A1 · Brook · 2018 [cited by examiner]
US 20180150725A1 · Tate · 2018 [cited by examiner]
US 20180157321A1 · Liu · 2018 [cited by examiner]
US 20190004325A1 · Connor · 2019 [cited by examiner]
US 20190012835A1 · Bleyer · 2019 [cited by examiner]
US 20190035363A1 · Schluessler · 2019 [cited by examiner]
US 20190243448A1 · Miller et al. · 2019 [cited by applicant]
US 20190251333A1 · Wang · 2019 [cited by examiner]
US 20190266418A1 · Xu · 2019 [cited by examiner]
US 20190273910A1 · Malaika · 2019 [cited by examiner]
US 20190295282A1 · Smolyanskiy · 2019 [cited by examiner]
US 20190302883A1 · Greer · 2019 [cited by examiner]
US 20190303759A1 · Farabet · 2019 [cited by examiner]
US 20190369403A1 · Leister · 2019 [cited by examiner]
US 20190379827A1 · Berkovich · 2019 [cited by examiner]
US 20200104457A1 · Chuang · 2020 [cited by examiner]
US 20200186764A1 · Wozniak · 2020 [cited by examiner]
US 20200193976A1 · Cartwright · 2020 [cited by examiner]
US 20200202628A1 · Jones · 2020 [cited by examiner]
US 20200250461A1 · Yang · 2020 [cited by examiner]
US 20200311945A1 · Lim · 2020 [cited by examiner]
US 20200409457A1 · Terrano · 2020 [cited by examiner]
US 20210097943A1 · Wyatt · 2021 [cited by examiner]
US 20210264674A1 · Shahrokni · 2021 [cited by examiner]
JP 2012213513 · 2012 [cited by applicant]
JP 2019517071 · 2019 [cited by applicant]
WO WO2018175625 · 2018 [cited by applicant]
Shuai Wang, “Manufacture Assembly Fault Detection Method based on Deep Learning and Mixed Reality,” Aug. 26, 2019, Proceeding of the IEEE International Conference on Information and Automation ,Wuyi Mountain, China, Aug… [cited by examiner]
Ismoilov Nusrat et al. ,“A Comparison of Regularization Techniques in Deep Neural Networks,” Nov. 18, 2018, Symmetry 2018,10,648, pp. 1-10. [cited by examiner]
Anuradha Kar,“A Review and Analysis of Eye-Gaze Estimation Systems, Algorithms and Performance Evaluation Methods in Consumer Platforms,” Sep. 6, 2017, IEEE Access vol. 5, 2017, Digital Object Identifier 10.1109/ACCESS.… [cited by examiner]
Joseph Lemley,“Efficient CNN Implementation for Eye-Gaze Estimation on Low-Power/Low-Quality Consumer Imaging Systems,” Jun. 28, 2018,arXiv:1806.10890v1 [cs.CV] Jun. 28, 2018, pp. 1-5. [cited by examiner]
Bin Li,“Etracker: A Mobile Gaze-Tracking System with Near-Eye Display Based on a Combined Gaze-Tracking Algorithm,” May 19, 2018, Sensors 2018, 18, 1626; doi: 10.3390/s18051626,pp. 1-14. [cited by examiner]
Jannick P. Rolland,“High-resolution inset head-mounted display,” Jan. 13, 1998,Applied Optics, vol. 37, No. 19,pp. 4183-4187. [cited by examiner]
Christy K. Sheehy,“High-speed, image-based eye tracking with a scanning laser ophthalmoscope,” Sep. 19, 2012,Biomedical Optics Express, vol. 3—No. 10 , pp. 4-6. [cited by examiner]
Alex Krizhevsky,“Imagenet classification with deep convolutional neural networks,” In Advances in neural information processing systems, May 24, 2017, Communications of the ACM,vol. 60,Issue 6,Jun. 2017,pp. 1-6. [cited by examiner]
Chen Gong, “Retinamatch: Efficient template matching of retina images for teleophthalmology,” Jun. 17, 2019, IEEE Transactions on Medical Imaging, 38(8):1993-2004, 2019, pp. 1993-2002. [cited by examiner]
Hong Hua,“Video-based eyetracking methods and algorithms in headmounted displays,” Jun. 8, 2006, Optics Express, 14(10): 2006,pp. 4333-4347. [cited by examiner]
Eric Seibel,“A full-color scanning fiber endoscope,” Feb. 15, 2006, In Optical fibers and sensors for medical diagnostics and treatment applications VI, vol. 6083, pp. 608303-1-608303-4. [cited by examiner]
D.W.F. van Krevelen,“A Survey of Augmented Reality Technologies, Applications and Limitations,” Jan. 26, 2010,The International Journal of Virtual Reality, 2010, 9(2),pp. 1-10. [cited by examiner]
English Translation of Foreign OA for JP Patent Appln. No. 2022-575741 dated Jul. 30, 2024. [cited by applicant]
Foreign Response to EESR for EP Patent Appln. No. 21821754.5 dated May 28, 2024. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/345,305 dated Aug. 10, 2023. [cited by applicant]
Amendment Response to NFOA for U.S. Appl. No. 17/345,305 dated Sep. 6, 2023. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/345,305 dated Sep. 20, 2023. [cited by applicant]
Anuradha Kar,“A Review and Analysis of Eye-Gaze Estimation Systems, Algorithms and Performance Evaluation Methods in Consumer Platforms,” Sep. 6, 2017, IEEE Access vol. 5, 2017, Digital Object Identifier 0.1109/ACCESS.2… [cited by applicant]
Alex Krizhevsky, “Imagenet classification with deep convolutional neural networks,” In Advances in neural information processing systems, OS/24/2017, Communications of the ACM, vol. 60,Issue 6,Jun. 2017,pp. 1-6. [cited by applicant]
Bajura, M. et al. Merging virtual objects with the real world: Seeing ultrasound imagery within the patient. ACM SIGGRAPH Computer Graphics, 26(2):203-210, 1992. [cited by applicant]
Bulling, A. et al. Wearable eog goggles: Seamless sensing and contextawareness in everyday environments. Journal of Ambient Intelligence and Smart Environments, 1(2):157-171, 2009. [cited by applicant]
Caudell, T.P. et al. Augmented reality: An application of heads-up display technology to manual manufacturing processes. In Hawaii International Conference on System Sciences, pp. 659-669, 1992. [cited by applicant]
Gong, C. et al. Real-time Retinal Localization for Eye-tracking in Head-mounted Displays. Fourth Workshop on Computer Vision for AR/VR, Jun. 15, 2020. [cited by applicant]
Gong, C. et al. Retinamatch: Efficient template matching of retina images for teleophthalmology. IEEE Transactions on Medical Imaging, 38(8):1993-2004, 2019. [cited by applicant]
Hua, H. et al. Video-based eyetracking methods and algorithms in head-mounted displays. Optics Express, 14(10):4328-4350, 2006. [cited by applicant]
Krizhevsky, A. et al. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097-1105, 2012. [cited by applicant]
Li, B. et al. High performance visual tracking with siamese region proposal network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8971-8980, 2018. [cited by applicant]
Rolland, J.P. et al. A survey of tracking technologies for virtual environments. In Fundamentals of wearable computers and augmented reality, pp. 83-128. CRC Press, 2001. [cited by applicant]
Rolland, J.P. et al. High-resolution inset head-mounted display. Applied optics, 37(19):4183-4193, 1998. [cited by applicant]
Seibel, E.J. et al. A full-color scanning fiber endoscope. In Optical fibers and sensors for medical diagnostics and treatment applications VI, vol. 6083, p. 608303. International Society for Optics and Photonics, 2006. [cited by applicant]
Sheehy, C.K. et al. High-speed, image-based eye tracking with a scanning laser ophthalmoscope. Biomedical optics express, 3(10):2611-2622, 2012. [cited by applicant]
STARE Project. Structured analysis of the retina. https://cecas.clemson.edu/˜ahoover/stare/. [cited by applicant]
Truong, P. et al. GLAMpoints: Greedily learned accurate match points. In Proceedings of the IEEE International Conference on Computer Vision, pp. 10731-10740, 2019. [cited by applicant]
Welch, G. et al. An introduction to the kalman filter. 1995. [cited by applicant]
Whitmire, E. et al. Eyecontact: scleral coil eye tracking for virtual reality. In Proceedings of the 2016 ACM International Symposium on Wearable Computers, pp. 184-191, 2016. [cited by applicant]
Widefield imaging module. https://business-lounge.heidelbergengineering.com/us/en/products/spectralis/widefield-imaging-module/. [cited by applicant]
PCT International Search Report and Written Opinion for International Appln. No. PCT/US21/37072, Applicant Magic Leap, Inc., dated Oct. 1, 2021. [cited by applicant]
Foreign OA for JP Patent Appln. No. 2022-575741 dated Jul. 30, 2024. [cited by applicant]
Pan, Z., et al., “Human Eye Tracking Based on CNN and Kalman Filtering,” Transactions on Edutainment XV, dated Apr. 27, 2019. [cited by applicant]
B. Li, J. Yan, W. Wu, Z. Zhu and X. Hu, “High Performance Visual Tracking with Siamese Region Proposal Network,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 89… [cited by applicant]
Foreign OA for KR Patent Appln. No. 10-2023-7001121 dated Oct. 18, 2024. [cited by applicant]
Foreign Response for JP Patent Appln. No. 2022-575741 dated Oct. 29, 2024. [cited by applicant]
Foreign Response for KR Patent Appln. No. 10-2023-7001121 dated Dec. 13, 2024. [cited by applicant]
Foreign NOA for JP Patent Appln. No. 2022-575741 dated Nov. 8, 2024. [cited by applicant]
Foreign NOA for IL Patent Appln. No. 298986 dated Dec. 24, 2024. [cited by applicant]
Foreign Exam Report for EP Patent Appln. No. 21821754.5 dated Feb. 24, 2025. [cited by applicant]
Foreign FOA for Korean Appln. No. 10-2023-7001121 dated Apr. 3, 2025. [cited by applicant]