IP Library Granted Patent US 12,487,664
Granted Patent B2
US 12,487,664 · App. 17/674,724 · Granted Dec 2, 2025

Eye tracking and gaze estimation using off-axis camera

Inventors: Zhengyang Wu (Bellewue, WA); Srivignesh Rajendran (San Francisco, CA); Tarrence van As (New York, NY); Joelle Zimmermann (Los Angeles, CA); Vijay Badrinarayanan (Mountain View, CA); Andrew Rabinovich (San Francisco, CA)
Assignee: Magic Leap, Inc.
G06F3/013G02B27/0093G06V10/267G06V10/774G06V10/82G06V40/193G06V40/197G02B2027/0138G02B27/0172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,487,664
App. No.
17/674,724
Granted
Dec 2, 2025
Kind
B2
Abstract

Techniques related to the computation of gaze vectors of users of wearable devices are disclosed. A neural network may be trained through first and second training steps. The neural network may include a set of feature encoding layers and a plurality of sets of task-specific layers that each operate on an output of the set of feature encoding layers. During the first training step, a first image of a first eye may be provided to the neural network, eye segmentation data may be generated using the neural network, and the set of feature encoding layers may be trained. During the second training step, a second image of a second eye may be provided to the neural network, network output data may be generated using the neural network, and the plurality of sets of task-specific layers may be trained.

Claims (55)

1 . A method of training a neural network, the method comprising:

performing a first training step including:

providing a first image of a first eye to the neural network as input, the neural network having a set of feature encoding layers connected to a plurality of sets of task-specific layers, the plurality of sets of task-specific layers including at least three sets of task-specific layers that operate on an output generated by the set of feature encoding layers, the plurality of sets of task-specific layers including:

a first set of task-specific layers that output two-dimensional (2D) pupil data,

a second set of task-specific layers that output eye segmentation data that includes a segmentation of an eye into a plurality of regions including one or more of a background region, a sclera region, a pupil region, or an iris region, and

a third of task-specific layers that output cornea center data;

generating, using the set of feature encoding layers and the second set of task-specific layers of the neural network and based on the first image of the first eye as input, eye segmentation data for the first eye that includes a segmentation of the first eye into the plurality of regions; and

training the set of feature encoding layers using the eye segmentation data for the first eye by modifying weights associated with the set of feature encoding layers; and

performing a second training step including:

providing a second image of a second eye to the neural network as input;

generating, using the set of feature encoding layers, the first set of task specific layers, and the third set of task-specific layers of the neural network and based on the second image of the second eye as input, network output data including 2D pupil data corresponding to the second eye and cornea center data corresponding to the second eye; and

training the plurality of sets of task-specific layers using the network output data by modifying weights associated with the plurality of sets of task-specific layers;

wherein the neural network is trained such that the set of feature encoding layers are trained during the first training step but are held fixed during the second training step.

2 . The method of claim 1 , wherein the first training step is performed during a first time duration and the second training step is performed during a second time duration that is after the first time duration.

3 . The method of claim 1 , wherein the plurality of regions includes one or more of a background region, a sclera region, a pupil region, or an iris region.

4 . The method of claim 1 , wherein performing the first training step further includes:

training the second set of task-specific layers using the eye segmentation data for the first eye.

5 . The method of claim 1 , wherein performing the first training step further includes:

receiving eye segmentation ground truth (GT) data; and

comparing the eye segmentation data to the eye segmentation GT data.

6 . The method of claim 1 , wherein the network output data includes glint detection data corresponding to the second eye.

7 . The method of claim 1 , wherein the network output data includes a blink prediction corresponding to the second eye.

8 . The method of claim 1 , wherein the network output data includes an eye expression classification corresponding to the second eye.

9 . The method of claim 1 , wherein the network output data includes eye segmentation data for the second eye that includes a second segmentation of the second eye into the plurality of regions.

10 . A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations for training a neural network, wherein the operations comprise:

performing a first training step including:

providing a first image of a first eye to the neural network as input, the neural network having a set of feature encoding layers connected to a plurality of sets of task-specific layers, the plurality of sets of task-specific layers including at least three sets of task-specific layers that operate on an output generated by the set of feature encoding layers, the plurality of sets of task-specific layers including:

a first set of task-specific layers that output two-dimensional (2D) pupil data,

a second set of task-specific layers that output eye segmentation data that includes a segmentation of an eye into a plurality of regions including one or more of a background region, a sclera region, a pupil region, or an iris region, and

a third of task-specific layers that output cornea center data;

generating, using the set of feature encoding layers and the second set of task-specific layers of the neural network and based on the first image of the first eye as input, eye segmentation data for the first eye that includes a segmentation of the first eye into the plurality of regions; and

training the set of feature encoding layers using the eye segmentation data for the first eye by modifying weights associated with the set of feature encoding layers; and

performing a second training step including:

providing a second image of a second eye to the neural network as input;

generating, using the set of feature encoding layers, the first set of task specific layers, and the third set of task-specific layers of the neural network and based on the second image of the second eye as input, network output data including 2D pupil data corresponding to the second eye and cornea center data corresponding to the second eye; and

training the plurality of sets of task-specific layers using the network output data by modifying weights associated with the plurality of sets of task-specific layers;

wherein the neural network is trained such that the set of feature encoding layers are trained during the first training step but are held fixed during the second training step.

11 . The non-transitory computer-readable medium of claim 10 , wherein the first training step is performed during a first time duration and the second training step is performed during a second time duration that is after the first time duration.

12 . A system comprising:

one or more processors; and

a non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for training a neural network, wherein the operations comprise:

performing a first training step including:

providing a first image of a first eye to the neural network as input, the neural network having a set of feature encoding layers connected to a plurality of sets of task-specific layers, the plurality of sets of task-specific layers including at least three sets of task-specific layers that operate on an output generated by the set of feature encoding layers, the plurality of sets of task-specific layers including:

a first set of task-specific layers that output two-dimensional (2D) pupil data,

a second set of task-specific layers that output eye segmentation data that includes a segmentation of an eye into a plurality of regions including one or more of a background region, a sclera region, a pupil region, or an iris region, and

a third of task-specific layers that output cornea center data;

generating, using the set of feature encoding layers and the second set of task-specific layers of the neural network and based on the first image of the first eye as input, eye segmentation data for the first eye that includes a segmentation of the first eye into the plurality of regions; and

training the set of feature encoding layers using the eye segmentation data for the first eye by modifying weights associated with the set of feature encoding layers; and

performing a second training step including:

providing a second image of a second eye to the neural network as input;

generating, using the set of feature encoding layers, the first set of task specific layers, and the third set of task-specific layers of the neural network and based on the second image of the second eye as input, network output data including 2D pupil data corresponding to the second eye and cornea center data corresponding to the second eye; and

training the plurality of sets of task-specific layers using the network output data by modifying weights associated with the plurality of sets of task-specific layers;

wherein the neural network is trained such that the set of feature encoding layers are trained during the first training step but are held fixed during the second training step.

13 . The system of claim 12 , wherein the first training step is performed during a first time duration and the second training step is performed during a second time duration that is after the first time duration.

14 . The system of claim 12 , wherein the plurality of regions includes one or more of a background region, a sclera region, a pupil region, or an iris region.

Assignments (1)
SECURITY INTEREST Recorded May 24, 2022
From: MOLECULAR IMPRINTS, INC.; MENTOR ACQUISITION ONE, LLC; MAGIC LEAP, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 060338/0665 →
Continuity (5)
Continuation PCTUS2020047046 · Aug 19, 2020
Provisional Application 62935584 · Nov 14, 2019
Provisional Application 62926241 · Oct 25, 2019
Provisional Application 62888953 · Aug 19, 2019
Related Publication 20220244781A1 · Aug 4, 2022
References Cited (73)
US 8878749B1 · Wu et al. · 2014 [cited by applicant]
US 11775058B2 · Badrinarayanan et al. · 2023 [cited by applicant]
US 20130050070A1 · Lewis et al. · 2013 [cited by applicant]
US 20150178939A1 · Bradski et al. · 2015 [cited by applicant]
US 20150346495A1 · Welch et al. · 2015 [cited by applicant]
US 20160026253A1 · Bradski et al. · 2016 [cited by applicant]
US 20160202756A1 · Wu · 2016 [cited by applicant]
US 20170119298A1 · Cheung · 2017 [cited by applicant]
US 20180053056A1 · Rabinovich et al. · 2018 [cited by applicant]
US 20180089834A1 · Spizhevoy · 2018 [cited by examiner]
US 20180137335A1 · Kim · 2018 [cited by examiner]
US 20180157892A1 · Han · 2018 [cited by examiner]
US 20180181809A1 · Ranjan et al. · 2018 [cited by applicant]
US 20180300897A1 · Woods et al. · 2018 [cited by applicant]
US 20190073533A1 · Chen et al. · 2019 [cited by applicant]
US 20190213434A1 · Zamfir · 2019 [cited by examiner]
US 20190244108A1 · Meyerson et al. · 2019 [cited by applicant]
US 20190295287A1 · De Villers-Sidani et al. · 2019 [cited by applicant]
US 20190302883A1 · Greer et al. · 2019 [cited by applicant]
US 20190303722A1 · Linden · 2019 [cited by applicant]
US 20200202561A1 · Liu · 2020 [cited by examiner]
US 20200348755A1 · Gebauer · 2020 [cited by examiner]
US 20210049410A1 · Dierkes et al. · 2021 [cited by applicant]
US 20210182554A1 · Badrinarayanan et al. · 2021 [cited by applicant]
US 20220100268A1 · Weinberg · 2022 [cited by examiner]
US 20220197376A1 · Boyle · 2022 [cited by examiner]
CN 106133648A · 2016 [cited by applicant]
CN 106503654A · 2017 [cited by applicant]
CN 108073889A · 2018 [cited by applicant]
CN 108809522A · 2018 [cited by applicant]
CN 109274812A · 2019 [cited by applicant]
CN 112400148A · 2021 [cited by applicant]
EP 3811182A · 2021 [cited by applicant]
JP 2016106668A · 2016 [cited by applicant]
JP 2017211891A · 2017 [cited by applicant]
WO 2018000020A1 · 2018 [cited by applicant]
WO 2018035531A1 · 2018 [cited by applicant]
WO 2018039269A1 · 2018 [cited by applicant]
WO 2018063451A1 · 2018 [cited by applicant]
WO 2019147677A1 · 2019 [cited by applicant]
WO 2019246613A1 · 2019 [cited by applicant]
“General Theory of Remote Gaze Estimation Using the Pupil Center and Corneal Reflections,” Elias Daniel Guestrin and Moshe Eizenman, IEEE Transactions on Biomedical Engineering vol. 53, No. 6 p. 1124-1133 (Year: 2006). [cited by examiner]
A free geometry model-independent neural eye-gaze tracking system, by Massimo Gneo, Maurizio Schmid, Silvia Conforto, Tommaso D'Alessio, pub Journal of Neuro Engineering and Rehabilitation, 9, 82, Nov. 16, 2012 (Year: 2… [cited by examiner]
CN201980041066.6, “Office Action”, Jan. 21, 2024, 7 pages. [no translation available]. [cited by applicant]
JP2020570526, “Office Action” and English translation, Dec. 28, 2023, 5 pages. [cited by applicant]
Bengio , “Learning Deep Architectures for AI”, Foundations and Trends in Machine Learning, vol. 2, No. 1, 2009, pp. 1-127. [cited by applicant]
Application No. EP20853899.1 , “Extended European Search Report”, Sep. 9, 2022, 12 pages. [cited by applicant]
Playout et al., “A Novel Weakly Supervised Multitask Architecture for Retinal Lesions Segmentation on Fundus Images”, IEEE Transactions on Medical Imaging, vol. 38, No. 10, Oct. 2019, pp. 2434-2444. [cited by applicant]
Playout et al., “A Novel Weakly Supervised Multitask Architecture for Retinal Lesions Segmentation on Fundus Images—Abstract”, IEEE Transactions on Medical Imaging, vol. 38, No. 10 Available Online at : URL:https://pubm… [cited by applicant]
Application No. EP19821828.1, Extended European Search Report, Mailed on Jun. 28, 2021, 7 pages. [cited by applicant]
Application No. PCT/US2019/038693, International Preliminary Report on Patentability, mailed on Dec. 30, 2020, 10 pages. [cited by applicant]
Application No. PCT/US2019/038693 , International Search Report and Written Opinion, mailed on Sep. 20, 2019, 7 pages. [cited by applicant]
Application No. PCT/US2020/047046 , International Preliminary Report on Patentability, mailed on Mar. 3, 2022, 9 pages. [cited by applicant]
Application No. PCT/US2020/047046 , International Search Report and Written Opinion, mailed on Nov. 9, 2020, 10 pages. [cited by applicant]
OpenCV: cv::SimpleBlobDetector Class Reference, 2D Features Framework/Feature Detection and Description, retrieved from internet: https://docs.opencv.org/3.4/20/d7a/classcv_1_1SimpleBlobDeterctor.html, on Aug. 19, 2019,… [cited by applicant]
Kitazumi et al., “Pupil Detection from Visible Light Corneal Image Using Deep Learning and Application to Gaze Estimation”, IPSJ SIG Technical Report, Computer Vision and Image Media (CVIM), Jan. 11, 2018, 8 pages. [Eng… [cited by applicant]
U.S. Appl. No. 17/129,669, “Notice of Allowance”, Jun. 7, 2023, 9 pages. [cited by applicant]
European Patent Application No. 19821828.1, “Office Action”, May 25, 2023, 7 pages. [cited by applicant]
Guestrin et al., “General Theory of Remote Gaze Estimation Using the Pupil Center and Corneal Reflections”, IEEE Transactions on Biomedical Engineering, vol. 53, No. 6, Jun. 2006, pp. 1124-1133. [cited by applicant]
Japanese Patent Application No. 2020-570526, “Office Action” and English translation, Jun. 30, 2023, 10 pages. [cited by applicant]
U.S. Appl. No. 17/129,669 , “Non-Final Office Action”, Feb. 15, 2023, 10 pages. [cited by applicant]
JP 2022-510817, “Office Action”, Jun. 17, 2024, 3 pages. [no translation available]. [cited by applicant]
Matsui et al., “Simultaneous Estimation of Facial Landmark and Attributes with Separation Multi-task Networks”, International Conference on Computer Vision Theory and Applications, Feb. 27, 2019, 8 pages. [cited by applicant]
Matsui et al., “Simultaneous estimation of facial organ points and face attributes by Separation Multi-tasks Networks”, The Institute of Electronics, Information and Communication Engineers, vol. 118, No. 362, Technical… [cited by applicant]
EP20853899.1, “Office Action”, Dec. 4, 2024, 10 pages. [cited by applicant]
CN201980041066.6, “Notice of Decision to Grant”, Oct. 30, 2024, 4 pages. [no translation available]. [cited by applicant]
CN202080059575.4, “Office Action”, Sep. 23, 2024, 10 pages. [no translation available]. [cited by applicant]
CN202080059575.4, “Office Action”, Dec. 12, 2024, 13 pages. [no translation available]. [cited by applicant]
Vandenhende et al., “Branched Multi-Task Networks: Deciding What Layers to Share”, [MTL] Branched Multi-Task Networks, Jun. 11, 2019, pp. 1-7. [cited by applicant]
CN202080059575.4, “Office Action and English translation”, Feb. 18, 2025, 16 pages. [cited by applicant]
CN202080059575.4, “Office Action and English translation”, Apr. 23, 2025, 20 pages. [cited by applicant]
CN202080059575.4, “Notice of Decision to Grant” and English translation, Aug. 27, 2025, 8 pages. [cited by applicant]
EP19821828.1, “Office Action”, Jun. 5, 2025, 5 pages. [cited by applicant]