IP Library › Granted Patent US 12,405,660
Granted Patent B2
US 12,405,660 · App. 16/872,069 · Granted Sep 2, 2025

Gaze estimation using one or more neural networks

Inventors: Michael Stengel (Cupertino, CA); Morgan McGuire (Waterloo, CA); Alexander Majercik (San Francisco, CA); David Luebke (Charlottesville, VA)
Assignee: NVIDIA Corporation
G06F3/013G06N3/08G06T7/246G06T7/70G06V10/82G06V20/20G06V40/19G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,660
App. No.
16/872,069
Granted
Sep 2, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to estimate user gaze. In at least one embodiment, one or more neural networks are used to determine coarse and fine gaze estimates for one or more users.

Claims (38)

1. One or more processors, comprising:

circuitry to use an amount by which a first image has changed from a second image to select whether to use one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils within the first image, wherein one or more second portions of the one or neural networks are used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image when the amount is below a threshold amount.

2. The one or more processors of claim 1 , wherein the circuitry to use the one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils further comprises passing a full eye image from one or more images of one or more eyes of one or more users through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

3. The one or more processors of claim 1 , wherein the circuitry to use the one or more second portions of one or more neural networks used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image further comprises passing smaller sub-regions of a full eye image through the one or more second portions of one or more neural networks, wherein the full eye image is the prior image that has passed through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

4. The one or more processors of claim 1 , wherein the one or more first portions of one or more neural networks and the one or more second portions of one or more neural networks use a same convolutional neural network (CNN)-based architecture but with different network parameters.

5. The one or more processors of claim 1 , wherein the threshold amount comprises an amount of lateral movement of a pupil in an image.

6. The one or more processors of claim 1 , wherein the second image includes one or more separate images for both eyes of one of one or more users, and wherein the circuitry is further to concatenate those separate images into a concatenated image with padding between representations of the eyes before providing the concatenated image as input to the one or more first portions of the one or more neural networks.

7. A system comprising:

one or more processors to use an amount by which a first image has changed from a second image to select whether to use one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils within the first image, wherein one or more second portions of the one or neural networks are used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image when the amount is below a threshold amount.

8. The system of claim 7 , wherein the one or more processors to use one or more first portions of the one or more neural networks to generate a coarse estimate of a location of one or more pupils further comprises passing a full eye image from one or more images of one or more eyes of one or more users through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

9. The system of claim 7 , wherein the one or more processors to use one or more second portions of the one or more neural networks used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image further comprises passing smaller sub-regions of a full eye image through the one or more second portions of one or more neural networks, wherein the full eye image is the prior image that has passed through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

10. The system of claim 7 , wherein the one or more first portions of one or more neural networks and the one or more second portions of one or more neural networks use a same convolutional neural network (CNN)-based architecture but with different network parameters.

11. The system of claim 7 , wherein the threshold amount comprises an amount of lateral movement of a pupil in an image.

12. The system of claim 7 , wherein the second image includes one or more separate images for both eyes of one of one or more users, and wherein the one or more processors are further to concatenate those separate images into a concatenated image with padding between representations of the eyes before providing the concatenated image as input to the one or more first portions of the one or more neural networks.

13. A method comprising:

using an amount by which a first image has changed from a second image to select whether to use one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils within the first image, wherein one or more second portions of the one or neural networks are used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image when the amount is below a threshold amount.

14. The method of claim 13 , wherein using the one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils further comprises passing a full eye image from one or more images of one or more eyes of one or more users through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

15. The method of claim 13 , wherein using the one or more second portions of one or more neural networks to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image further comprises passing smaller sub-regions of a full eye image through the one or more second portions of one or more neural networks, wherein the full eye image is the prior image that has passed through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

16. The method of claim 13 , wherein the one or more first portions of one or more neural networks and the one or more second portions of one or more neural networks use a same convolutional neural network (CNN)-based architecture but with different network parameters.

17. The method of claim 13 , wherein

threshold amount comprises an amount of lateral movement of a pupil in an image.

18. The method of claim 13 , wherein the second image includes one or more separate images for both eyes of one of one or more users, and further comprising concatenating those separate images into a concatenated image with padding between representations of the eyes before providing the concatenated image as input to the one or more first portions of the one or more neural networks.

19. A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use an amount by which a first image has changed from a second image to select whether to use one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils within the first image, wherein one or more second portions of the one or neural networks are used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image when the amount is below a threshold amount.

20. The non-transitory machine-readable medium of claim 19 , wherein the one or more processors to use one or more first portions of the one or more neural networks to generate a coarse estimate of a location of one or more pupils further comprises passing a full eye image from one or more images of one or more eyes of one or more users through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

21. The non-transitory machine-readable medium of claim 19 , wherein the one or more processors to use one or more second portions of the one or more neural networks used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image further comprises passing smaller sub-regions of a full eye image through the one or more second portions of one or more neural networks, wherein the full eye image is the prior image that has passed through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

22. The non-transitory machine-readable medium of claim 19 , wherein the one or more first portions of one or more neural networks and the one or more second portions of one or more neural networks use a same convolutional neural network (CNN)-based architecture but with different network parameters.

23. The non-transitory machine-readable medium of claim 19 , wherein the

threshold amount comprises an amount of lateral movement of a pupil in an image.

24. The non-transitory machine-readable medium of claim 20 , wherein the second image includes one or more separate images for both eyes of one of one or more users, and wherein the one or more processors are further to concatenate those separate images into a concatenated image with padding between representations of the eyes before providing the concatenated image as input to the one or more first portions of the one or more neural networks.

25. A gaze estimation system, comprising:

one or more processors to use an amount by which a first image has changed from a second image to select whether to use one or more first portions of one or more neural networks to generate a coarse estimate of a location of one or more pupils within the first image, wherein one or more second portions of the one or neural networks are used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image when the amount is below a threshold amount.

26. The gaze estimation system of claim 25 , wherein the one or more processors to use one or more first portions of the one or more neural networks to generate a coarse estimate of a location of one or more pupils further comprises passing a full eye image from one or more images of one or more eyes of one or more users through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

27. The gaze estimation system of claim 25 , wherein the one or more processors to use one or more second portions of the one or more neural networks used to generate a fine estimate of the location of the one or more pupils within the first image using a prior estimate from the second image further comprises passing smaller sub-regions of a full eye image through the one or more second portions of one or more neural networks, wherein the full eye image is the prior image that has passed through the one or more first portions of one or more neural networks to generate smaller sub-regions of the full eye image.

28. The gaze estimation system of claim 25 , wherein the one or more first portions of one or more neural networks and the one or more second portions of one or more neural networks use a same convolutional neural network (CNN)-based architecture but with different network parameters.

29. The gaze estimation system of claim 25 , wherein the threshold amount comprises an amount of lateral movement of a pupil in an image.

30. The gaze estimation system of claim 25 , wherein the second image includes one or more separate images for both eyes of one of one or more users, and wherein the one or more processors are further to concatenate those separate images into a concatenated image with padding between representations of the eyes before providing the concatenated image as input to the one or more first portions of the one or more neural networks.

31. The one or more processors of claim 1 , wherein the circuitry to select whether to use the one or more first portions of the one or more neural networks, is further to use only the one or more second portions of the one or more neural networks based on the threshold amount, wherein the threshold amount comprises the amount by which the pupil location changes from the first image to the second image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2020
From: STENGEL, MICHAEL; MCGUIRE, MORGAN; MAJERCIK, ALEXANDER; LUEBKE, DAVID
To: NVIDIA CORPORATION
Reel/Frame 052639/0906 →
Continuity (1)
Related Publication 20210350550A1 · Nov 11, 2021
References Cited (34)
US 7415126B2 · Breed · 2008 [cited by examiner]
US 9898082B1 · Greenwald · 2018 [cited by examiner]
US 10204264B1 · Gallagher · 2019 [cited by examiner]
US 11074714B2 · De Villers-Sidani · 2021 [cited by examiner]
US 11321865B1 · Kim et al. · 2022 [cited by applicant]
US 20160328015A1 · Ha · 2016 [cited by examiner]
US 20170046616A1 · Socher · 2017 [cited by examiner]
US 20170188823A1 · Ganesan · 2017 [cited by examiner]
US 20170372487A1 · Lagun · 2017 [cited by examiner]
US 20180314325A1 · Gibson · 2018 [cited by examiner]
US 20190259174A1 · De · 2019 [cited by examiner]
US 20190303724A1 · Linden · 2019 [cited by examiner]
US 20200005511A1 · Kavidayal · 2020 [cited by examiner]
US 20200129063A1 · McGrath et al. · 2020 [cited by applicant]
US 20200348755A1 · Gebauer · 2020 [cited by examiner]
CN 106575357A · 2017 [cited by applicant]
CN 110309906A · 2019 [cited by applicant]
WO 2015154882A1 · 2015 [cited by applicant]
WO 2019154511A1 · 2019 [cited by applicant]
WO WO2019147677A1 · 2019 [cited by examiner]
Krafka, Kyle, et al. “Eye tracking for everyone.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. (Year: 2016). [cited by examiner]
Google translation of CA3038584A1 (Year: 2019). [cited by examiner]
Fuhl et al., “PupilNet: Convolutional Neural Networks for Robust Pupil Detection,” Jan. 19, 2016, 10 pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Patent Application No. 2106670.9 dated Feb. 8, 2022, 12 pages. [cited by applicant]
Xu et al., “Accelerating Convolutional Neural Networks for Continuous Mobile Vision via Cache Reuse,” Dec. 1, 2017, 12 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106670.9, mailed Jun. 2, 2023, 5 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106670.9, mailed Mar. 8, 2023, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110494125.4, mailed Jun. 19, 2024, 21 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106670.9, mailed Jul. 5, 2024, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110494125.4, mailed Feb. 28, 2025, 13 pages. [cited by applicant]
Notice of Intenttion to Grant for United Kingdom Application No. GB2103370.9, mailed Nov. 13, 2024, 2 pages. [cited by applicant]
Office Action for Chinese Application No. 202110494125.4, mailed Nov. 29, 2024, 13 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202110494125.4, mailed Apr. 30, 2025, 15 pages. [cited by applicant]
Cited By (1)
US 12,669,867