IP Library Granted Patent US 12,541,249
Granted Patent B2
US 12,541,249 · App. 18/279,117 · Granted Feb 3, 2026

Implicit calibration from screen content for gaze tracking

Inventors: Dmitry Lagun (Mountain View, CA); Gautam Prasad (Mountain View, CA); Pezhman Firoozfam (Mountain View, CA); Jimin Pi (Mountain View, CA)
Assignee: GOOGLE LLC
G06F3/013G06T7/70G06T11/60G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,249
App. No.
18/279,117
Granted
Feb 3, 2026
Kind
B2
Abstract

The technology relates to methods and systems for implicit calibration for gaze tracking. This can include receiving, by a neural network module, display content that is associated with presentation on a display screen ( 1202 ). The neural network module may also receive uncalibrated gaze information, in which the uncalibrated gaze information includes an uncalibrated gaze trajectory that is associated with a viewer gaze of the display content on the display screen ( 1204 ). A selected function is applied by the neural network module to the uncalibrated gaze information and the display content to generate a user-specific gaze function ( 1206 ). The user-specific gaze function has one or more personalized parameters. And the neural network module can then apply the user-specific gaze function to the uncalibrated gaze information to generate calibrated gaze information associated with the display content on the display screen ( 1208 ). Training and testing information may alternatively be created for implicit gaze calibration ( 1000 ).

Claims (34)

1 . A computer-implemented method of performing implicit gaze calibration for gaze tracking, the method comprising:

receiving, by a neural network module, display content that is associated with presentation on a display screen;

receiving, by the neural network module, uncalibrated gaze information, the uncalibrated gaze information including an uncalibrated gaze trajectory that is associated with a viewer gaze of the display content on the display screen;

applying, by the neural network module, a selected function to the uncalibrated gaze information and the display content to generate a user-specific gaze function, the user-specific gaze function having one or more personalized parameters; and

applying, by the neural network module, the user-specific gaze function to the uncalibrated gaze information to generate calibrated gaze information associated with the display content on the display screen.

2 . The method of claim 1 , wherein the selected function is a linear or polynomial function.

3 . The method of claim 1 , wherein the uncalibrated gaze information further includes timestamp information for when the display content was collected.

4 . The method of claim 1 , wherein the uncalibrated gaze information further includes at least one of screen orientation information, camera focal length, aspect ratio, or resolution information.

5 . The method of claim 1 , wherein the one or more personalized parameters of the user-specific gaze function are estimated from collected data.

6 . The method of claim 1 , wherein applying the selected function to the uncalibrated gaze information and the display content to generate a user-specific gaze function includes:

generating temporal information and dimensional information at an encoder block of the neural network;

generating a context vector from the temporal information and the dimensional information in a self-attention block of the neural network; and

applying the context vector to the uncalibrated gaze information to generate the calibrated gaze information, in a decoder block of the neural network.

7 . The method of claim 6 , wherein the temporal information encompasses a selected time interval associated with a gaze along the display screen.

8 . The method of claim 7 , wherein the temporal information is encoded by looking through an entire sequence of gaze measurements and screen content pixels associated with the entire sequence.

9 . The method of claim 6 , wherein applying the context vector to the uncalibrated gaze information comprises multiplying the context vector with an array of data from the uncalibrated gaze information.

10 . The method of claim 6 , wherein applying the context vector to the uncalibrated gaze information includes applying the uncalibrated gaze information and the context vector using a plurality of fully connected layers of the neural network.

11 . The method of claim 1 , wherein the display content comprises synthetic content.

12 . The method of claim 11 , wherein the synthetic content includes at least one of synthetic text or synthetic graphical information.

13 . The method of claim 11 , wherein the synthetic content corresponds to a dataset of gaze trajectories for a group of users over a selected number of unique user interfaces.

14 . The system of claim 1 , wherein application of the selected function to the uncalibrated gaze information and the display content to generate a user-specific gaze function includes:

generation of temporal information and dimensional information at an encoder block of the neural network;

generation of a context vector from the temporal information and the dimensional information in a self-attention block of the neural network; and

application of the context vector to the uncalibrated gaze information to generate the calibrated gaze information, in a decoder block of the neural network.

15 . A system comprising one or more processors and one or more storage devices storing instructions, wherein when the instructions are executed by the one or more processors, the one or more processors implement a method of implicit gaze calibration for gaze tracking comprising:

receiving, by a neural network module of the one or more processors, display content that is associated with presentation on a display screen;

receiving, by the neural network module, uncalibrated gaze information, the uncalibrated gaze information including an uncalibrated gaze trajectory that is associated with a viewer gaze of the display content on the display screen;

applying, by the neural network module, a selected function to the uncalibrated gaze information and the display content to generate a user-specific gaze function, the user-specific gaze function having one or more personalized parameters; and

applying, by the neural network module, the user-specific gaze function to the uncalibrated gaze information to generate calibrated gaze information associated with the display content on the display screen.

16 . The system of claim 15 , wherein the uncalibrated gaze information further includes timestamp information for when the display content was collected.

17 . The system of claim 15 , wherein the uncalibrated gaze information further includes at least one of screen orientation information, camera focal length, aspect ratio, or resolution information.

18 . The system of claim 15 , wherein the one or more personalized parameters of the user-specific gaze function are estimated from collected data.

19 . The system of claim 15 , wherein the display content comprises synthetic content.

20 . The system of claim 15 , wherein the synthetic content either includes at least one of synthetic text or synthetic graphical information, or corresponds to a dataset of gaze trajectories for a group of users over a selected number of unique user interfaces.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: LAGUN, DMITRY; PRASAD, GAUTAM; FIROOZFAM, PEZHMAN; PI, JIMIN
To: GOOGLE LLC
Reel/Frame 064733/0595 →
Continuity (1)
Related Publication 20240126365A1 · Apr 18, 2024
References Cited (18)
US 10127680B2 · Lagun et al. · 2018 [cited by applicant]
US 10379612B1 · Bonnier · 2019 [cited by examiner]
US 10846877B2 · Lagun et al. · 2020 [cited by applicant]
US 10890969B2 · Yuan et al. · 2021 [cited by applicant]
US 20140282646A1 · McCoy · 2014 [cited by examiner]
US 20180059782A1 · San Agustin Lopez et al. · 2018 [cited by applicant]
US 20190018483A1 · Aleem et al. · 2019 [cited by applicant]
US 20190064513A1 · Bagherpour · 2019 [cited by examiner]
US 20190259174A1 · De et al. · 2019 [cited by applicant]
US 20210255485A1 · Xu · 2021 [cited by examiner]
CN 111949131A · 2020 [cited by applicant]
WO 2016111880A1 · 2016 [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US21/28367 dated Mar. 2, 2022 (21 pages). [cited by applicant]
Partial Search Report and Provisional Opinion for Application No. PCT/US21/28367 dated Jan. 31, 2022. [cited by applicant]
“Gaze Interaction Library”, Gaze Interaction Library—Windows Community Toolkit | Microsoft Docs, https://docs.microsoft.com/en-us/windows/communitytoolkit/gaze/gazeinteractionlibrary, Jan. 22, 2021, pp. 1-10. [cited by applicant]
“Tracking and Visualizing Faces”, Tracking and Visualizing Faces | Apple Developer Documentation, https://developer.apple.com/documentation/arkit/tracking_and_visualizing_faces#see-also, Jan. 22, 2021, pp. 1-6. [cited by applicant]
Montabone, Sebastian , et al., “Human Detection Using a Mobile Platform and Novel Features Derived From a Visual Saliency Mechanism”, Preprint submitted to Image and Vision Computing, Jan. 9, 2009, 28 pgs. [cited by applicant]
Park, Seonwook , et al., “Towards End-to-end Video-based Eye-Tracking”, arXiv:2007.13120v1 [cs.CV] Jul. 26, 2020, 27 pgs. [cited by applicant]