IP Library Granted Patent US 12,566,493
Granted Patent B2
US 12,566,493 · App. 18/221,227 · Granted Mar 3, 2026

Methods and systems for eye-gaze location detection and accurate collection of eye-gaze data

Inventor: Aric Katz (Haifa, IL)
Assignee: Amplio Learning Technologies Holdings LLC.
G06F3/013G06T7/70G06V10/25G06V40/161G10L15/02G06T2207/30201G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,493
App. No.
18/221,227
Granted
Mar 3, 2026
Kind
B2
Abstract

Disclosed herein are methods and system for eye-gaze location detection and accurate collection of eye-gaze data from low resolution, low frame rate cameras. The method includes calibrating a machine learning model by constraining an area of interest to a single axis. Presenting to the user one or more lines of words and/or images on a screen, capturing by a camera one or more images of the user's eye-gaze, looking at said lines of words and/or images, applying the calibrated machine learning model on the one or more images of the user's eye-gaze, constraining the area of interest to a single-axis, and detecting an eye-gaze location on the screen with a word-level accuracy. The method further includes combining eye-gaze and audio data collection by measuring a time difference between a first timestamp of a word eye-gaze location detection and a second timestamp of the word first phoneme pronunciation by a user.

Claims (50)

1 . A method for eye-gaze location detection, comprising the steps of:

calibrating one or more machine learning models trained to receive one or more images of an eye-gaze of a user on a screen and detect the eye-gaze location on the screen, by constraining an area of interest to a single axis, wherein calibrating the one or more machine learning models comprises:

presenting to the user a plurality of single-axis calibration points on the screen;

for each single-axis calibration point:

capturing by a camera at least one image of the user's eye-gaze looking at the single-axis calibration point; and

providing the at least one image to the corresponding machine learning model, thereby calibrating each machine learning model to be constrained to a single-axis area of interest;

calibrating a first machine learning model with x-axis calibration points and a second machine learning model with y-axis calibration points;

presenting to the user one or more lines of words and/or images on the screen;

capturing by the camera one or more images of the user's eye-gaze, looking at said lines of words and/or images;

applying the first and second calibrated machine learning models in parallel on the one or more images of the user's eye gaze looking at said lines of words and/or images, constraining the area of interest to an x-axis for the first machine learning model and a y-axis for the second machine learning model;

detecting an eye-gaze location on the screen with a word-level accuracy for each of the first and second machine learning models; and

combining the detected eye-gaze locations from the first and second machine learning models to obtain a single detected eye-gaze location on the screen with increased accuracy.

2 . The method of claim 1 , wherein the one or more machine learning models are configured to be trained on two-dimensional data.

3 . The method of claim 1 , wherein the one or more machine learning models are configured to be trained on a single-axis data.

4 . The method of claim 1 , wherein the word-level accuracy is between 2-3 cm.

5 . The method of claim 1 , wherein the first machine learning model is configured to be trained on the x-axis and the second machine learning model is configured to be trained on the y-axis.

6 . The method of claim 1 , wherein the camera is a low resolution and low frame rate camera.

7 . The method of claim 1 , wherein the camera is a web camera.

8 . The method of claim 1 , further comprising combining eye-gaze and audio data collection, comprising:

presenting a text to the user with a plurality of lines of words and/or images;

recording a speech sample of the user reading at least one of the presented lines;

measuring a first timestamp of a word detected by the eye-gaze location, and a second timestamp of a first phoneme of the word pronunciation; and

comparing the first and second timestamps of the word, providing the time difference between eye-gaze detection and the first phoneme pronunciation, thereby allowing inferring and/or detecting regression disorders.

9 . The method of claim 8 , wherein the regression disorders are ADD or ADHD, dyslexia, Reading Comprehension Deficits, Visual Processing Disorders and Reading Speed and Fluency Disorders.

10 . The method of claim 1 , further comprising the step of detecting face position by a facial posture framework and providing the face position as an input to the one or more machine learning models, instead of the step of calibrating the one or more machine learning models.

11 . A system for eye-gaze detection, comprising:

a screen;

a camera configured to capture images of a user's eye-gaze looking at the screen; and

a processor executing a code configured to:

calibrate one or more machine learning models trained to receive one or more images of an eye-gaze of the user on the screen and detect the eye-gaze location on the screen, by constraining an area of interest to a single axis, wherein calibrating the one or more machine learning models comprises:

presenting to the user a plurality of single-axis calibration points on the screen;

for each single-axis calibration point:

capturing by the camera at least one image of the user's eye-gaze looking at the single-axis calibration point; and

providing the at least one image to the corresponding machine learning model, thereby calibrating each machine learning model to be constrained to a single-axis area of interest;

wherein a first machine learning model is calibrated with x-axis calibration points and a second machine learning model is calibrated with y-axis calibration points;

present to the user one or more lines of words and/or images on the screen;

capture by the camera one or more images of the user's eye-gaze, looking at said lines of words and/or images;

apply the first and second calibrated machine learning models in parallel on the one or more captured images of the user's eye gaze looking at said lines of words and/or images, constraining the area of interest to an x-axis for the first machine learning model and a y-axis for the second machine learning model;

detect an eye-gaze location on the screen with a word-level accuracy for each of the first and second machine learning models; and

combine the detected eye-gaze locations from the first and second machine learning models to obtain a single detected eye-gaze location on the screen with increased accuracy.

12 . The system of claim 11 , wherein the one or more machine learning models were trained on two-dimensional data.

13 . The system of claim 11 , wherein the one or more machine learning models were trained on single axis data.

14 . The system of claim 11 , wherein the camera is a low resolution and low frame rate camera.

15 . The system of claim 11 , wherein the camera is a web camera.

16 . The system of claim 11 , wherein the processor further executes a code for combining eye-gaze and audio data collection, the code is configured to:

present a text to the user with a plurality of lines of words and/or images;

record a speech sample of the user reading at least one of the presented lines;

measuring a first timestamp of a word detected by the eye-gaze location, and a second timestamp of a first phoneme of the word pronunciation; and

compare the first and second timestamps of the word, providing the time difference between eye-gaze detection and the first phoneme pronunciation, thereby allowing inferring and/or detecting regression disorders.

17 . The system of claim 11 , further comprising a facial posture framework for detecting face position wherein the code is configured to receive the face position as an input to the one or more machine learning models, instead of calibrating the one or more machine learning models.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2024
From: AMPLIO LEARNING TECHNOLOGIES LTD.
To: AMPLIO LEARNING TECHNOLOGIES HOLDINGS LLC.
Reel/Frame 068812/0325 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2023
From: KATZ, ARIC
To: AMPLIO LEARNING TECHNOLOGIES LTD.
Reel/Frame 064460/0009 →
Continuity (2)
Provisional Application 63388760 · Jul 13, 2022
Related Publication 20240019931A1 · Jan 18, 2024
References Cited (6)
US 4950069A · Hutchinson · 1990 [cited by examiner]
US 11619993B2 · Sharma · 2023 [cited by examiner]
US 20140226131A1 · Lopez · 2014 [cited by examiner]
US 20200364539A1 · Anisimov · 2020 [cited by examiner]
US 20220300072A1 · Arar · 2022 [cited by examiner]
Sinha et al., “Eyegaze Tracking In Handheld Devices,” 2018 Fifth International Conference on Emerging Applications of Information Technology (EAIT), Kolkata, India, 2018, pp. 1-5, doi: 10.1109/EAIT.2018.8470402 (Year: 2… [cited by examiner]