IP Library › Granted Patent US 12,566,495
Granted Patent B2
US 12,566,495 · App. 18/511,872 · Granted Mar 3, 2026

Image analysis using gaze tracking and utterance dictation

Inventors: Junjian Huang (Birmingham, AL); Chanler Megan Crowe Cantor (Madison, AL); Steven Michael Thomas (Madison, AL)
Assignees: Intuitive Research and Technology Corporation; The UAB Research Foundation
G06F3/013G06F40/205G06V20/70G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,495
App. No.
18/511,872
Granted
Mar 3, 2026
Kind
B2
Abstract

Techniques for evaluating images are disclosed. An image is displayed. Data corresponding to a gaze fixation of a user who is viewing the image is obtained. An image artifact is identified based on the gaze fixation. The image artifact is at a location in the image where the gaze fixation occurred. A recording of the user's utterance is accessed. This recording is recorded during an overlapping time period with when the gaze fixation occurred. The recording is transcribed into text, which is then parsed. A key term is extracted from the parsed text. A determination is made as to whether the key term corresponds to the image artifact identified based on the gaze fixation.

Claims (49)

1 . A method for facilitating an evaluation of an image by comparing an identified image artifact included in the image against a gaze fixation that is directed toward the image artifact and against a recording of an utterance that is also associated with the image artifact, said method comprising:

displaying an image;

while the image is displayed, tracking a gaze of a user who is viewing the image, wherein tracking the gaze includes identifying a gaze fixation of the user with respect to the image;

providing an instruction to the user to prompt the user's gaze to follow a specific gaze pattern when the image is displayed, wherein an image artifact is overlapped by the specific gaze pattern;

determining that the image artifact is represented within the image at a location in the image where the gaze fixation occurred;

determining whether the user's gaze fixated on the image artifact while the user's gaze followed the specific gaze pattern;

while the image is displayed, recording an utterance of the user, said utterance being recorded during an overlapping time period with when the gaze fixation occurred;

transcribing the recorded utterance into text;

in response to parsing the text, extracting at least one key term from the parsed text; and

determining whether the at least one key term accurately describes the image artifact where the gaze fixation occurred relative to the image.

2 . The method of claim 1 , wherein the image is a medical image.

3 . The method of claim 1 , wherein the gaze fixation of the user occurs in response to the user's gaze residing on the image artifact for at least a threshold time period.

4 . The method of claim 3 , wherein the threshold time period is at least 500 milliseconds.

5 . The method of claim 1 , wherein tracking the user's gaze is performed via a wearable gaze detector.

6 . The method of claim 1 , wherein tracking the user's gaze is performed via a device that is remote relative to the user such that the device is not a wearable device.

7 . The method of claim 1 , wherein tracking the user's gaze is performed via a headset that renders virtualized content.

8 . The method of claim 1 , wherein tracking the user's gaze is performed at a sampling rate that is between about 30 Hertz (Hz) and about 120 Hz.

9 . The method of claim 1 , wherein said method is performed by a handheld device.

10 . The method of claim 1 , wherein the image artifact is tagged with metadata describing subject matter of the image artifact, wherein determining whether the at least one key term accurately describes the image artifact is performed by comparing the metadata against the at least one key term.

11 . The method of claim 1 , wherein the image artifact is associated with label data describing subject matter of the image artifact.

12 . The method of claim 11 , wherein the label data is generated using a machine learning engine.

13 . The method of claim 11 , wherein the label data is manually provided via user input.

14 . A computer system comprising:

one or more processors; and

one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computer system to:

display an image;

while the image is displayed, track a gaze of a user who is viewing the image, wherein tracking the gaze includes identifying a gaze fixation of the user with respect to the image;

provide an instruction to the user to prompt the user's gaze to follow a specific gaze pattern when the image is displayed, wherein an image artifact is overlapped by the specific gaze pattern;

determine that the image artifact is represented within the image at a location in the image where the gaze fixation occurred;

determine whether the user's gaze fixated on the image artifact while the user's gaze followed the specific gaze pattern;

while the image is displayed, record an utterance of the user, said utterance being recorded during an overlapping time period with when the gaze fixation occurred;

transcribe the recorded utterance into text;

in response to parsing the text, extract a key term from the parsed text; and

determine whether the key term corresponds to the image artifact, which is identified based on the gaze fixation.

15 . The computer system of claim 14 , wherein at least one of a visual cue or an audio cue is provided to the user while the user's gaze is being tracked.

16 . The computer system of claim 14 , wherein touch input is further received, and wherein the touch input, the key term, and data for the gaze fixation is used to determine whether the image artifact is accurately being described.

17 . A method comprising:

displaying an image;

obtaining data corresponding to a gaze fixation of a user who is viewing the image;

providing an instruction to the user to prompt a gaze of the user to follow a specific gaze pattern when the image is displayed, wherein an image artifact is overlapped by the specific gaze pattern;

determining that the image artifact is represented within the image and that the image artifact is at a location in the image where the gaze fixation occurred;

determining whether the user's gaze fixated on the image artifact while the user's gaze followed the specific pattern;

accessing a recording of an utterance of the user, said recording being recorded during an overlapping time period with when the gaze fixation occurred;

transcribing the recording into text;

in response to parsing the text, extracting a key term from the parsed text; and

determining whether the key term corresponds to the image artifact identified based on the gaze fixation.

18 . The method of claim 17 , wherein the method further includes providing a second instruction to prompt the user to look at a second image artifact.

19 . The method of claim 18 , wherein the method further includes determining that the user's gaze fixated on the second image artifact.

20 . The method of claim 19 , wherein the method further includes at least one of: providing an audio respond stating that the user successfully fixated on the second image artifact or providing a visual alert stating that the user successfully fixated on the second image artifact.

Assignments (3)
SECURITY INTEREST Recorded Jul 20, 2026
From: INTUITIVE RESEARCH AND TECHNOLOGY CORPORATION
To: REGIONS BANK
Reel/Frame 076014/0667 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: HUANG, JUNJIAN
To: THE UAB RESEARCH FOUNDATION
Reel/Frame 065591/0693 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: CANTOR, CHANLER MEGAN CROWE; THOMAS, STEVEN MICHAEL
To: INTUITIVE RESEARCH AND TECHNOLOGY CORPORATION
Reel/Frame 065591/0751 →
Continuity (2)
Provisional Application 63427388 · Nov 22, 2022
Related Publication 20240168549A1 · May 23, 2024
References Cited (10)
US 7203649B1 · Linebarger · 2007 [cited by examiner]
US 20070078642A1 · Weng · 2007 [cited by examiner]
US 20170262051A1 · Tall · 2017 [cited by examiner]
US 20170286398A1 · Hunt · 2017 [cited by examiner]
US 20190362557A1 · Lacey · 2019 [cited by examiner]
US 20200128231A1 · Pace · 2020 [cited by examiner]
US 20200400957A1 · Van Heugten · 2020 [cited by examiner]
US 20210034702A1 · Sizemore · 2021 [cited by examiner]
US 20220092754A1 · Hakim · 2022 [cited by examiner]
US 20240169494A1 · Strandborg · 2024 [cited by examiner]