IP Library Granted Patent US 12,406,487
Granted Patent B2
US 12,406,487 · App. 18/006,078 · Granted Sep 2, 2025

Systems and methods for training machine-learned visual attention models

Inventors: Xuhui Jia (Seattle, WA); Raviteja Vemulapalli (Seattle, WA); Bradley Ray Green (Bellevue, WA); Bardia Doosti (Bloomington, IN); Ching-Hui Chen (Shoreline, WA); Yukon Zhu (Shoreline, WA)
Assignee: GOOGLE LLC
G06V10/82G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,487
App. No.
18/006,078
Granted
Sep 2, 2025
Kind
B2
Abstract

Systems and methods of the present disclosure are directed to a method for training a machine-learned visual attention model. The method can include obtaining image data that depicts a head of a person and an additional entity. The method can include processing the image data with an encoder portion of the visual attention model to obtain latent head and entity encodings. The method can include processing the latent encodings with the visual attention model to obtain a visual attention value and processing the latent encodings with a machine-learned visual location model to obtain a visual location estimation. The method can include training the models by evaluating a loss function that evaluates differences between the visual location estimation and a pseudo visual location label derived from the image data and between the visual attention value and a ground truth visual attention label.

Claims (86)

1. A computer-implemented method for training a machine-learned visual attention model, the method comprising:

obtaining, by a computing system comprising one or more computing devices, image data and an associated ground truth visual attention label, wherein the image data depicts at least a head of a person and an additional entity;

processing, by the computing system, the image data with an encoder portion of the machine-learned visual attention model to obtain a latent head encoding and a latent entity encoding;

processing, by the computing system, the latent head encoding and the latent entity encoding with the machine-learned visual attention model to obtain a visual attention value indicative of whether a visual attention of the person is focused on the additional entity;

processing, by the computing system, the latent head encoding and the latent entity encoding with a machine-learned three-dimensional visual location model to obtain a three-dimensional visual location estimation, wherein the three-dimensional visual location estimation comprises an estimated three-dimensional spatial location of the visual attention of the person;

evaluating, by the computing system, a loss function that evaluates a difference between the three-dimensional visual location estimation and a pseudo visual location label derived from the image data and a difference between the visual attention value and the ground truth visual attention label; and

respectively adjusting, by the computing system, one or more parameters of the machine-learned visual attention model and the machine-learned three-dimensional visual location model based at least in part on the loss function.

2. The computer-implemented method of claim 1 , wherein:

the head of the person and the additional entity are respectively defined within the image data by a head bounding box and an entity bounding box;

obtaining, by the computing system, the image data further comprises generating, by the computing system, a spatial encoding feature vector based at least in part on a plurality of image data characteristics of the image data, wherein the spatial encoding feature vector comprises a two-dimensional spatial encoding and a three-dimensional spatial encoding; and

the spatial encoding feature vector is input alongside the latent space head encoding and the latent space entity encoding to the machine-learned visual attention model to obtain the visual attention value.

3. The computer-implemented method of claim 2 , wherein:

the two-dimensional spatial encoding describes one or more of the plurality of image data characteristics; and

the plurality of image data characteristics comprise:

respective two-dimensional location coordinates within the image data for each of the head bounding box and the entity bounding box; and

a height value and a width value of the image data.

4. The computer-implemented method of claim 2 , wherein:

the plurality of image data characteristics comprise:

respective two-dimensional location coordinates within the image data for each of the head bounding box and the entity bounding box;

an estimated camera focal length corresponding to the image data, wherein the estimated camera focal length is based at least in part on a height value and a width value of the image data;

respective depth estimates for each of the head of the person and the entity, wherein the respective estimated depths are based at least in part on the estimated camera focal length; and

the three-dimensional spatial encoding describes a pseudo three-dimensional relative position of both the head of the person and the additional entity.

5. The computer-implemented method of claim 4 , wherein the pseudo visual location label is based at least in part on the three-dimensional spatial encoding.

6. The computer-implemented method of claim 1 , wherein the additional entity comprises at least a portion of:

an object;

a person;

a direction;

a machine-readable visual encoding;

a surface; or

a space.

7. The computer-implemented method of claim 1 , wherein:

the additional entity comprises a head of a second person; and

the visual attention value is indicative of whether both the visual attention of the person is focused on the head of the second person and a visual attention of the second person is focused on the head of the person.

8. The computer-implemented method of claim 7 , wherein the three-dimensional visual location estimation comprises the estimated three-dimensional spatial location of the visual attention of the person and an estimated three-dimensional spatial location of the visual attention of the second person.

9. The computer-implemented method of claim 1 , wherein the visual attention value is a binary value.

10. The computer-implemented method of claim 1 , wherein at least one of the machine-learned visual attention model or the machine-learned three-dimensional visual location model comprises one or more convolutional neural networks.

11. The computer-implemented method of claim 1 , wherein:

the additional entity comprises a head of a second person; and

the method further comprises:

obtaining, by the computing system, second image data depicting at least a third head of a third person and a fourth head of a fourth person;

processing, by the computing system, the second image data with the machine-learned visual attention model to obtain a second visual attention value, wherein the second visual attention value is indicative of whether both a visual attention of the third person is focused on the fourth person and a visual attention of the fourth person is focused on the third person; and

determining, by the computing system based at least in part on the visual attention value, that the third person and the fourth person are looking at each other.

12. A computing system for visual attention tasks, comprising:

one or more processors;

a machine-learned visual attention model, the machine-learned visual attention model configured to:

receive image data depicting at least a head of a person and an additional entity; and

generate, based on the image data, a visual attention value, wherein the visual attention value is indicative of whether a visual attention of the person is focused on the additional entity; and

one or more tangible, non-transitory computer readable media storing computer-readable instructions that when executed by the one or more processors cause the one or more processors to perform operations, the operations comprising:

obtaining image data that depicts the at least the head of the person and the additional entity;

processing the image data with the machine-learned visual attention model to obtain the visual attention value indicative of whether the visual attention of the person is focused on the additional entity, wherein the machine-learned visual attention model is trained based at least in part on an output of a machine-learned three-dimensional visual location model, wherein the output of the machine-learned three-dimensional visual location model comprises an estimated three-dimensional spatial location of the visual attention of at least the person; and

determining, based at least in part on the visual attention value, whether the person is looking at the additional entity.

13. The computing system of claim 12 , wherein the additional entity comprises:

an object;

a person;

a direction;

a machine-readable visual encoding

a surface; or

a space.

14. The computing system of claim 12 , wherein determining, based at least in part on the visual attention value, whether the person is looking at the additional entity comprises determining, based at least in part on the visual attention value, that the person is not looking at the additional entity.

15. The computing system of claim 12 , wherein:

the additional entity comprises at least a head of a second person;

the visual attention value is indicative of whether both the visual attention of the person is focused on the head of the second person and a visual attention of the second person is focused on the head of the person; and

determining, based at least in part on the visual attention value, whether the person is looking at the additional entity comprises determining, based at least in part on the visual attention value, that the person and the second person are looking at each other.

16. The computing system of claim 15 , wherein the operations further comprise labeling the image data with a label that indicates that the image data depicts two people looking at each other.

17. The computing system of claim 12 , wherein:

the additional entity is a moving vehicle; and

the operations further comprise providing, based at least in part on whether the person is looking at the additional entity, one or more instructions configured to execute an action.

18. The computing system of claim 12 , wherein:

the additional entity is a machine-readable visual encoding descriptive of one or more actions performable by the computing system; and

the operations further comprise performing the one or more actions indicated by the machine-readable visual encoding.

19. The computing system of claim 18 , wherein the one or more actions comprise at least one of:

providing data to a second computing system via one or more networks;

retrieving data from the second computing system via the one or more networks;

establishing a secure connection with the second computing system, the secure connection configured to facilitate one or more secure transactions;

providing image data to a display device of a user of the computing system, the image data depicting at least one of:

one or more augmented reality objects, wherein the one or more augmented reality objects are two-dimensional or three-dimensional;

one or more two-dimensional images;

one or more portions of text;

a virtual reality environment;

a webpage; or

a video; or

providing data to a computing device of the user of the computing system.

20. One or more tangible, non-transitory computer readable media storing computer-readable instructions that when executed by one or more processors cause the one or more processors to perform operations, the operations comprising:

obtaining image data that depicts at least a head of a person and an additional entity;

processing the image data with a machine-learned visual attention model to obtain a visual attention value, wherein the visual attention value is indicative of whether a visual attention of the person is focused on the additional entity, wherein the machine-learned visual attention model is trained based at least in part on an output of a three-dimensional visual location model, wherein the output of the three-dimensional visual location model comprises an estimated three-dimensional spatial location of the visual attention of at least the person; and

determining, based at least in part on the visual attention value, whether the person is looking at the additional entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2023
From: JIA, XUHUI; VEMULAPALLI, RAVITEJA; ZHU, YUKUN; GREEN, BRADLEY RAY; DOOSTI, BARDIA; CHEN, CHING-HUI
To: GOOGLE LLC
Reel/Frame 063062/0758 →
Continuity (1)
Related Publication 20230281979A1 · Sep 7, 2023
References Cited (32)
US 8830164B2 · Sakata · 2014 [cited by examiner]
US 10269177B2 · Frueh · 2019 [cited by examiner]
US 10353465B2 · Qin et al. · 2019 [cited by applicant]
US 10740985B2 · Sommerlade · 2020 [cited by examiner]
US 10922889B2 · Palos · 2021 [cited by examiner]
US 11527082B2 · Arora · 2022 [cited by examiner]
US 11694419B2 · Haro · 2023 [cited by examiner]
US 11704931B2 · Aleem · 2023 [cited by examiner]
US 11886634B2 · Arar · 2024 [cited by examiner]
US 12184931B2 · Takaki · 2024 [cited by examiner]
US 12277267B2 · Silva · 2025 [cited by examiner]
US 12277738B2 · Ghebremusse · 2025 [cited by examiner]
US 20190051057A1 · Sommerlade · 2019 [cited by examiner]
US 20190212815A1 · Zhang · 2019 [cited by examiner]
US 20240242458A1 · Sommerlade et al. · 2024 [cited by applicant]
CN 111259713 · 2020 [cited by applicant]
EP 2515206A1 · 2012 [cited by examiner]
KR 20160031183 · 2016 [cited by applicant]
WO WO2020242085A1 · 2020 [cited by examiner]
Asteriadis et al., “Visual Focus of Attention in Non-calibrated Environments using Gaze Estimation” (pp. 293-316). (Year: 2014). [cited by examiner]
Hillaire et al., “Using a Visual Attention Model to Improve Gaze Tracking Systems in Interactive 3D Applications” (pp. 1830-1841). (Year: 2010). [cited by examiner]
Smith et al., “Tracking the Visual Focus of Attention for a Varying Number of Wandering People” (pp. 1-15). (Year: 2007). [cited by examiner]
Machine Translated Chinese Search Report Corresponding to Application No. 2020801043046 on Dec. 12, 2024. [cited by applicant]
Marin-Jimenez et al., “LAEO-Net: Revisiting People Looking At Each Other in Videos”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 9 pages. [cited by applicant]
Zhang et al., “It's Written All Over Your Face: Full-Face Appearance-Based Gaze Estimation”, IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017, 27 pages. [cited by applicant]
Aung et al., “Who Are They Looking At? Automatic Eye Gaze Following for Classroom Observation Video Analysis”, 2018, Database Compendex, Proceedings of the 11 [cited by applicant]
Chong et al., “Detecting Attended Visual Targets in Video”, Jun. 13, 2020, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5395-5405, XP033805577. [cited by applicant]
International Search Report for Application No. PCT/US2020/044717, mailed on Apr. 26, 2021, 2 pages. [cited by applicant]
Lian et al., “Believe it or Not, We Know What You Are Looking At!”, arxiv.org, Jul. 4, 2019, Cornell University Library, XP081438057. [cited by applicant]
Mukherjee et al., “Deep Head Pose: Gaze-Direction Estimation in Multimodal Video”, IEEE Transactions on Multimedia, IEEE Service Center, vol. 17, No. 11, Nov. 1, 2015, pp. 2094-2107, XP011588123. [cited by applicant]
Ping et al., “Where and Why are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks”, Jun. 18, 2018, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6801-6809, XP033473598. [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2020/044717, mailed Feb. 16, 2023, 10 pages. [cited by applicant]