IP Library Granted Patent US 12,212,813
Granted Patent B2
US 12,212,813 · App. 17/620,082 · Granted Jan 28, 2025

Systems and methods for characterizing joint attention during real world interaction

Inventors: Leanne Chukoskie (San Diego, CA); Pamela Cosman (La Jolla, CA); Pranav Venuprasad (Framingham, MA); Anurag Paul (Sunnyvale, CA); Tushar Dobhal (Bentonville, AR); Tu Nguyen (Tay Ninh Province, VN); Andrew Gilman (Tauranga, NZ)
Assignee: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
H04N21/44218G06F3/013
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,212,813
App. No.
17/620,082
Granted
Jan 28, 2025
Kind
B2
Abstract

Systems, devices, and methods are disclosed for characterizing joint attention. A method includes dynamically obtaining video streams of participants; dynamically obtaining gaze streams; dynamically providing a cue to the participants to view the object; dynamically detecting a joint gaze based on the gaze streams focusing on the object over a time interval; and dynamically providing feedback based on detecting the joint gaze.

Claims (66)

1. A computer-implemented method comprising:

obtaining a first gaze stream, wherein the first gaze stream comprises first gaze location data that identifies where a first pupil of a first participant is looking in a real-world environment as a function of time and a first gaze confidence score specifying an accuracy of the first gaze location data;

obtaining a second gaze stream, wherein the second gaze stream comprises second gaze location data that identifies where a second pupil of the first participant is looking in the real-world environment as a function of time and a second gaze confidence score specifying an accuracy of the second gaze location data;

based on the first and second gaze confidence scores, rectifying the first and second gaze location data into a first rectified gaze stream that identifies where the first participant is looking in the real-world environment as a function of time;

obtaining a third gaze stream, wherein the third gaze stream comprises third gaze location data that identifies where a first pupil of a second participant is looking in the real-world environment as a function of time and a third gaze confidence score specifying an accuracy of the third gaze location data;

obtaining a fourth gaze stream, wherein the fourth gaze stream comprises fourth gaze location data that identifies where a second pupil of the second participant is looking in the real-world environment as a function of time and a fourth gaze confidence score specifying an accuracy of the fourth gaze location data;

based on the third and fourth gaze confidence scores, rectifying the third and fourth gaze location data into a second rectified gaze stream that identifies where the second participant is looking in the real-world environment as a function of time;

providing a cue to the first and second participants to view an object in the real-world environment;

detecting joint attention between the first participant and the second participant based on the rectified first and second gaze streams indicating that both the first participant and the second participant are focusing on the object over a threshold time interval within a time window of predetermined length; and

providing feedback based on detecting the joint attention.

2. The computer-implemented method of claim 1 , further comprising filtering the first gaze location data based on a Kalman filter and the first gaze confidence score.

3. The computer-implemented method of claim 1 , further comprising:

generating a visualization of a bounding box surrounding the object; and

displaying the visualization on a head-mounted display when the object is in a current frame of a video stream of the real-world environment from a perspective of the first participant.

4. The computer-implemented method of claim 1 , further comprising:

detecting the object in a current frame of a video stream of the real-world environment from a perspective of the first participant using a convolutional neural network;

generating a visualization of a bounding box surrounding the object in the current frame; and

displaying the visualization on a head-mounted display when the object is in the current frame.

5. The computer-implemented method of claim 1 , wherein the cue comprises one of an audio cue and a visual cue.

6. The computer-implemented method of claim 1 , wherein the feedback comprises one of an audio cue and a visual cue.

7. A system comprising:

one or more processors; and

non-transitory storage medium coupled to the one or more processors to store instructions, which when executed by the one or more processors, cause the system to:

obtain a first gaze stream, wherein the first gaze stream comprises first gaze location data that identifies where a first pupil of a first participant is looking in a real-world environment as a function of time and a first gaze confidence score specifying an accuracy of the first gaze location data;

obtaining a second gaze stream, wherein the second gaze stream comprises second gaze location data that identifies where a second pupil of the first participant is looking in the real-world environment as a function of time and a second gaze confidence score specifying an accuracy of the second gaze location data;

based on the first and second gaze confidence scores, rectify the first and second gaze location data into a first rectified gaze stream that identifies where the first participant is looking in the real-world environment as a function of time;

obtain a third gaze stream, wherein the third gaze stream comprises third gaze location data that identifies where a first pupil of a second participant is looking in the real- world environment as a function of time and a third gaze confidence score specifying an accuracy of the third gaze location data;

obtain a fourth gaze stream, wherein the fourth gaze stream comprises fourth gaze location data that identifies where a second pupil of the second participant is looking in the real-world environment as a function of time and a fourth gaze confidence score specifying an accuracy of the fourth gaze location data;

based on the third and fourth gaze confidence scores, rectify the third and fourth gaze location data into a second rectified gaze stream that identifies where the second participant is looking in the real-world environment as a function of time;

provide a cue to the first and second participants to view an object in the real-world environment;

detect joint attention between the first participant and the second participant based on the rectified first and second gaze streams indicating that both the first participant and the second participant are focusing on the object over a threshold time interval within a time window of predetermined length; and

provide feedback based on detecting the joint attention.

8. The system of claim 7 , wherein the non-transitory storage medium is coupled to the one or more processors to store additional instructions, which when executed by the one or more processors, further cause the system to:

filter the first gaze location data based on a Kalman filter and the first gaze confidence score; and

filter the second gaze location data based on a Kalman filter and the second gaze confidence score.

9. The system of claim 7 , wherein the non-transitory storage medium is coupled to the one or more processors to store additional instructions, which when executed by the one or more processors, further cause the system to:

generate a visualization of a bounding box surrounding the object; and

display the visualization on a see-through display when the object is in a current frame of a video stream of the real-world environment from a perspective of the first participant.

10. The system of claim 7 , wherein the non-transitory storage medium is coupled to the one or more processors to store additional instructions, which when executed by the one or more processors, further cause the system to:

detect the object in a current frame of a video stream of the real-world environment from a perspective of the first participant using a convolutional neural network;

generate a visualization of a bounding box surrounding the object in the current frame; and

display the visualization on a see-through display when the object is in the current frame.

11. The system of claim 7 , wherein the non-transitory storage medium is coupled to the one or more processors to store additional instructions, which when executed by the one or more processors, cause the system to detect a gaze triad from at least one participant of the first and second participant, wherein the gaze triad comprises a look to a face before and after a look to the object.

12. The system of claim 7 , wherein detecting the joint attention comprises applying a runlength filter, wherein the runlength filter comprises a window size parameter specifying a number of frames to consider together and a hit threshold specifying a minimum number of hits in the number of frames to be identified as a gaze.

13. A computer-implemented method comprising:

obtaining a first video stream comprising a video of a real-world environment from a perspective of a first participant;

obtaining a second video stream comprising a video of a real-world environment from a perspective of a second participant;

a first gaze stream, wherein the first gaze stream comprises first gaze location data that identifies where a first pupil of the first participant is looking in a real-world environment as a function of time and a first gaze confidence score specifying an accuracy of the first gaze location data;

obtaining a second gaze stream, wherein the second gaze stream comprises second gaze location data that identifies where a second pupil of the first participant is looking in the real-world environment as a function of time and a second gaze confidence score specifying an accuracy of the second gaze location data;

based on the first and second gaze confidence scores, rectifying the first and second gaze location data into a first rectified gaze stream that identifies where the first participant is looking in the real-world environment as a function of time;

obtaining a third gaze stream, wherein the third gaze stream comprises third gaze location data that identifies where a first pupil of the second participant is looking in the real-world environment as a function of time and a third gaze confidence score specifying an accuracy of the third gaze location data;

obtaining a fourth gaze stream, wherein the fourth gaze stream comprises fourth gaze location data that identifies where a second pupil of the second participant is looking in the real-world environment as a function of time and a fourth gaze confidence score specifying an accuracy of the fourth gaze location data;

based on the third and fourth gaze confidence scores, rectifying the third and fourth gaze location data into a second rectified gaze stream that identifies where the second participant is looking in the real-world environment as a function of time;

providing a cue to the first and second participants to view an object in the real-world environment;

detecting joint attention between the first participant and the second participant based on the rectified first and second gaze streams indicating that both the first participant and the second participant are focusing on the object over a threshold time interval within a time window of predetermined length; and

providing feedback based on detecting the joint attention.

14. The computer-implemented method of claim 13 , further comprising:

generating a bounding box surrounding the object;

displaying the bounding box via a first see-through display when the object is in a current frame of the first video stream; and

displaying the bounding box via a second see-through display when the object is in a current frame of the second video stream.

15. The computer-implemented method of claim 14 , wherein displaying the bounding box via the first see-through display when the object is in the current frame of the first video stream comprises changing a size of the bounding box based on at least one of:

a current size of the bounding box;

the first gaze confidence score; and

a depth of the object from the first participant.

16. The computer-implemented method of claim 13 , further comprising detecting a gaze triad, wherein the gaze triad comprises a look to a face before and after a look to the object.

17. The computer-implemented method of claim 13 , further comprising filtering the first gaze location data based on a Kalman filter and the first gaze confidence score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: CHUKOSKIE, LEANNE; COSMAN, PAMELA; VENUPRASAD, PRANAV; PAUL, ANURAG; DOBHAL, TUSHAR; NGUYEN, TU; GILMAN, ANDREW
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 058532/0968 →
Continuity (2)
Provisional Application 62866573 · Jun 25, 2019
Related Publication 20220360847A1 · Nov 10, 2022
References Cited (14)
US 20020115925A1 · Avrin · 2002 [cited by examiner]
US 20050004496A1 · Pilu · 2005 [cited by examiner]
US 20060262140A1 · Kujawa · 2006 [cited by examiner]
US 20130278631A1 · Border et al. · 2013 [cited by applicant]
US 20140361971A1 · Sala · 2014 [cited by examiner]
US 20150268821A1 · Ramsby · 2015 [cited by examiner]
US 20160262613A1 · Klin · 2016 [cited by examiner]
US 20170004573A1 · Hussain · 2017 [cited by examiner]
US 20170061034A1 · Ritchey et al. · 2017 [cited by applicant]
US 20180088340A1 · Amayeh · 2018 [cited by examiner]
US 20180239523A1 · Rosenberg · 2018 [cited by examiner]
US 20200097076A1 · Alcaide · 2020 [cited by examiner]
International Search Report and Written Opinion mailed Sep. 4, 2020 for International Application No. PCT/US2020/039536. [cited by applicant]
International Preliminary Report on Patentability dated May 21, 2021 for International Application No. PCT/US2020/039536. [cited by applicant]