IP Library Granted Patent US 12,670,722
Granted Patent B2
US 12,670,722 · App. 18/077,974 · Granted Jun 30, 2026

Self-supervised compositional feature representation for video understanding

Inventors: Zhipeng Bao (Pittsburgh, PA); Pavel Tokmakov (Santa Monica, CA); Adrien David Gaidon (San Jose, CA); Allan Jabri (Toronto, CA); Yuxiong Wang (Champaign, IL); Martial Hebert (Pittsburgh, PA)
Assignee: TOYOTA JIDOSHA KABUSHIKI KAISHA
G06V20/58G06T7/215G06T7/246G06V10/7715G06V10/7753G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,722
App. No.
18/077,974
Filed
Dec 8, 2022
Granted
Jun 30, 2026
Kind
B2
Art Unit
2667
USPC
382/103
Abstract

A method of compositional feature representation learning for video understanding is described. The method includes individually processing a sequence of video frames received as an input of a feature map network to generate a plurality of feature maps. The method also includes binding the plurality of feature maps to a fixed set of slot variables using an attention model according to a motion segmentation signal. The method further includes combining slot states corresponding to the fixed set of slot variables into a combined feature map. The method also includes decoding the combined feature map to form a reconstructed sequence of video frames, in which objects discovered in the reconstructed sequence of video frames are identified.

Claims (30)

1 . A method of compositional feature representation learning for video understanding, comprising:

individually processing a sequence of video frames received as an input of a feature map network to generate a plurality of feature maps;

binding the plurality of feature maps to a fixed set of slot variables using an attention model according to a motion segmentation signal;

combining slot states corresponding to the fixed set of slot variables into a combined feature map;

decoding the combined feature map to form a reconstructed sequence of video frames, in which dynamic objects discovered in the reconstructed sequence of video frames are identified by 2D bounding boxes estimated from 2D instance masks estimated to focus on the dynamic objects discovered instead of non-dynamic objects in the reconstructed sequence of video frames; and

controlling an ego vehicle to follow a planned trajectory adjusted for collision avoidance in response to the dynamic objects discovered in the reconstructed sequence of video frames in a scene captured by and surrounding the ego vehicle.

2 . The method of claim 1 , in which binding the plurality of feature maps comprises separating the moving objects from static objects in the plurality of feature maps based on the motion segmentation signal.

3 . The method of claim 1 , further comprising using motion supervision for training the attention model to assign the fixed set of slot variables distinguishing a foreground of a sequence of video frames from a background of the sequence of video frames.

4 . The method of claim 1 , further comprises segmenting the objects discovered in the sequence of video frames.

5 . The method of claim 1 , further comprises tracking the objects discovered in the sequence of video frames.

6 . The method of claim 1 , in which decoding the combined feature map comprises embedding features within the reconstructed sequence of video frames.

7 . A non-transitory computer-readable medium having program code recorded thereon of compositional feature representation learning for video understanding, the program code being executed by a processor and comprising:

program code to individually process a sequence of video frames received as an input of a feature map network to generate a plurality of feature maps;

program code to bind the plurality of feature maps to a fixed set of slot variables using an attention model according to a motion segmentation signal;

program code to combine slot states corresponding to the fixed set of slot variables into a combined feature map;

program code to decode the combined feature map to form a reconstructed sequence of video frames, in which dynamic objects discovered in the reconstructed sequence of video frames are identified by 2D bounding boxes estimated from 2D instance masks estimated to focus on the dynamic objects discovered instead of non-dynamic objects in the reconstructed sequence of video frames; and

program code to control an ego vehicle to follow a planned trajectory adjusted for collision avoidance in response to the dynamic objects discovered in the reconstructed sequence of video frames in a scene captured by and surrounding the ego vehicle.

8 . The non-transitory computer-readable medium of claim 7 , in which the program code to bind the plurality of feature maps comprises program code to separate the moving objects from static objects in the plurality of feature maps based on the motion segmentation signal.

9 . The non-transitory computer-readable medium of claim 7 , further comprising program code to use motion supervision for training the attention model to assign the fixed set of slot variables, distinguishing a foreground of a sequence of video frames from a background of the sequence of video frames.

10 . The non-transitory computer-readable medium of claim 7 , further comprises program code to segment the objects discovered in the sequence of video frames.

11 . The non-transitory computer-readable medium of claim 7 , further comprises program code to track the objects discovered in the sequence of video frames.

12 . The non-transitory computer-readable medium of claim 7 , in which the program code to decode the combined feature map comprises program code to embed features within the reconstructed sequence of video frames.

13 . A system of compositional feature representation learning for video understanding, the system comprising:

an individual feature map generation module to individually process a sequence of video frames received as an input of a feature map network to generate a plurality of feature maps;

an attention model to bind the plurality of feature maps to a fixed set of slot variables using the attention model according to a motion segmentation signal;

a combined feature map generation module to combine slot states corresponding to the fixed set of slot variables into a combined feature map;

an object discovery module to decode the combined feature map to form a reconstructed sequence of video frames, in which dynamic objects discovered in the reconstructed sequence of video frames are identified by 2D bounding boxes estimated from 2D instance masks estimated to focus on the dynamic objects discovered instead of non-dynamic objects in the reconstructed sequence of video frames; and

a controller to control an ego vehicle to follow a planned trajectory adjusted for collision avoidance in response to the dynamic objects discovered in the reconstructed sequence of video frames in a scene captured by and surrounding the ego vehicle.

14 . The system of claim 13 , in which the attention model is trained to separate the moving objects from static objects in the plurality of feature maps based on the motion segmentation signal.

15 . The system of claim 14 , in which the attention model is further trained to assign the fixed set of slot variables, distinguishing a foreground of a sequence of video frames from a background of the sequence of video frames using motion supervision.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 075616/0029 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2026
From: HEBERT, MARTIAL
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 074896/0864 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2026
From: WANG, YUXIONG
To: THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
Reel/Frame 074898/0376 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2026
From: JABRI, ALLAN
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 074898/0779 →
Continuity (2)
Provisional Application 63288463 · Dec 10, 2021
Related Publication 20230252796A1 · Aug 10, 2023
References Cited (21)
US 8903128B2 · Shet et al. · 2014 [cited by applicant]
US 9454819B1 · Seetharaman · 2016 [cited by examiner]
US 11816533B2 · Ren · 2023 [cited by examiner]
US 12067785B2 · Ambrus · 2024 [cited by examiner]
US 20150317519A1 · Gurbuz · 2015 [cited by examiner]
US 20200394458A1 · Yu et al. · 2020 [cited by applicant]
US 20210319242A1 · Cholakkal · 2021 [cited by examiner]
US 20210374416A1 · Zablotskaia · 2021 [cited by examiner]
US 20210383199A1 · Weissenborn · 2021 [cited by examiner]
US 20210390710A1 · Zhang · 2021 [cited by examiner]
US 20210406560A1 · Park · 2021 [cited by examiner]
US 20230123899A1 · Iqbal · 2023 [cited by examiner]
US 20230154198A1 · Makansi · 2023 [cited by examiner]
CN 111860485A · 2020 [cited by applicant]
CN 113012254A · 2021 [cited by applicant]
CN 113204010A · 2021 [cited by applicant]
WO 2021165569A1 · 2021 [cited by applicant]
Yang, Charig et al. “Self-supervised Video Object Segmentation by Motion Grouping”, Oct. 17, 2021 [retrieved on Feb. 28, 2025], 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 202… [cited by examiner]
Yang, Charig et al. “Self-supervised Video Object Segmentation by Motion Grouping”, Oct. 17, 2021, 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 7157-7168. Retrieved f… [cited by examiner]
Jabri, et al., “Space-Time Correspondence as a Contrastive Random Walk”, arXiv:2006.14613v2, Dec. 3, 2020. [cited by applicant]
Wang, et al., “Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning”, arXiv:2009.05769v4, Apr. 22, 2021. [cited by applicant]