IP Library › Granted Patent US 12,254,691
Granted Patent B2
US 12,254,691 · App. 17/111,352 · Granted Mar 18, 2025

Cooperative-contrastive learning systems and methods

Inventors: Nishant Rai (Stanford, CA); Ehsan Adeli Mosabbeb (Menlo Park, CA); Kuan-Hui Lee (Los Altos, CA); Adrien Gaidon (Los Altos, CA); Juan Carlos Niebles (Mountain View, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
G06V20/41G06N20/00G06V30/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,691
App. No.
17/111,352
Granted
Mar 18, 2025
Kind
B2
Abstract

Systems and methods for multi-view cooperative contrastive self-supervised learning, may include receiving a plurality of video sequences, the video sequences comprising a plurality of image frames; applying selected images of a first and second video sequence of the plurality of video sequences to a plurality of different encoders to derive a plurality of embeddings for different views of the selected images of the first and second video sequences; determining distances of the derived plurality of embeddings for the selected images of the first and second video sequences; detecting inconsistencies in the determined distances; and predicting semantics of a future image based on the determined distances.

Claims (36)

1. A method for multi-view self-supervised learning, comprising:

receiving a plurality of video sequences, the video sequences comprising a plurality of image frames;

applying selected images of a first and second video sequence of the plurality of video sequences to a plurality of different encoders to derive a plurality of embeddings for different views of the selected images of the first and second video sequences, the plurality of embeddings comprising RGB embeddings, flow embeddings, and KeyPoint embeddings;

determining distances of the derived plurality of embeddings for the selected images of the first and second video sequences;

detecting inconsistencies between distances of the RGB embeddings, distances of the flow embeddings, and distances of the KeyPoint embeddings outside a threshold distance; and

predicting semantics of a future image based on the determined distances.

2. The method of claim 1 , further comprising partitioning the received plurality of video sequences into disjoint blocks; and wherein applying selected images of a first and second video sequence of the plurality of video sequences to a plurality of respective encoders to derive a plurality of embeddings, comprises:

transforming the disjoint blocks into corresponding latent representations for the disjoint blocks; and

generating context representations from the latent representations.

3. The method of claim 2 , wherein predicting future semantics based on the determined distances comprises capturing contextual semantics and frame level semantics to predict a latent state of future images.

4. The method of claim 1 , further comprising using predicted semantics of a future image to predict semantics of an image that is farther into the future than the future image.

5. The method of claim 1 , further comprising computing a noise contrastive estimation loss over at least some of the plurality of embeddings.

6. The method of claim 5 , wherein the noise contrastive estimation loss comprises an entropy loss distinguishing a positive pair of embeddings from negative pairs of embeddings in a video sequence.

7. The method of claim 1 , wherein the plurality of different encoders for the selected images of the first and second video sequences comprise at least two of a flow encoder, an RGB encoder and a keypoint encoder.

8. The method of claim 1 , wherein determining distances comprises determining view-specific distances for the different views, and wherein the method further comprises synchronizing the distances across all views.

9. The method of claim 8 , further comprising enforcing a consistency loss between distances from each of the different views.

10. The method of claim 8 , further comprising deriving discriminative scores based on pairs of embeddings that comprise positive and negative distances.

11. A system for multi-view self-supervised learning, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations, the operations comprising:

receiving a plurality of video sequences, the video sequences comprising a plurality of image frames;

applying selected images of a first and second video sequence of the plurality of video sequences to a plurality of different encoders to derive a plurality of embeddings for different views of the selected images of the first and second video sequences, the plurality of different encoders comprising RGB embeddings, flow embeddings, and KeyPoint embeddings;

determining distances of the derived plurality of embeddings for the selected images of the first and second video sequences;

detecting inconsistencies between distances of RGB embeddings, distances of the flow embeddings, and distances of KeyPoint embeddings outside a threshold distance; and

predicting semantics of a future image based on the determined distances.

12. The system of claim 11 , wherein the operations further comprise partitioning the received plurality of video sequences into disjoint blocks; and wherein applying selected images of a first and second video sequence of the plurality of video sequences to a plurality of respective encoders to derive a plurality of embeddings, comprises:

transforming the disjoint blocks into corresponding latent representations for the disjoint blocks; and

generating context representations from the latent representations.

13. The system of claim 12 , wherein predicting future semantics based on the determined distances comprises capturing contextual semantics and frame level semantics to predict a latent state of future images.

14. The system of claim 11 , wherein the operations further comprise using predicted semantics of a future image to predict semantics of an image that is farther into the future than the future image.

15. The system of claim 11 , wherein the operations further comprise computing a noise contrastive estimation loss over at least some of the plurality of embeddings.

16. The system of claim 15 , wherein the noise contrastive estimation loss comprises an entropy loss distinguishing a positive pair of embeddings from negative pairs of embeddings in a video sequence.

17. The system of claim 11 , wherein the plurality of different encoders for the selected images of the first and second video sequences comprise at least two of a flow encoder, an RGB encoder and a keypoint encoder.

18. The system of claim 11 , wherein determining distances comprises determining view-specific distances for the different views, and wherein the operations further comprise synchronizing the distances across all views.

19. The system of claim 18 , wherein the operations further comprise enforcing a consistency loss between distances from each of the different views.

20. The system of claim 18 , wherein the operations further comprise deriving discriminative scores based on pairs of embeddings that comprise positive and negative distances.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 071027/0780 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2021
From: RAI, NISHANT; MOSABBEB, EHSAN ADELI; NIEBLES, JUAN CARLOS
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 054922/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: LEE, KUAN-HUI; GAIDON, ADRIEN
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 054539/0519 →
Continuity (1)
Related Publication 20220180101A1 · Jun 9, 2022
References Cited (13)
US 20180124425A1 · Van Leuven · 2018 [cited by applicant]
US 20200202074A1 · Ghulati · 2020 [cited by applicant]
US 20200211206A1 · Wang · 2020 [cited by examiner]
US 20220044056A1 · Liu · 2022 [cited by examiner]
US 20220114365A1 · Lukác · 2022 [cited by examiner]
Tschannen et al., “Self-Supervised Learning of Video-Induced Visual Invariances,” Google Research, Brain Team, arXiv:1912.02783v2 [cs.CV]; Apr. 1, 2020, 18 pages. [cited by applicant]
Tian et al., “Contrastive Multiview Coding,” arXiv:1906.05849v4 [cs.CV]; Mar. 11, 2020, pp. 1-16. [cited by applicant]
Luo et al., “Unsupervised Learning of Long-Term Motion Dynamics for Videos,” Stanford University, arXiv:1701.01821v3 [cs.CV] Apr. 11, 2017, pp. 1-10. [cited by applicant]
Mathieu et al., “Deep multi-scale video prediction beyond mean square error,” https://arxiv.org/pdf/1511.05440.pdf%5D; Feb. 26, 2016, pp. 1-14. [cited by applicant]
Han et al., “Video Representation Learning by Dense Predictive Coding,” https://arxiv.org/pdf/1909.04656.pdf, Sep. 27, 2019, pp. 1-13. [cited by applicant]
Vondrick et al., “Generating Videos with Scene Dynamics,” https://arxiv.org/pdf/1609.02612.pdf; Oct. 26, 2016, pp. 1-10. [cited by applicant]
Bojar, Daniel, “Using Deep Learning to Classify Relationship State with DeepConnection,” https://towardsdatascience.com/using-deep-learning-to-classifyrelationship-state-with-deepconnection-227e9124c72?gi=50c3b167e4d3; … [cited by applicant]
Sayed et al., “Cross and Learn: Cross-Modal Self-Supervision,” https://arxiv.org/pdf/1811.03879.pdf; Apr. 29, 2019, pp. 1-15. [cited by applicant]