IP Library › Granted Patent US 12,277,171
Granted Patent B2
US 12,277,171 · App. 18/179,617 · Granted Apr 15, 2025

Video retrieval techniques using video contrastive learning

Inventors: Xiao Xia Mao (Shanghai, CN); Wei Jun Zheng (Shanghai, CN); Shi Hui Gui (Shanghai, CN); Xiao Feng Ji (Shanghai, CN)
Assignee: International Business Machines Corporation
G06F16/78G06F40/30G06V10/761G06V10/774G06V10/82G06V20/70G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,171
App. No.
18/179,617
Granted
Apr 15, 2025
Kind
B2
Abstract

A method, computer system, and a computer program product are provided for training a neural network for finding queried videos. Two pairs of video clips and associated text are obtained from a first dataset and a second dataset. The first dataset is used to train two video encoders by providing the video clips to the encoders as input and providing the outputs to a cosine similarity calculator. The second dataset is used to train a multi-mentor paradigm with two mentors. A first mentor and a second mentor are each provided the pair of textual data inputs. The first mentor provides a similarity value comparison, and the second mentor provides a word mover distance. Using the output from the multi-mentor paradigm and the encoders, a contrastive loss is calculated and used to provide contrastive learning of video features by differentiating similarity and dissimilarity of the video clips.

Claims (36)

1. A method of training a neural network for finding and retrieving queried videos, comprising:

obtaining two video clips from a first dataset and providing the two video clips to two video encoders for training;

providing an output of each of the two video encoders to a cosine similarity calculator;

training a multi-mentor paradigm having at least two mentors by obtaining two textual inputs from a second dataset, wherein a first mentor is provided each textual input to provide a similarity value comparison and a second mentor is provided said two textual inputs to provide a word mover distance (WMD); and

using said output from said multi-mentor paradigm and said encoders, calculate a contrastive loss used to provide contrastive learning of video features for differentiating similarity and dissimilarity of video clips.

2. The method of claim 1 , wherein said first dataset comprises pairs of video and text clips that are overlapping on a timeline.

3. The method of claim 1 , wherein said first dataset is used to construct a positive video and text pair for contrastive learning.

4. The method of claim 1 , wherein said second dataset comprises a corpus having pairs of substantially similar texts.

5. The method of claim 4 , wherein said corpus can include an online or an offline document.

6. The method of claim 1 , wherein said first and second mentor each identify a similarity between said two inputted texts from a perspective of text semantics.

7. The method of claim 1 , wherein said multi-mentor paradigm has a feedforward network with multiple layers, and said feedforward network combines a plurality of outputs from all mentors in order to generate a label.

8. The method of claim 1 , wherein said encoders use spatio-temporal tokens within each video to provide a video feature vector used in said cosine similarity calculator.

9. The method of claim 1 , wherein said encoders are part of a Siamese neural network.

10. The method of claim 9 , wherein a same weight is given to video clips provided to said encoders to compute a comparable output vector.

11. The method of claim 10 , wherein said weights are adjusted during said training.

12. The method of claim 1 , wherein said neural network is used to find and retrieve a queried video from a corpus based on a search request.

13. A computer system for training a neural network to find and retrieve queried videos, comprising:

one or more non-transitory computer readable storage media storing program instruction to be executed by a processor;

program instructions, stored on at least one of the one or more non-transitory computer-readable storage media for execution by at least one of the one or more processors via at least one of the one or more memories to update, wherein the computer system is enabled to perform the steps:

obtaining two video clips from a first dataset and providing it to two video encoders for training;

providing an output of each encoder and providing it to a cosine similarity calculator;

training a multi-mentor paradigm having at least two mentors by obtaining two textual inputs from a second dataset; wherein a first mentor is provided each textual input to provide a similarity value comparison and said second mentor is provided said two textual inputs to provide a word mover distance (WMD);

using said output from said multi-mentor paradigm and said encoders, calculate a contrastive loss used to provide contrastive learning of video features for differentiating similarity and dissimilarity of video clips.

14. The computer system of claim 13 , wherein said first dataset comprises pairs of video and text clips that are overlapping on a timeline and said second dataset comprises a corpus having pairs of substantially similar texts.

15. The computer system of claim 13 , wherein said first and second mentor each identify similarity between said two inputted texts from a perspective of text semantics.

16. The computer system of claim 13 , wherein said multi-mentor paradigm have a feedforward network with multiple layers and said feedforward network combines a plurality of outputs from all mentors in order to generate a label.

17. A computer program product for training a neural network for finding and retrieving queried videos, comprising:

one or more non-transitory computer readable storage media storing program instruction to be executed by a processor;

program instructions, stored on at least one of the one or more non-transitory computer-readable storage media, to update the program instructions comprising:

obtaining two video clips from a first dataset and providing it to two video encoders for training;

providing an output of each encoder and providing it to a cosine similarity calculator;

training a multi-mentor paradigm having at least two mentors by obtaining two textual inputs from a second dataset; wherein a first mentor is provided each textual input to provide a similarity value comparison and said second mentor is provided said two textual inputs to provide a word mover distance (WMD);

using said output from said multi-mentor paradigm and said encoders, calculate a contrastive loss used to provide contrastive learning of video features for differentiating similarity and dissimilarity of video clips.

18. The computer program product of claim 17 , wherein said first dataset comprises pairs of video and text clips that are overlapping on a timeline and said second dataset comprises a corpus having pairs of substantially similar texts.

19. The computer program product of claim 17 , wherein said first and second mentor each identify similarity between said two inputted texts from a perspective of text semantics.

20. The computer program product of claim 17 , wherein said multi-mentor paradigm have a feedforward network with multiple layers, and said feedforward network combines a plurality of outputs from all mentors in order to generate a label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: MAO, XIAO XIA; ZHENG, WEI JUN; GUI, SHI HUI; JI, XIAO FENG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 062905/0980 →
Continuity (1)
Related Publication 20240303272A1 · Sep 12, 2024
References Cited (18)
US 10642892B2 · Xiao · 2020 [cited by applicant]
US 20070255755A1 · Zhang · 2007 [cited by applicant]
US 20200142928A1 · Mei · 2020 [cited by applicant]
US 20200372066A1 · Saggi · 2020 [cited by examiner]
US 20220309278A1 · Gan · 2022 [cited by applicant]
US 20230281247A1 · Lee · 2023 [cited by examiner]
CN 114742018A · 2022 [cited by applicant]
CN 115129934A · 2022 [cited by applicant]
CN 115408558A · 2022 [cited by applicant]
RU 2647696C2 · 2018 [cited by applicant]
WO 2017114388A1 · 2017 [cited by applicant]
Arnab, et al., “ViViT: A Video Vision Transformer”, arXiv:2103.15691v2 [cs.CV], Nov. 1, 2021, 14 pgs., <https://arxiv.org/pdf/2103.15691.pdf>. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pgs., <https://arxiv.org/pdf/1810.04805.pdf>. [cited by applicant]
Google, “Google Images”, Google.com, [accessed Jan. 13, 2023], 1 pg., Retrieved from the Internet: <https://images.google.com/?gws_rd=ssl>. [cited by applicant]
Reimers, et al., “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks”, arXiv:1908.10084v1 [cs.CL], Aug. 27, 2019, 11 pgs., <https://arxiv.org/pdf/1908.10084.pdf>. [cited by applicant]
Tolstikhin, et al., “MLP-Mixer: An all-MLP Architecture for Vision”, arXiv:2105.01601v4 [cs.CV], Jun. 11, 2021, 16 pgs., <https://arxiv.org/pdf/2105.01601.pdf>. [cited by applicant]
Xu, et al., “VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding”, ACM, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6787-6800, Nov. 7-11, 2021, <htt… [cited by applicant]
Yang, et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, arXiv:1906.08237v2 [cs.CL], Jan. 2, 2… [cited by applicant]