IP Library › Granted Patent US 11,816,889
Granted Patent B2
US 11,816,889 · App. 17/216,605 · Granted Nov 14, 2023

Unsupervised video representation learning

Inventors: Chuang Gan (Cambridge, MA); Dakuo Wang (Cambridge, MA); Antonio Jose Jimeno Yepes (Melbourne, AU); Bo Wu (Cambridge, MA)
Assignee: International Business Machines Corporation
G06V20/41G06F16/735G06F18/214G06F18/24G06N3/04G06N3/088G06V10/40G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,816,889
App. No.
17/216,605
Granted
Nov 14, 2023
Kind
B2
Abstract

Unsupervised learning for video classification. One or more features from one or more video clips are extracted using a spatial-temporal encoder. The one or more extracted features are processed, using a video instance discrimination task, to generate a classification label, the classification label indicating whether two of the video clips are from a same video. The one or more extracted features are processed, using a pair-wise speed discrimination task, to generate a comparison label, the comparison label indicating a relative playback speed between two given video clips. A search is performed in a video database for a video that is similar to a given video based on the comparison label.

Claims (34)

1. A method comprising:

extracting, using a spatial-temporal encoder, one or more features from one or more video clips;

processing, using a video instance discrimination task, the one or more extracted features to generate a classification label, the classification label indicating whether two of the video clips are from a same video;

processing, using a pair-wise speed discrimination task, the one or more extracted features to generate a comparison label, the comparison label indicating a relative playback speed between two given video clips; and

searching, in a video database, for a video that is similar to a given video clip in terms of playback speed based on the comparison label generated by the pair-wise speed discrimination task and that is from a same video as the given video clip based on the classification label generated by the video instance discrimination task.

2. The method of claim 1 , wherein the spatial-temporal encoder is based on a spatial-temporal neural network.

3. The method of claim 1 , wherein the video instance discrimination task is based on a model g a of a video instance neural network.

4. The method of claim 3 , the method further comprising training the model g a using a database of training videos and corresponding training video clips to distinguish video clips derived from the same video from video clips derived from different videos.

5. The method of claim 1 , wherein the processing, using the video instance discrimination task, the one or more extracted features further generates a loss a .

6. The method of claim 1 , wherein the pair-wise speed discrimination task is based on a model g b of a pair-wise speed discrimination neural network.

7. The method of claim 6 , the method further comprising training the model g b using a database of training videos and corresponding training video clips to identify a difference in playback speed between two video clips.

8. The method of claim 1 , wherein the processing, using the pair-wise speed discrimination task, the one or more extracted features further generates a loss m .

9. The method of claim 1 , wherein the searching operation is further based on the classification label.

10. An apparatus comprising:

a memory; and

at least one processor, coupled to said memory, and operative to perform operations of:

extracting, using a spatial-temporal encoder, one or more features from one or more video clips;

processing, using a video instance discrimination task, the one or more extracted features to generate a classification label, the classification label indicating whether two of the video clips are from a same video;

processing, using a pair-wise speed discrimination task, the one or more extracted features to generate a comparison label, the comparison label indicating a relative playback speed between two given video clips; and

searching, in a video database, for a video that is similar to a given video clip in terms of playback speed based on the comparison label generated by the pair-wise speed discrimination task and that is from a same video as the given video clip based on the classification label generated by the video instance discrimination task.

11. The apparatus of claim 10 , wherein the spatial-temporal encoder is based on a spatial-temporal neural network.

12. The apparatus of claim 10 , wherein the video instance discrimination task is based on a model g a of a video instance neural network.

13. The apparatus of claim 12 , the operations further comprising training the model g a using a database of training videos and corresponding training video clips to distinguish video clips derived from the same video from video clips derived from different videos.

14. The apparatus of claim 10 , wherein the processing, using the video instance discrimination task, the one or more extracted features further generates a loss a .

15. The apparatus of claim 10 , wherein the pair-wise speed discrimination task is based on a model g b of a pair-wise speed discrimination neural network.

16. The apparatus of claim 15 , the operations further comprising training the model g b using a database of training videos and corresponding training video clips to identify a difference in playback speed between two video clips.

17. The apparatus of claim 10 , wherein the processing, using the pair-wise speed discrimination task, the one or more extracted features further generates a loss m .

18. The apparatus of claim 10 , wherein the searching operation is further based on the classification label.

19. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method of:

extracting, using a spatial-temporal encoder, one or more features from one or more video clips;

processing, using a video instance discrimination task, the one or more extracted features to generate a classification label, the classification label indicating whether two of the video clips are from a same video;

processing, using a pair-wise speed discrimination task, the one or more extracted features to generate a comparison label, the comparison label indicating a relative playback speed between two given video clips; and

searching, in a video database, for a video that is similar to a given video clip in terms of playback speed based on the comparison label generated by the pair-wise speed discrimination task and that is from a same video as the given video clip based on the classification label generated by the video instance discrimination task.

20. The computer program product of claim 19 , wherein the searching operation is further based on the classification label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2021
From: GAN, CHUANG; WANG, DAKUO; JIMENO YEPES, ANTONIO JOSE; WU, BO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055759/0431 →
Continuity (1)
Related Publication 20220309278A1 · Sep 29, 2022