IP Library › Granted Patent US 11,676,370
Granted Patent B2
US 11,676,370 · App. 17/317,202 · Granted Jun 13, 2023

Self-supervised cross-video temporal difference learning for unsupervised domain adaptation

Inventors: Gaurav Sharma (Newark, CA); Jinwoo Choi (Blacksburg, VA)
Assignee: NEC Corporation
G06V10/765G06F18/211G06F18/217G06F18/2155G06F18/24G06F18/253G06N3/04G06N3/08G06V10/764G06V10/7753G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,370
App. No.
17/317,202
Granted
Jun 13, 2023
Kind
B2
Abstract

A method is provided for Cross Video Temporal Difference (CVTD) learning. The method adapts a source domain video to a target domain video using a CVTD loss. The source domain video is annotated, and the target domain video is unannotated. The CVTD loss is computed by quantizing clips derived from the source and target domain videos by dividing the source domain video into source domain clips and the target domain video into target domain clips. The CVTD loss is further computed by sampling two clips from each of the source domain clips and the target domain clips to obtain four sampled clips including a first source domain clip, a second source domain clip, a first target domain clip, and a second target domain clip. The CVTD loss is computed as |(second source domain clip−first source domain clip)−(second target domain clip−first target domain clip)|.

Claims (36)

1. A computer-implemented method for Cross Video Temporal Difference (CVTD) learning for unsupervised domain adaptation, comprising:

adapting a source domain video to a target domain video using a CVTD loss, wherein the source domain video is annotated, and the target domain video is unannotated,

wherein the CVTD loss is computed by

quantizing clips derived from the source domain video and the target domain video by dividing the source domain video into a plurality of source domain clips and the target domain video into a plurality of target domain clips;

sampling two clips from each of the plurality of source domain clips and the plurality of target domain clips to obtain four sampled clips comprising a first sampled source domain clip, a second sampled source domain clip, a first sampled target domain clip, and a second sampled target domain clip; and

computing, by a clip encoder convolutional neural network, the CVTD loss as |(second sampled source domain clip−first sampled source domain clip)−(second sampled target domain clip−first sampled target domain clip)|.

2. The computer-implemented method of claim 1 , further comprising concatenating, by the clip encoder convolutional neural network, features of the four sampled clips to obtain a set of concatenated features used to calculate the CVTD loss.

3. The computer-implemented method of claim 1 , wherein the CVTD loss is an L 2 loss between a predicted CVTD value versus a true CVTD value for the four sampled clips.

4. The computer-implemented method of claim 1 , wherein the source domain video is adapted to the target domain video further using a source classification loss that inputs the plurality of source domain clips and optimized predictions of classes thereon.

5. The computer-implemented method of claim 4 , wherein the source classification loss is a cross entropy loss.

6. The computer-implemented method of claim 4 , wherein the source classification loss comprises per class sigmoid losses.

7. The computer-implemented method of claim 1 , wherein the source domain video is adapted to the target domain video further using a domain adversarial loss.

8. The computer-implemented method of claim 1 , wherein the CVTD loss is applied to learn representations capable of predicting temporal differences between the plurality of source domain clips and the plurality of target domain clips.

9. The computer-implemented method of claim 1 , wherein said sampling step samples the two clips from each of the plurality of source domain clips and the plurality of target domain clips to obtain the four sampled clips using importance-based sampling that excludes background only images.

10. The computer-implemented method of claim 9 , wherein during training, a relative importance of the pluralities of source and target domain clips are updated based on a cross-entropy of the plurality of source domain videos and an entropy of predictions of the plurality of target domain videos.

11. A computer program product for Cross Video Temporal Difference (CVTD) learning for unsupervised domain adaptation, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

adapting a source domain video to a target domain video using a CVTD loss, wherein the source domain video is annotated, and the target domain video is unannotated,

wherein the CVTD loss is computed by

quantizing clips derived from the source domain video and the target domain video by dividing the source domain video into a plurality of source domain clips and the target domain video into a plurality of target domain clips;

sampling two clips from each of the plurality of source domain clips and the plurality of target domain clips to obtain four sampled clips comprising a first sampled source domain clip, a second sampled source domain clip, a first sampled target domain clip, and a second sampled target domain clip; and

computing, by a clip encoder convolutional neural network, the CVTD loss as |(second sampled source domain clip−first sampled source domain clip)−(second sampled target domain clip−first sampled target domain clip)|.

12. The computer program product of claim 11 , wherein the method further comprises concatenating, by the clip encoder convolutional neural network, features of the four sampled clips to obtain a set of concatenated features used to calculate the CVTD loss.

13. The computer program product of claim 11 , wherein the CVTD loss is an L 2 loss between a predicted CVTD value versus a true CVTD value for the four sampled clips.

14. The computer program product of claim 11 , wherein the source domain video is adapted to the target domain video further using a source classification loss that inputs the plurality of source domain clips and optimized predictions of classes thereon.

15. The computer program product of claim 14 , wherein the source classification loss is a cross entropy loss.

16. The computer program product of claim 14 , wherein the source classification loss comprises per class sigmoid losses.

17. The computer program product of claim 11 , wherein the source domain video is adapted to the target domain video further using a domain adversarial loss.

18. The computer program product of claim 11 , wherein the CVTD loss is applied to learn representations capable of predicting temporal differences between the plurality of source domain clips and the plurality of target domain clips.

19. The computer program product of claim 11 , wherein said sampling step samples the two clips from each of the plurality of source domain clips and the plurality of target domain clips to obtain the four sampled clips using importance-based sampling that excludes background only images.

20. A computer processing system for Cross Video Temporal Difference (CVTD) learning for unsupervised domain adaptation, comprising:

a memory device for storing program code; and

a hardware processor operatively coupled to the memory device for running the program code to adapt a source domain video to a target domain video using a CVTD loss, wherein the source domain video is annotated, and the target domain video is unannotated,

wherein the hardware processor computes the CVTD loss by

quantizing clips derived from the source domain video and the target domain video by dividing the source domain video into a plurality of source domain clips and the target domain video into a plurality of target domain clips;

sampling two clips from each of the plurality of source domain clips and the plurality of target domain clips to obtain four sampled clips comprising a first sampled source domain clip, a second sampled source domain clip, a first sampled target domain clip, and a second sampled target domain clip; and

computing, using a clip encoder convolutional neural network, the CVTD loss as |(second sampled source domain clip−first sampled source domain clip)−(second sampled target domain clip−first sampled target domain clip)|.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2023
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 063213/0514 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: SHARMA, GAURAV; CHOI, JINWOO
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 056201/0937 →
Continuity (2)
Provisional Application 63030336 · May 27, 2020
Related Publication 20210374481A1 · Dec 2, 2021