IP Library Granted Patent US 12,266,157
Granted Patent B2
US 12,266,157 · App. 17/712,617 · Granted Apr 1, 2025

Temporal augmentation for training video reasoning system

Inventors: Farley Lai (Santa Clara, CA); Asim Kadav (Mountain View, CA)
Assignee: NEC Corporation
G06V10/7747G06V10/62G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,157
App. No.
17/712,617
Granted
Apr 1, 2025
Kind
B2
Abstract

A method for augmenting video sequences in a video reasoning system is presented. The method includes randomly subsampling a sequence of video frames captured from one or more video cameras, randomly reversing the subsampled sequence of video frames to define a plurality of sub-sequences of randomly reversed video frames, training, in a training mode, a video reasoning model with temporally augmented input, including the plurality of sub-sequences of randomly reversed video frames, to make predictions over temporally augmented target classes, updating parameters of the video reasoning model by a machine leaning algorithm, and deploying, in an inference mode, the video reasoning model in the video reasoning system to make a final prediction related to a human action in the sequence of video frames.

Claims (37)

1. A method for augmenting video sequences in a video reasoning system, the method comprising:

randomly subsampling a sequence of video frames captured from one or more video cameras;

randomly reversing the subsampled sequence of video frames to define a plurality of sub-sequences of randomly reversed video frames;

training, in a training mode, a video reasoning model with temporally augmented input, including the plurality of sub-sequences of randomly reversed video frames, to make predictions over temporally augmented target classes;

updating parameters of the video reasoning model by a machine learning algorithm by discarding a final prediction to retain the temporally augmented target classes with corresponding original classes; and

deploying, in an inference mode, the video reasoning model, with overfitting suppression through a learned temporal order of video frames to determine the original classes, in the video reasoning system to make the final prediction related to classify a human action in the sequence of video frames.

2. The method of claim 1 , wherein a target class is offset by a total number of original classes when the subsampled sequence of video frames is randomly reversed.

3. The method of claim 1 , wherein the reversing of the subsampled sequence of video frames is implemented to classify a doubled number of class categories.

4. The method of claim 1 , wherein the video reasoning model learns from temporal ordering for classification.

5. The method of claim 1 , wherein taking a softmax of logits of original classes only is employed to arrive at the final prediction.

6. The method of claim 1 , wherein combining logits of original classes and corresponding augmented classes for a softmax value is employed to arrive at the final prediction.

7. The method of claim 1 , wherein the final prediction is discarded if a top class does not belong to any original classes.

8. A non-transitory computer-readable storage medium comprising a computer-readable program for explaining sensor time series data in natural language, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of:

randomly subsampling a sequence of video frames captured from one or more video cameras;

randomly reversing the subsampled sequence of video frames to define a plurality of sub-sequences of randomly reversed video frames;

training, in a training mode, a video reasoning model with temporally augmented input, including the plurality of sub-sequences of randomly reversed video frames, to make predictions over temporally augmented target classes;

updating parameters of the video reasoning model by a machine learning algorithm by discarding a final prediction to retain the temporally augmented target classes with corresponding original classes; and

deploying, in an inference mode, the video reasoning model, with overfitting suppression through a learned temporal order of video frames to determine the original classes, in a video reasoning system to make the final prediction related to classify a human action in the sequence of video frames.

9. The non-transitory computer-readable storage medium of claim 8 , wherein a target class is offset by a total number of original classes when the subsampled sequence of video frames is randomly reversed.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the reversing of the subsampled sequence of video frames is implemented to classify a doubled number of class categories.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the video reasoning model learns from temporal ordering for classification.

12. The non-transitory computer-readable storage medium of claim 8 , wherein taking a softmax of logits of original classes only is employed to arrive at the final prediction.

13. The non-transitory computer-readable storage medium of claim 8 , wherein combining logits of original classes and corresponding augmented classes for a softmax value is employed to arrive at the final prediction.

14. The non-transitory computer-readable storage medium of claim 8 , wherein the final prediction is discarded if a top class does not belong to any original classes.

15. A system for explaining sensor time series data in natural language, the system comprising:

a memory; and

one or more processors in communication with the memory configured to:

randomly subsample a sequence of video frames captured from one or more video cameras;

randomly reverse the subsampled sequence of video frames to define a plurality of sub-sequences of randomly reversed video frames;

train, in a training mode, a video reasoning model with temporally augmented input, including the plurality of sub-sequences of randomly reversed video frames, to make predictions over temporally augmented target classes;

update parameters of the video reasoning model by a machine learning algorithm by discarding a final prediction to retain the temporally augmented target classes with corresponding original classes; and

deploy, in an inference mode, the video reasoning model, with overfitting suppression through a learned temporal order of video frames to determine the original classes, in the video reasoning system to make the final prediction related to classify a human action in the sequence of video frames.

16. The system of claim 15 , wherein a target class is offset by a total number of original classes when the subsampled sequence of video frames is randomly reversed.

17. The system of claim 15 , wherein the reversing of the subsampled sequence of video frames is implemented to classify a doubled number of class categories.

18. The system of claim 15 , wherein the video reasoning model learns from temporal ordering for classification.

19. The system of claim 15 , wherein taking a softmax of logits of original classes only is employed to arrive at the final prediction.

20. The system of claim 15 , wherein combining logits of original classes and corresponding augmented classes for a softmax value is employed to arrive at the final prediction.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 070345/0687 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2022
From: LAI, FARLEY; KADAV, ASIM
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 059490/0758 →
Continuity (2)
Provisional Application 63171215 · Apr 6, 2021
Related Publication 20220319157A1 · Oct 6, 2022
References Cited (10)
US 20100067865A1 · Saxena et al. · 2010 [cited by applicant]
US 20200242425A1 · Yamada · 2020 [cited by examiner]
US 20220303560A1 · Sridhar · 2022 [cited by examiner]
D. Wei, J. Lim, A. Zisserman and W. T. Freeman, “Learning and Using the Arrow of Time,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 8052-8060, doi: 10.1109/CVP… [cited by examiner]
W. Li, W. Nie and Y. Su, “Human Action Recognition Based on Selected Spatio-Temporal Features via Bidirectional LSTM,” in IEEE Access, vol. 6, pp. 44211-44220, 2018, doi: 10.1109/ACCESS.2018.2863943 (Year: 2018). [cited by examiner]
B. Fernando, H. Bilen, E. Gavves and S. Gould, “Self-Supervised Video Representation Learning with Odd-One-Out Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, … [cited by examiner]
Xu et al., Quadratic Video Interpolation, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. (Year: 2019). [cited by examiner]
Dwibedi et al., Counting OUt Time: Class Agnostic Video Repetition Counting in the Wild, 2020 (Year: 2020). [cited by examiner]
Falcon, Alex, Oswald Lanz, and Giuseppe Serra. “Data augmentation techniques for the Video Question Answering task.” In European Conference on Computer Vision, pp. 511-525. Springer, Cham, 2020. [cited by applicant]
Wang, Jason, and Luis Perez. “The effectiveness of data augmentation in image classification using deep learning.” Convolutional Neural Networks Vis. Recognit 11 (2017): 1-8. [cited by applicant]