IP Library Granted Patent US 12,266,174
Granted Patent B2
US 12,266,174 · App. 17/862,667 · Granted Apr 1, 2025

Few-shot action recognition

Inventors: Biplob Debnath (Princeton, NJ); Srimat Chakradhar (Manalapan, NJ); Oliver Po (San Jose, CA); Asim Kadav (Mountain View, CA); Farley Lai (Santa Clara, CA); Farhan Asif Chowdhury (Albuquerque, NM)
Assignee: NEC Corporation
G06V20/41G06N3/08G06V10/764G06V10/82G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,174
App. No.
17/862,667
Granted
Apr 1, 2025
Kind
B2
Abstract

Methods and systems of training a neural network include training a feature extractor and a classifier using a first set of training data that includes one or more base cases. The classifier is trained with few-shot adaptation using a second set of training data, smaller than the first set of training data, while keeping parameters of the feature extractor constant.

Claims (20)

1. A computer-implemented method of training a neural network, comprising:

training a feature extractor and a classifier, the feature extractor including a time segmentation network, using a first set of training data that includes one or more base cases, wherein the training data includes video and the time segmentation network extracts features from a plurality of frames from the video and uses segmental consensus to generate a feature vector corresponding to the plurality of frames; and

training the classifier with few-shot adaptation using a second set of training data, smaller than the first set of training data, while keeping parameters of the feature extractor constant.

2. The method of claim 1 , wherein the feature extractor is a video feature extractor and the classifier is an action identification model, and wherein the first set of training data includes base case training examples and the second set of training data includes novel training examples for an action that is not represented in the first set of training data.

3. The method of claim 1 , wherein the segmental consensus is a pooling operation.

4. The method of claim 1 , wherein the feature extractor includes a temporal shift.

5. The method of claim 4 , wherein the training data includes video, and wherein the temporal shift shifts information for a channel of a first frame of the video to a second frame of the video.

6. The method of claim 1 , wherein the classifier performs linear classification.

7. The method of claim 1 , wherein the classifier performs cosine similarity-based classification.

8. A system for training a neural network, comprising:

a hardware processor; and

a memory that stores a computer program, which, when executed by the hardware processor, causes the hardware processor to:

train a feature extractor and a classifier, the feature extractor including a time segmentation network, using a first set of training data that includes one or more base cases, wherein the training data includes video and the time segmentation network extracts features from a plurality of frames from the video and uses segmental consensus to generate a feature vector corresponding to the plurality of frames; and

train the classifier with few-shot adaptation using a second set of training data, smaller than the first set of training data, while keeping parameters of the feature extractor constant.

9. The system of claim 8 , wherein the feature extractor is a video feature extractor and the classifier is an action identification model, and wherein the first set of training data includes base case training examples and the second set of training data includes novel training examples for an action that is not represented in the first set of training data.

10. The system of claim 8 , wherein the segmental consensus is a pooling operation.

11. The system of claim 8 , wherein the feature extractor includes a temporal shift.

12. The system of claim 11 , wherein the training data includes video, and wherein the temporal shift shifts information for a channel of a first frame of the video to a second frame of the video.

13. The system of claim 8 , wherein the classifier performs linear classification.

14. The system of claim 8 , wherein the classifier performs cosine similarity-based classification.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 070345/0687 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2022
From: DEBNATH, BIPLOB; CHAKRADHAR, SRIMAT; PO, OLIVER; KADAV, ASIM; LAI, FARLEY; CHOWDHURY, FARHAN ASIF
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 060482/0982 →
Continuity (3)
Provisional Application 63220623 · Jul 12, 2021
Related Publication 20230008303A1 · Jan 12, 2023
Related Publication 20230049770A1 · Feb 16, 2023
References Cited (14)
US 9330171B1 · Shetty · 2016 [cited by examiner]
US 20210366126A1 · Chen · 2021 [cited by examiner]
US 20220036538A1 · Steiman · 2022 [cited by examiner]
US 20220129677A1 · Kale · 2022 [cited by examiner]
US 20220172700A1 · Xiong · 2022 [cited by examiner]
US 20220248296A1 · Merwaday · 2022 [cited by examiner]
US 20220253729A1 · Vashist · 2022 [cited by examiner]
US 20220272255A1 · Xiong · 2022 [cited by examiner]
Zhu et al., “Few-shot Action Recognition with Prototype-centered Attentive Learning.” arXiv:2101.08085v4 [cs.CV], Mar. 28, 2021, pp. 1-10. [cited by applicant]
Cao et al., “Few-Shot Video Classification via Temporal Alignment”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020 (pp. 10618-10627). [cited by applicant]
Lin et al., “TSM: Temporal Shift Module for Efficient Video Understanding”, InProceedings of the IEEE/CVF International Conference on Computer Vision, Nov. 2019 (pp. 7083-7093). [cited by applicant]
Zhu et al., “Compound Memory Networks for Few-shot Video Classification”, InProceedings of the European Conference on Computer Vision (ECCV), Sep. 2018 (pp. 751-766). [cited by applicant]
Perrett et al., “Temporal-Relational Cross Transformers for Few-Shot Action Recognition”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, (pp. 475-484). [cited by applicant]
Chen et al., “A Closer Look at Few-Shot Classification”, arXiv:1904.04232v2 [cs.CV] Jan. 12, 2020, pp. 1-17. [cited by applicant]