IP Library Granted Patent US 12,567,255
Granted Patent B2
US 12,567,255 · App. 18/366,931 · Granted Mar 3, 2026

Few-shot video classification

Inventors: Kai Li (Plainsboro, NJ); Renqiang Min (Princeton, NJ); Haifeng Xia (New Orleans, LA)
Assignee: NEC Corporation
G06V20/41G06V10/774G06V20/46G06V20/48
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,255
App. No.
18/366,931
Granted
Mar 3, 2026
Kind
B2
Abstract

Methods and systems for video processing include enriching an input video feature from an input video frame set using a meta-action bank video sub-actions to generate enriched features. Reinforced image representation is performed using reinforcement learning to compare support image frames and query image frames and determine an importance of the input video frame. A classification is performed on the input video frame based on the importance and the enriched features to generate a label. An action is performed responsive to the generated label.

Claims (96)

1 . A computer-implemented method for video processing, comprising:

enriching an input video feature from an input video frame set using a meta-action bank of a plurality of video sub-actions to generate enriched features, wherein enriching the video frame feature is performed as:

f

˜

*

j

s

/

q

=

(

1

-

ρ

)

f

*

j

s

/

q

+

ρ

l

=

1

κ

α

j

l

d

l

e

where p is a parameter balancing between accepting external knowledge with keeping original representations, a is a cosine similarity between the video frame feature f *j s/q , d l e is a projection coefficient on a basis within the meta-action bank e, κ is a number of bases in the meta-action bank, and s/q indicates that the term is treated the same for support frames s and query frames q;

performing reinforced image representation using reinforcement learning to compare support image frames and query image frames and determine an importance of the input video frame;

performing a classification on the input video frame based on the importance and the enriched features to generate a label; and

performing an action responsive to the generated label.

2 . The method of claim 1 , wherein the reinforcement learning uses a policy based on a feature of the input video frame, an aggregated feature of remaining frames of the input video frame set, and a separate video.

3 . The method of claim 2 , wherein the aggregated feature is a weighted average of features of the remaining frames.

4 . The method of claim 1 , wherein the meta-action bank includes information for a plurality of different action categories to establish an action space.

5 . The method of claim 1 , wherein the label identifies an action within the input video frame set.

6 . The method of claim 1 , wherein performing the action includes an action selected from the group consisting of generating additional information relating to the label and performing a security action relating to the label.

7 . A system for video processing, comprising:

a hardware processor; and

a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:

enrich an input video feature from an input video frame set using a meta-action bank of a plurality of video sub-actions to generate enriched features, wherein enrichment of the video frame feature is performed as:

f

˜

*

j

s

/

q

=

(

1

-

ρ

)

f

*

j

s

/

q

+

ρ

l

=

1

κ

α

j

l

d

l

e

where p is a parameter balancing between accepting external knowledge with keeping original representations, a is a cosine similarity between the video frame feature f *j s/q , d l e is a projection coefficient on a basis within the meta-action bank e, κ is a number of bases in the meta-action bank, and s/q indicates that the term is treated the same for support frames s and query frames q;

perform reinforced image representation using reinforcement learning to compare support image frames and query image frames and determine an importance of the input video frame;

perform a classification on the input video frame based on the importance and the enriched features to generate a label; and

perform an action responsive to the generated label.

8 . The system of claim 7 , wherein the reinforcement learning uses a policy based on a feature of the input video frame, an aggregated feature of remaining frames of the input video frame set, and a separate video.

9 . The system of claim 8 , wherein the aggregated feature is a weighted average of features of the remaining frames.

10 . The system of claim 7 , wherein the meta-action bank includes information for a plurality of different action categories to establish an action space.

11 . The system of claim 7 , wherein the label identifies an action within the input video frame set.

12 . The system of claim 7 , wherein the computer program further causes the hardware processor to perform the action as an action selected from the group consisting of generating additional information relating to the label and performing a security action relating to the label.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2026
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 073431/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2023
From: LI, KAI; MIN, RENQIANG; XIA, HAIFENG
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 064520/0855 →
Continuity (3)
Provisional Application 63424159 · Nov 10, 2022
Provisional Application 63397460 · Aug 12, 2022
Related Publication 20240054782A1 · Feb 15, 2024
References Cited (9)
US 9922411B2 · Aydin · 2018 [cited by examiner]
US 10885341B2 · Chen · 2021 [cited by examiner]
US 12217137B1 · Fakoor · 2025 [cited by examiner]
US 20160335224A1 · Wohlberg · 2016 [cited by examiner]
US 20170308754A1 · Torabi · 2017 [cited by examiner]
US 20190258953A1 · Lang · 2019 [cited by examiner]
US 20210124987A1 · Gan · 2021 [cited by examiner]
US 20230117307A1 · Kim · 2023 [cited by examiner]
Cao et al., “Few-Shot Video Classification via Temporal Alignment”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2020, Jun. 2020, pp. 10618-10627. [cited by applicant]