IP Library › Granted Patent US 12,511,899
Granted Patent B2
US 12,511,899 · App. 18/449,393 · Granted Dec 30, 2025

Cross-domain few-shot video classification with optical-flow semantics

Inventors: Kai Li (Plainsboro, NJ); Renqiang Min (Princeton, NJ); Haifeng Xia (New Orleans, LA)
Assignee: NEC Corporation
G06V20/41G06T7/246G06V10/774G06V10/776G06V10/82G06V20/46G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,899
App. No.
18/449,393
Granted
Dec 30, 2025
Kind
B2
Abstract

Methods and systems for video processing include extracting flow features and appearance features from frames of a video stream. The flow features are processed using a flow model that is trained on a first set of training data. An output of the flow model is processed using a sub-network that is trained on the first set of training data and a second set of domain-specific training data to generate a flow parameter. The appearance features are processed using an appearance model that is trained on the first set of training data and that further processes the appearance features using the flow parameter, to classify the frames of the video stream. An action is performed responsive to the classified frames.

Claims (34)

1 . A computer-implemented method for video processing, comprising:

extracting flow features and appearance features from frames of a video stream;

processing the flow features using a flow model that is trained on a first set of training data;

processing an output of the flow model using a sub-network that is trained on the first set of training data and a second set of domain-specific training data to generate a flow parameter;

processing the appearance features using an appearance model that is trained on the first set of training data and that further processes the appearance features using the flow parameter, to classify the frames of the video stream; and

performing an action responsive to the classified frames.

2 . The method of claim 1 , wherein processing the appearance features includes a convolutional network that processes the appearance features in parallel with the flow parameter.

3 . The method of claim 2 , wherein processing the appearance features includes adding a product of flow parameter and the appearance features to an output of the convolutional network.

4 . The method of claim 1 , wherein the appearance model includes a series of blocks, each of which includes a convolutional network and a respective instance of the flow parameter.

5 . The method of claim 1 , wherein the flow features characterize motion information relating to the frames of the video stream.

6 . The method of claim 1 , wherein the appearance features characterize pixel value information from the frames of the video stream.

7 . The method of claim 1 , wherein performing the action includes an action selected from the group consisting of generating additional information relating to the classified frames and performing a security action relating to the classified frames.

8 . A system for video processing, comprising:

a hardware processor; and

a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:

extract flow features and appearance features from frames of a video stream;

process the flow features using a flow model that is trained on a first set of training data;

process an output of the flow model using a sub-network that is trained on the first set of training data and a second set of domain-specific training data to generate a flow parameter;

process the appearance features using an appearance model that is trained on the first set of training data and that further processes the appearance features using the flow parameter, to classify the frames of the video stream; and

perform an action responsive to the classified frames.

9 . The system of claim 8 , wherein the computer program further causes the hardware processor to use a convolutional network that processes the appearance features in parallel with the flow parameter.

10 . The system of claim 9 , wherein the computer program further causes the hardware processor to add a product of flow parameter and the appearance features to an output of the convolutional network.

11 . The system of claim 8 , wherein the appearance model includes a series of blocks, each of which includes a convolutional network and a respective instance of the flow parameter.

12 . The system of claim 8 , wherein the flow features characterize motion information relating to the frames of the video stream.

13 . The system of claim 8 , wherein the appearance features characterize pixel value information from the frames of the video stream.

14 . The system of claim 8 , wherein the computer program further causes the hardware processor to perform action selected from the group consisting of generating additional information relating to the classified frames and performing a security action relating to the classified frames.

15 . A computer-implemented method for training a neural network model, comprising:

training a flow model to process flow features of frames of a video stream, a sub-network to process outputs of the flow model, and an appearance model to process appearance features of the frames with a flow parameter from the sub-network, using a shared first set of training data; and

tuning the sub-network while the flow model and the appearance model are held fixed, using a second set of domain-specific training data.

16 . The method of claim 15 , wherein the appearance model includes a convolutional network that processes the appearance features in parallel with the flow parameter.

17 . The method of claim 16 , wherein the appearance model adds a product of flow parameter and the appearance features to an output of the convolutional network.

18 . The method of claim 15 , wherein the appearance model includes a series of blocks, each of which includes a convolutional network and a respective instance of the flow parameter.

19 . The method of claim 15 , further comprising extracting flow features from the frames to characterize motion information relating to the frames.

20 . The method of claim 15 , further comprising extracting appearance features from the frames to characterize pixel value information from the frames.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 072938/0913 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2023
From: LI, KAI; MIN, RENQIANG; XIA, HAIFENG
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 064590/0925 →
Continuity (2)
Provisional Application 63397959 · Aug 15, 2022
Related Publication 20240054783A1 · Feb 15, 2024
References Cited (5)
US 11074454B1 · Vijayanarasimhan · 2021 [cited by examiner]
US 20190333198A1 · Wang · 2019 [cited by examiner]
US 20210232825A1 · Tang · 2021 [cited by examiner]
Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition, Sun et al., 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (Year: 2018). [cited by examiner]
Liu et al., “A Multi-Mode Modulator for Multi-Domain Few-Shot Classification”, In Proceedings of the IEEE/CVF International Conference on Computer Vision 2021, Oct. 2021, pp. 8453-8462. [cited by applicant]