IP Library › Granted Patent US 12,743,891
Granted Patent B2
US 12,743,891 · App. 17/225,924 · Granted Sep 22, 2026

End-to-end action recognition in intelligent video analysis and edge computing systems

Inventors: Subhashree Radhakrishnan (San Jose, CA); Farzin Aghdasi (East Palo Alto, CA)
Assignee: NVIDIA Corporation
G06V20/58G06F18/2411G06N3/045G06N3/08G06T1/20G06V40/25G06T2207/20132G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,891
App. No.
17/225,924
Granted
Sep 22, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to perform action recognition. In at least one embodiment, action recognition is performed using one or more neural networks and hardware accelerators, in which the one or more neural networks are processed based on, for example, one or more quantization and pruning processes.

Claims (48)

1 . A method, comprising:

obtaining a plurality of frames of a video;

determining one or more objects represented in the plurality of frames;

causing a parallel processing unit to calculate a flow field representing movements of the one or more objects represented in one or more pixels among frames of the plurality of frames;

obtaining one or more cropped movements by performing a first set of cropping operations on the flow field based at least in part on a set of bounding boxes corresponding to the objects;

obtaining one or more cropped frames by performing a second set of cropping operations on the plurality of frames based at least in part on the set of bounding boxes; and

causing a neural network to classify one or more actions performed by the one or more objects represented in the plurality of frames based at least in part on the one or more cropped movements and the one or more cropped frames.

2 . The method of claim 1 , further comprising:

generating the set of bounding boxes corresponding to the one or more objects represented in the plurality of frames.

3 . The method of claim 2 , further comprising:

causing a second neural network to generate the set of bounding boxes corresponding to the one or more objects represented in the plurality of frames.

4 . The method of claim 1 , further comprising:

determining one or more values based at least in part on one or more kernels of the neural network; and

removing a set of kernels from the neural network based at least in part on the one or more values.

5 . The method of claim 1 , wherein the one or more actions include at least a sit action, a walk action, a run action, or a climb stairs action.

6 . The method of claim 1 , wherein the neural network comprises one or more quantized weights.

7 . A processor, comprising:

one or more circuits to:

identify one or more objects depicted in one or more frames from video data;

calculate, using the one or more frames of the video data, one or more flow fields and one or more bounding boxes corresponding to the one or more objects;

obtain one or more cropped movements by performing a first set of cropping operations on the one or more flow fields based, at least in part, on a set of bounding boxes corresponding to the objects;

obtain one or more cropped frames by performing a second set of cropping operations on the one or more frames based at least in part on the set of bounding boxes; and

cause one or more neural networks to classify one or more actions represented in the one or more frames based, at least in part, on the one or more cropped movements and the one or more cropped frames.

8 . The processor of claim 7 , wherein the one or more circuits are further to:

use a first neural network to calculate the set of bounding boxes; and

use a second neural network to classify the one or more actions.

9 . The processor of claim 8 , wherein the one or more circuits are further to:

calculate one or more Ll-norm values of one or more kernels of the first neural network and the second neural network; and

remove a set of kernels from the first neural network and the second neural network based at least in part on the one or more Ll-norm values.

10 . The processor of claim 9 , wherein the set of kernels are associated with a set of Ll-norm values above a threshold.

11 . The processor of claim 7 , wherein the one or more circuits further cause the one or more flow fields to be calculated using one or more hardware accelerators.

12 . The processor of claim 8 , wherein the first neural network and the second neural network comprise a set of quantized weights.

13 . The processor of claim 7 , wherein the one or more circuits are to obtain video data captured by one or more autonomous vehicle systems, wherein the video data is comprised of the one or more frames.

14 . A computing device, comprising:

a hardware accelerator; and

memory comprising instructions executable by one or more processors of the computing device to at least:

obtain a plurality of frames;

cause the hardware accelerator to calculate a flow field representing movements of one or more objects represented by one or more pixels among frames of the plurality of frames;

use a first neural network to determine the one or more objects represented in the plurality of frames;

obtain one or more cropped movements by performing a first set of cropping operations on the flow field based at least in part on a set of bounding boxes corresponding to the objects;

obtain one or more cropped frames by performing a second set of cropping operations on the plurality of frames based at least in part on the set of bounding boxes; and

cause a second neural network to use the one or more cropped movements and the one or more cropped frames to classify one or more actions.

15 . The computing device of claim 14 , wherein the first neural network and the second neural network comprise one or more weights processed through one or more quantization processes.

16 . The computing device of claim 15 , wherein the one or more quantization processes include a conversion of the one or more weights from a first representation to a second representation.

17 . The computing device of claim 15 , wherein the one or more quantization processes are based at least in part on one or more Kullback-Leibler divergence values.

18 . The computing device of claim 16 , wherein the first representation is a 32-bit floating point representation and the second representation is an 8-bit integer representation.

19 . The computing device of claim 14 , wherein the instructions further include instructions executable by the one or more processors to at least obtain the plurality of frames from one or more image capturing hardware of the computing device.

20 . The computing device of claim 14 , wherein the hardware accelerator comprises one or more parallel processing units.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2021
From: RADHAKRISHNAN, SUBHASHREE; AGHDASI, FARZIN
To: NVIDIA CORPORATION
Reel/Frame 055870/0541 →
Continuity (1)
Related Publication 20220327318A1 · Oct 13, 2022
References Cited (22)
US 11282367B1 · Aquino · 2022 [cited by examiner]
US 20140201126A1 · Zadeh · 2014 [cited by examiner]
US 20160174902A1 · Georgescu et al. · 2016 [cited by applicant]
US 20200175326A1 · Shen et al. · 2020 [cited by applicant]
US 20200361083A1 · Mousavian et al. · 2020 [cited by applicant]
US 20210255860A1 · Morrison · 2021 [cited by examiner]
CN 110998594A · 2020 [cited by applicant]
CN 111950693A · 2020 [cited by applicant]
JP 2020042646A · 2020 [cited by applicant]
International Search Report And Written Opinion for Application No. PCT/US2022/022924, mailed Jul. 7, 2022, filed Mar. 31, 2022, 15 pages. [cited by applicant]
Xu et al., “Can Humans Fly? Action Understanding with Multiple Classes of Actors,” CVPR, Jun. 7, 2015, 10 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Lin et al., “TSM: Temporal Shift Module for Efficient Video Understanding,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 11 pages. [cited by applicant]
Molchanov et al., Pruning Convolutional Neural Networks For Resource Efficient Inference, ICLR, 2017, 12 pages. [cited by applicant]
NVIDIA, “DeepStream SDK,” retrieved from https://developer.nvidia.com/deepstream-sdk, 2019, 17 pages. [cited by applicant]
NVIDIA, “NVIDIA Optical Flow SDK,” retrieved from https://developer.nvidia.com/opticalflow-sdk, 2020, 6 pages. [cited by applicant]
Simonyan et al., “Two-Stream Convolutional Networks for Action Recognition in Videos,” Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Office Action for Chinese Application No. 202280003851.4, mailed Dec. 26, 2025, 18 pages. [cited by applicant]
Office Action for Chinese Application No. 202280003851.4, mailed Jul. 7, 2026, 20 pages. [cited by applicant]
Machine translation for CN 110998594A. [cited by applicant]