IP Library › Granted Patent US 12,579,843
Granted Patent B2
US 12,579,843 · App. 18/014,722 · Granted Mar 17, 2026

Method and system of image processing for action classification

Inventors: Anbang Yao (Beijing, CN); Shandong Wang (Beijing, CN); Ming Lu (Beijing, CN); Yuqing Hou (Beijing, CN); Yangyuxuan Kang (Beijing, CN); Yurong Chen (Beijing, CN)
Assignee: Intel Corporation
G06V40/23G06T7/20G06V10/44G06V10/82G06T2207/20044G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,843
App. No.
18/014,722
Granted
Mar 17, 2026
Kind
B2
Abstract

A method and system of image processing for action classification uses fine-grained motion-attributes.

Claims (44)

1 . At least one non-transitory machine-readable medium comprising a plurality of instructions to cause at least one programmable circuit to at least:

determine scores based on image data of frames of a video sequence;

determine action attribute values as an output of a neural network based on input data associated with a lower resolution version of the image data, wherein respective different combinations of the action attribute values represent one of corresponding different actions associated with the image data or non-action associated with the image data; and

determine whether any of the different actions are associated with the image data based on the scores and the action attribute values.

2 . The medium of claim 1 , wherein the actions correspond to motions associated with a sport.

3 . The medium of claim 1 , wherein one or more of the action attribute values are a measure of a speed that a body part moves between at least two of the frames.

4 . The medium of claim 1 , wherein one or more of the action attribute values are one or more positions of points on at least one person relative to a position of at least one other point on the at least one person.

5 . The medium of claim 1 , wherein one or more of the action attribute values represent an angle between body parts of at least one person or an angle between a body part and a reference line.

6 . The medium of claim 1 , wherein the neural network includes a prediction neural network that outputs likelihoods of multiple actions.

7 . The medium of claim 6 , wherein the instructions are to cause one or more of the at least one programmable circuit to determine whether to use a first one of the frames for action recognition based on the likelihoods.

8 . A computer-implemented system comprising:

at least one memory to store image data of frames of a video sequence;

instructions; and

at least one processor circuit to be programmed based on the instructions to:

determine initial scores based on first feature vectors, the first feature vectors based on the image data;

determine action attribute values as an output of a neural network based on input data, the input data including second features vectors based on a lower resolution version of the image data, wherein respective different combinations of the action attribute values represent one of corresponding different actions associated with the image data or non-action associated with the image data; and

determine whether any of the actions are associated with the image data based on the initial scores and the action attribute values.

9 . The system of claim 8 , wherein one or more of the at least one processor circuit is to use the action attribute values to determine which frames to use for action recognition.

10 . The system of claim 9 , wherein one or more of the at least one processor circuit is to determine whether any of the actions are associated with the image data based on a second neural network, the second neural network to process the first feature vectors, the initial scores, and the action attribute values as input to generate likelihoods of which one of the actions are associated with the image data.

11 . The system of claim 8 , wherein one or more of the at least one processor circuit is to:

compare likelihoods of respective ones of the actions output from a second neural network for a first one of the frames to at least one criterion; and

use the first one of the frames for action-related processing based on at least one of the likelihoods passing the at least one criterion.

12 . The system of claim 11 , wherein one or more of the at least one processor circuit is to discontinue use of a second frame associated with no likelihoods that meet the at least one criterion.

13 . At least one non-transitory machine-readable medium comprising a plurality of instructions to cause at least one programmable circuit to at least:

determine action prediction initial scores of individual frames of a video sequence based on feature vectors associated with respective ones of the individual frames;

determine action attribute values as an output of a first neural network based on input data associated with image data of the individual frames, wherein respective different combinations of the action attribute values represent one of corresponding different actions associated with the image data or non-action associated with the image data;

input ones of the action attribute values, ones of the action prediction initial scores and ones of the feature vectors for a first one of the frames to a second neural network to obtain probabilities of individual actions associated with the first one of the frames;

determine whether the first one of the frames is a relevant frame that indicates an action occurrence based on at least one of the probabilities meeting at least one criterion;

discard at least some of the ones of the action prediction initial scores based on the first one of the frames not being determined to be the relevant frame; and

output at least some of the ones of the action prediction initial scores based on the first one of the frames being determined to be the relevant frame.

14 . The medium of claim 13 , wherein the second neural network is a long short-term memory (LSTM) neural network including a first layer that is a fully connected layer, a second layer that is a rectifier linear unit (ReLU) layer, a third layer that is a fully connected layer, and a fourth layer that uses a sigmoid activation function to form the probabilities.

15 . The medium of claim 13 , wherein the feature vectors are first feature vectors, the input data is based on second feature vectors, and the instructions are to cause one or more of the at least one programmable circuit to generate the second feature vectors based on a lower resolution version of the image data.

16 . The medium of claim 13 , wherein the instructions are to cause one or more of the at least one programmable circuit to output the probabilities of the first one of the frames based on the first one of the frames being determined to be the relevant frame.

17 . The medium of claim 13 , wherein the instructions are to cause one or more of the at least one programmable circuit to output a highest probability among the probabilities of the first one of the frames based on the first one of the frames being determined to be the relevant frame.

18 . A method comprising:

determining action prediction initial scores of individual frames of a video sequence based on image data of frames of the video sequence, the image data of a first resolution;

determining action attribute values as an output of a neural network based on input data associated with a version of the image data having a second resolution different than the first resolution, wherein respective different combinations of the action attribute values represent one of corresponding different actions associated with the image data or non-action associated with the image data; and

combining ones of the action prediction initial scores of a same action and of multiple frames forming a segment of the video sequence to determine a single representative action score to output for further action-related processing.

19 . The method of claim 18 , including:

inputting the action attribute values of a single frame into a second neural network that outputs probabilities corresponding respectively to different actions; and

designating one of the frames to be a relevant frame based on at least one probability of the one of the frames meeting at least one criterion.

20 . The method of claim 19 , including combining ones of the probabilities of the same action and of the multiple frames forming the segment of the video sequence to determine a single representative probability to output for the further action-related processing.

21 . The method of claim 18 , wherein one or more of the attributes is a measure of a speed that a body part moves between at least two of the frames.

22 . The method of claim 18 , wherein one or more of the action attribute values represent an angle between body parts of at least one person or an angle between a body part and a reference line.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2025
From: YAO, ANBANG; WANG, SHANDONG; LU, MING; HOU, YUQING; KANG, YANGYUXUAN; CHEN, YURONG
To: INTEL CORPORATION
Reel/Frame 073154/0449 →
Continuity (1)
Related Publication 20230274580A1 · Aug 31, 2023
References Cited (28)
US 20200202119A1 · Wang et al. · 2020 [cited by applicant]
US 20200237266A1 · Qiao et al. · 2020 [cited by applicant]
CN 107122798 · 2017 [cited by applicant]
CN 110348364 · 2019 [cited by applicant]
WO WO2017155126A1 · 2017 [cited by examiner]
WO 2019072243 · 2019 [cited by applicant]
WO WO2021178692A1 · 2021 [cited by examiner]
WO WO2022026886A1 · 2022 [cited by examiner]
International Search Report and Written Opinion for PCT Application No. PCT/CN2020/109253, dated May 11, 2021. [cited by applicant]
Feichtenhofer, C., et al., “Convolutional Two-Stream Network Fusion for Video Action Recognition”, arXiv:1604.06573v2; (CVPR 2016). [cited by applicant]
Feichtenhofer, C., et al., “SlowFast Networks for Video Recognition”, arXiv:1812.03982v3; (ICCV2019). [cited by applicant]
Hochreiter, S., et al., “Long Short-term Memory”, Neural Computation 9(8), 1735-80; Dec. 1997. [cited by applicant]
Kay, W., et al., “The kinetics human action video dataset”, arXiv:1705.06950v1; 2017. [cited by applicant]
Kuehne, H., et al., “Hmdb: a large video database for human motion recognition”, ICCV (2011). [cited by applicant]
Li, D., et al., “HBONet: Harmonious Bottleneck on Two Orthogonal Dimensions”, arXiv:1908.03888v1; ICCV 2019. [cited by applicant]
Liu, S., et al., “FSD-10: A Dataset for Competitive Sports Content Analysis Analysis”, arXiv:2002.00312v1; IEEE 2020. [cited by applicant]
Parmar, P., et al., “Action quality assessment across multiple actions”, arXiv:1812.06367v2; WACV 2019. [cited by applicant]
Pirsiavash, H., et al., “Assessing the quality of actions”, ECCV 2014. [cited by applicant]
Qiu, Z., et al., “Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks”, arXiv:1711.10305v1; (ICCV 2017). [cited by applicant]
Simonyan, K., et al., “Two-Stream Convolutional Networks for Action Recognition in Videos”, arXiv:1406.2199v2; (NIPS 2014). [cited by applicant]
Soomro, K., et al., “Ucf101: A dataset of 101 human actions classes from videos in the wild”, arXiv:1212.0402v1; 2012. [cited by applicant]
Tran, D., et al., “Learning Spatiotemporal Features with 3D Convolutional Networks”, arXiv:1412.0767v4, Oct. 7, 2015; (ICCV 2015). [cited by applicant]
Wang, L., et al., “Temporal Segment Networks: Towards Good Practices for Deep Action Recognition”, arXiv:1608.00859v1; (ECCV 2018). [cited by applicant]
Wang, X., et al., “Non-local neural networks”, arXiv:1711.07971v3; (CVPR 2018). [cited by applicant]
Williams, R., et al., “Simple statistical gradient-following algorithms for connectionist reinforcement learning”, Machine Learning, 8 (3-4): 229-256, 1992. [cited by applicant]
Yan, S., et al., “Spatial temporal graph convolutional networks for skeleton-based action recognition”, arXiv:1801.07455v2; (AAAI 2018). [cited by applicant]
Zolfaghari, M. , et al., “ECO: efficient convolutional network for online video understanding”, arXiv:1804.09066v2; (ECCV 2018). [cited by applicant]
International Preliminary Report on Patentability for PCT Patent Application No. PCT/CN2020/109253, dated Feb. 23, 2023. [cited by applicant]