IP Library Granted Patent US 12,198,397
Granted Patent B2
US 12,198,397 · App. 17/586,284 · Granted Jan 14, 2025

Keypoint based action localization

Inventors: Asim Kadav (Mountain View, CA); Farley Lai (Santa Clara, CA); Hans Peter Graf (South Amboy, NJ); Yi Huang (San Diego, CA)
Assignee: NEC Corporation
G06V10/26G06T7/251G06V10/44G06V10/82G06V20/58G06V40/10G08G1/166G06T2207/30261
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,397
App. No.
17/586,284
Granted
Jan 14, 2025
Kind
B2
Abstract

A computer-implemented method is provided for action localization. The method includes converting one or more video frames into person keypoints and object keypoints. The method further includes embedding position, timestamp, instance, and type information with the person keypoints and object keypoints to obtain keypoint embeddings. The method also includes predicting, by a hierarchical transformer encoder using the keypoint embeddings, human actions and bounding box information of when and where the human actions occur in the one or more video frames.

Claims (31)

1. A computer-implemented method for action localization, comprising:

converting one or more video frames into person keypoints and object keypoints;

embedding position, timestamp, instance, and type information with the person keypoints and object keypoints to obtain keypoint embeddings; and

predicting, by a hierarchical transformer encoder using the keypoint embeddings, human actions and bounding box information of when and where the human actions occur in the one or more video frames, the embedding including converting the keypoints to tokens, and the predicting including projecting the tokens to embedding metrics and summing the embedding metrics to obtain an output keypoint embedding.

2. The computer-implemented method of claim 1 , wherein said converting converts the one or more video frames into person keypoints in a form of human joint names for each detected person.

3. The computer-implemented method of claim 2 , wherein said converting further comprises selecting a top N out of detected persons based on person detection confidence scores.

4. The computer-implemented method of claim 1 , wherein said converting comprises extracting the object keypoints by subsampling a contour of an object mask detected by a Mask R-CNN.

5. The computer-implemented method of claim 4 , wherein said converting further comprises selecting top N out of detected objects based on object detection confidence scores.

6. The computer-implemented method of claim 1 , further comprising learning atomic actions from the person keypoints and the object keypoints.

7. The computer-implemented method of claim 1 , wherein the position information comprises a down-sampled spatial location of each pixel coordinate.

8. The computer-implemented method of claim 1 , wherein the timestamp information comprises a difference between a keypoint timestamp and a beginning keyframe timestamp.

9. The computer-implemented method of claim 1 , wherein the instance information comprises a spatial correlation between the person keypoints and a person instance.

10. The computer-implemented method of claim 1 , wherein the type information comprises a human body part name.

11. The computer-implemented method of claim 1 , wherein the position, timestamp, instance, and type information comprise representative tokens that are linearly projected to a respective embedding metric and summed to obtain an output keypoint embedding through a transformer based Keypoint Embedding Network.

12. The computer-implemented method of claim 1 , further comprising controlling a vehicle system for accident avoidance responsive to the predicted human actions and the bounding box information.

13. The computer-implemented method of claim 1 , further comprising controlling a robotic system for collision avoidance responsive to the predicted human actions and the bounding box information.

14. A computer program product for action localization, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

converting, by a processor device of the computer, one or more video frames into person keypoints and object keypoints;

embedding, by the processor device, position, timestamp, instance, and type information with the person keypoints and object keypoints to obtain keypoint embeddings; and

predicting, by a hierarchical transformer encoder of the computer using the embedded keypoints, human actions and bounding box information of when and where the human actions occur in the one or more video frames, the embedding including converting the keypoints to tokens, and the predicting including projecting the tokens to embedding metrics and summing the embedding metrics to obtain an output keypoint embedding.

15. The computer program product of claim 14 , wherein said converting converts the one or more video frames into person keypoints in a form of human joint names for each detected persons.

16. The computer program product of claim 15 , wherein said converting further comprises selecting a top N out of the detected persons based on person detection confidence scores.

17. The computer program product of claim 14 , wherein said converting comprises extracting the object keypoints by subsampling a contour of a mask detected by a Mask R-CNN.

18. The computer program product of claim 17 , wherein said converting further comprises selecting top N out of detected objects based on object detection confidence scores.

19. The computer program product of claim 14 , further comprising learning atomic actions in the person keypoints and the object keypoints.

20. A computer processing system for action localization, comprising:

a memory device for storing program code;

a processor device operatively coupled to the memory device for running the program code for:

converting one or more video frames into person keypoints and object keypoints;

embedding position, timestamp, instance, and type information with the person keypoints and object keypoints to obtain keypoint embeddings; and

predicting, using a hierarchical transformer encoder that inputs the keypoint embeddings, human actions and bounding box information of when and where the human actions occur in the one or more video frames, the embedding including converting the keypoints to tokens, and the predicting including projecting the tokens to embedding metrics and summing the embedding metrics to obtain an output keypoint embedding.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: KADAV, ASIM; LAI, FARLEY; GRAF, HANS PETER; HUANG, YI
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 058796/0968 →
Continuity (2)
Provisional Application 63142602 · Jan 28, 2021
Related Publication 20220237884A1 · Jul 28, 2022
References Cited (36)
US 9437009B2 · Medioni · 2016 [cited by examiner]
US 20200074678A1 · Ning · 2020 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20200359064A1 · Zhou · 2020 [cited by examiner]
US 20200394413A1 · Bhanu · 2020 [cited by examiner]
US 20210076105A1 · Parmar · 2021 [cited by examiner]
Feichtenhofer, Christoph, et al. “Slowfast networks for video recognition”, InProceedings of the IEEE/CVF International conference on computer vision. Oct. 2019, pp. 6202-6211. [cited by applicant]
Cao, Zhe, et al. “OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields”, IEEE transactions on pattern analysis and machine intelligence. Jul. 17, 2019, pp. 172-186. [cited by applicant]
Carreira, Joao, et al. “Quo vadis, action recognition? a new model and the kinetics dataset” Inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Jul. 2017, pp. 6299-6308. [cited by applicant]
Devlin, Jacob, et al. “Bert: Pre-training of deep bidirectional transformers for language understanding”, InProceedings of NAACL-HLT. May 2019, pp. 4171-4186. [cited by applicant]
Du, Wenbin, et al. “Rpan: An end-to-end recurrent pose-attention network for action recognition in videos”, InProceedings of the IEEE International Conference on Computer Vision. Oct. 2017, pp. 3725-3734. [cited by applicant]
Du, Yong, et al. “Hierarchical recurrent neural network for skeleton based action recognition”, InProceedings of the IEEE conference on computer vision and pattern recognition. Jun. 2015, pp. 1110-1118. [cited by applicant]
Feichtenhofer, Christoph, et al. “Convolutional two-stream network fusion for video action recognition”, InProceedings of the IEEE conference on computer vision and pattern recognition. Jun. 2016, pp. 1933-1941. [cited by applicant]
Feng, Yutong, et al. “Relation Modeling in Spatio-Temporal Action Localization”, arXiv preprint arXiv:2106.08061. Jun. 16, 2021, pp. 1-6. [cited by applicant]
Gavrilyuk, Kirill, et al. “Actor-transformers for group activity recognition”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jun. 2020, pp. 839-848. [cited by applicant]
Girdhar, Rohit, et al. “Video action transformer network”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jun. 2019, pp. 244-253. [cited by applicant]
Gkioxari, Georgia, et al. “Detecting and recognizing human-object interactions”, InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Jun. 2018, pp. 8359-8367. [cited by applicant]
Gu, Chunhui, et al. “Ava: A video dataset of spatio-temporally localized atomic visual actions”, InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Jun. 2018, pp. 6047-6056. [cited by applicant]
He, Kaiming, et al. “Mask r-cnn”, InProceedings of the IEEE international conference on computer vision. Oct. 2017, pp. 2961-2969. [cited by applicant]
Jhuang, Hueihan, et al. “Towards understanding action recognition”, InProceedings of the IEEE international conference on computer vision. Dec. 2013, pp. 3192-3199. [cited by applicant]
Lin, Ji, et al. “Tsm: Temporal shift module for efficient video understanding”, InProceedings of the IEEE/CVF International Conference on Computer Vision. Oct. 2019, pp. 7083-7093. [cited by applicant]
Liu, Ziyu, et al. “Disentangling and unifying graph convolutions for skeleton-based action recognition”, InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. Jun. 2020, pp. 143-152. [cited by applicant]
Ma, Chih-Yao, et al. “TS-LSTM and temporal-inception: Exploiting spatiotemporal dynamics for activity recognition”, Signal Processing: Image Communication. Feb. 1, 2019, pp. 76-87. [cited by applicant]
Ma, Chih-Yao, et al. “Attend and interact: Higher-order object interactions for video understanding”, InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Jun. 2018, pp. 6790-6800. [cited by applicant]
Obinata, Yuya, et al. “Temporal Extension Module for Skeleton-Based Action Recognition”, In2020 25th International Conference on Pattern Recognition (ICPR). Jan. 10, 2021, pp. 534-540. [cited by applicant]
Pan, Junting, et al. “Actor-context-actor relation network for spatio-temporal action localization”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Jun. 2021, pp. 464-474. [cited by applicant]
Shao, Hao, et al. “Temporal interlacing network”, InProceedings of the AAAI Conference on Artificial Intelligencevol. 34, No. 07. Apr. 3, 2020, pp. 11966-11973. [cited by applicant]
Snower, Michael, et al. “15 keypoints is all you need”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jun. 2020, pp. 6738-6748. [cited by applicant]
Song, Sijie, et al. “An end-to-end spatio-temporal attention model for human action recognition from skeleton data”, InProceedings of the AAAI conference on artificial intelligence, vol. 31, No. 1. Feb. 12, 2017, pp. 42… [cited by applicant]
Wang, Fei, et al. “Can WiFi estimate person pose?”, arXiv preprint arXiv:1904.00277. Apr. 2, 2019, pp. 1-11. [cited by applicant]
Wang, Jingdong, et al. “Deep high-resolution representation learning for visual recognition”, IEEE transactions on pattern analysis and machine intelligence. Mar. 13, 2020, pp. 1-23. [cited by applicant]
Wang, Limin, et al. “Temporal segment networks: Towards good practices for deep action recognition”, InEuropean conference on computer vision, Springer, Cham. Oct. 8, 2016, pp. 20-36. [cited by applicant]
Wang, Saiwen, et al. “Interacting with soli: Exploring fine-grained dynamic gesture recognition in the radio-frequency spectrum”, InProceedings of the 29th Annual Symposium on User Interface Software and Technology. Oct… [cited by applicant]
Wang, Wei, et al. “Pose-based two-stream relational networks for action recognition in videos”, arXiv preprint arXiv:1805.08484. May 22, 2018, pp. 1-15. [cited by applicant]
Wu, Zuxuan, et al. “Multi-stream multi-class fusion of deep networks for video classification”, InProceedings of the 24th ACM international conference on Multimedia. Oct. 1, 2016, pp. 791-800. [cited by applicant]
Yan, Sijie, et al. Spatial temporal graph convolutional networks for skeleton-based action recognition. InThirty-second AAAI conference on artificial intelligence. Apr. 27, 2018, pp. 7444-7452. [cited by applicant]
Cited By (1)
US 12,626,507