IP Library › Granted Patent US 12,562,005
Granted Patent B2
US 12,562,005 · App. 18/561,524 · Granted Feb 24, 2026

Neural network system for 3D pose estimation

Inventors: Tao Wang (Singapore, SG); Jianfeng Zhang (Singapore, SG); Shuicheng Yan (Singapore, SG); Jiashi Feng (Singapore, SG)
Assignee: Shopee IP Singapore Private Limited
G06V40/20G06T7/55G06V10/82G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,005
App. No.
18/561,524
Granted
Feb 24, 2026
Kind
B2
Abstract

A system for 3D pose estimation of people is described, comprising an encoder neural network configured to extract multi-view image features from 2D input images obtained from different camera views, and a decoder neural network configured to receive the multi-view image features and predict 3D joint locations of the people in the 2D input images. The decoder neural network comprising a projective-attention mechanism configured to: determine a 2D projection of a predicted 3D joint location for the camera views and assign them as anchor points; apply an adaptive deformable sampling strategy to gather localized context information of the camera views and to learn deformable offsets, and based on the deformable offsets, determine the deformable points for the anchor points; generate attention weights based on the multi-view image features at the anchor points; and apply the attention weights to aggregate the multi-view image features at the deformable points.

Claims (18)

1 . A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement a neural network system for predicting 3D poses of a plurality of persons in a scene based on 2D input images, each of the 2D input images obtained from a different camera view, the different camera views capturing a visual of the plurality of persons in a scene from different perspectives;

the neural network system comprising an encoder neural network and a decoder neural network, the encoder neural network configured to extract multi-view image features from the 2D input images, the decoder neural network configured to receive the multi-view image features as input, and predict 3D joint locations of each of the plurality of persons;

wherein the decoder neural network comprises decoder layers to regressively refine the predicted 3D joint locations, and wherein each of the decoder layers comprises a projective-attention mechanism, the projective-attention mechanism configured to:

determine a 2D projection of a predicted 3D joint location for each of the different camera views, wherein the predicted 3D joint location is obtained from a preceding decoder layer or is determined by linear methods based on input joint queries;

assign the 2D projection as an anchor point for each of the different camera views;

apply an adaptive deformable sampling strategy to gather localized context information of the different camera views to learn deformable offsets, and based on the deformable offsets, determine the deformable points for each of the anchor points;

generate attention weights based on the multi-view image features at the anchor points; and

apply the attention weights to aggregate the multi-view image features at the deformable points.

2 . The system of claim 1 wherein the input joint queries are joint query embeddings, the joint query embeddings comprising hierarchical embedding of joint queries and person queries that encode a person-joint relation; and wherein each of the decoder layers further comprises a self-attention mechanism configured to perform self-attention on the joint query embeddings.

3 . The system of claim 2 wherein the joint query embeddings are augmented with scene level information that are specific to each of the 2D input images.

4 . The system of claim 1 further comprising a positional encoder, the positional encoder configured to encode the camera ray directions for each of the different camera views into the multi-view image features.

5 . The system of claim 4 wherein the camera ray directions are generated with camera parameters of the different camera views; and wherein encoding the camera ray directions for each of the different camera views into the multi-view image features comprises concatenating channel-wisely the camera ray directions to the corresponding image features, and applying a standard convolution to obtain updated image representations.

6 . The system of claim 5 wherein the projective-attention mechanism is further configured to: receive the updated image representations as input;

generate attention weights based on the updated image representations at the anchor points; and

apply the attention weights to aggregate the updated image representations at the deformable points.

7 . The system of claim 1 wherein each of the decoder layers further comprises a feed forward network block, the feed forward network block configured to apply feed-forward regression to predict the 3D joint locations, and confidence scores of the predicted 3D joint locations.

8 . The system of claim 1 wherein the decoder neural network comprises at least four decoder layers.

9 . The system of claim 1 wherein the encoder neural network comprises a convolutional neural network or a transformer based component.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2025
From: GARENA ONLINE PRIVATE LIMITED
To: SHOPEE IP SINGAPORE PRIVATE LIMITED
Reel/Frame 071546/0281 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: ZHANG, JIANFENG; WANG, TAO; YAN, SHUICHENG; FENG, JIASHI
To: GARENA ONLINE PRIVATE LIMITED
Reel/Frame 065589/0102 →
Priority Claims (3)
SG 10202105170T · May 18, 2021 · national
SG 10202105716X · May 29, 2021 · national
SG 10202108302R · Jul 29, 2021 · national
Continuity (1)
Related Publication 20240233441A1 · Jul 11, 2024
References Cited (64)
US 9053562B1 · Rabin · 2015 [cited by examiner]
US 9418480B2 · Issa · 2016 [cited by examiner]
US 9904867B2 · Fathi · 2018 [cited by examiner]
US 11357581B2 · Mozes · 2022 [cited by examiner]
US 20140081659A1 · Nawana · 2014 [cited by examiner]
US 20190094542A1 · Langner · 2019 [cited by examiner]
US 20190347817A1 · Ferrantelli · 2019 [cited by examiner]
US 20200066029A1 · Chen · 2020 [cited by examiner]
US 20200084353A1 · Wacey · 2020 [cited by examiner]
US 20200234466A1 · Holzer · 2020 [cited by applicant]
US 20210295548A1 · Veiga · 2021 [cited by examiner]
US 20210390767A1 · Johnson · 2021 [cited by examiner]
US 20220274114A1 · Vazirani · 2022 [cited by examiner]
CN 109977827A · 2019 [cited by applicant]
Zhu et al , Deformable DETR: Deformable Transformers for End-to-End Object Detection, Mar. 18, 2021 (Year: 2021). [cited by examiner]
Vasileios Belagiannis, CVF, 3D Pictorial Structures for Multiple Human Pose Estimation, Computer Aided Medical Procedures, Technische Universität München, Germany, CVPR, 2014, pp. 1-8. [cited by applicant]
Richard Hartley and Andrew Zisserman, Cambridge, Multiple View Geometry in computer vision, Second Edition, Cambridge University Press< New York 2000, 2003, pp. 1-673. [cited by applicant]
Vasileios Belagiannis, 3D Pictorial Structures Revisited: Multiple Human Pose Estimation, pp. 1-14, 1929-1942, 2015. [cited by applicant]
Nicolas Carion, End-to-End Object Detection with Transformers, May 28, 2020, pp. 1-26. [cited by applicant]
Jifeng Dai, Deformable Convolutional Networks, CVF, Microsoft Research Asia, pp. 764-773, ICCV, 2017. [cited by applicant]
Alexey Dosovitskiy, An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale Published as a conference paper at ICLR 2021, Jun. 3, 2021, pp. 1-22. [cited by applicant]
Sara Ershadi-Nasab, Multiple human 3D pose estimation from Multiview images, CrossMark, Revised: Jul. 9, 2017 / Accepted: Aug. 20, 2017, © Springer Science+Business Media, LLC 2017, pp. 1-29. [cited by applicant]
Kehong Gong, PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation, National University of Singapore, CVF, 2021, pp. 8575-8584. [cited by applicant]
He Chen, Multi-person 3D Pose Estimation in Crowded Scenes Based on Multi-View Geometry, The Johns Hopkins University, USA, National University of Singapore, Singapore, Jul. 21, 2020, pp. 1-17. [cited by applicant]
Junting Dong, Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views, CVF, pp. 7792-7801, 2019. [cited by applicant]
Kaiming He, Deep Residual Learning for Image Recognition, CVF, pp. 770-778, 2016. [cited by applicant]
Yihui He, Epipolar Transformers, CVF, 2020, pp. 7779-7788. [cited by applicant]
Congzhentao Huang, End-to-end Dynamic Matching Network for Multi-view Multi-person 3d Pose Estimation, University of Technology Sydney, Sydney, Australia, pp. 1-16, 2020. [cited by applicant]
Karim Iskakov, Learnable Triangulation of Human Pose, Samsung AI Center, Moscow, May 14, 2019, pp. 1-9. [cited by applicant]
Wen Jiang, Coherent Reconstruction of Multiple Humans from a Single Image, University of Pennsylvania, Jun. 15, 2020, pp. 1-10. [cited by applicant]
Yifan Jiang, TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up, University of Texas at Austin, Dec. 9, 2021, pp. 1-19. [cited by applicant]
Hanbyul Joo, Panoptic Studio: A Massively Multiview System for Social Motion Capture, Yaser Sheikh The Robotics Institute, Carnegie Mellon University, pp. 1-9, 2015. [cited by applicant]
Hanbyul Joo, Panoptic Studio: A Massively Multiview System for Social Interaction Capture, https://domedb.perception.cs.cmu.edu, Dec. 9, 2016, pp. 1-14. [cited by applicant]
Abdolrahim Kadkhodamohammadi, A generalizable approach for multi-view 3D human pose regression, Oct. 8, 2019, pp. 1-15. [cited by applicant]
Angjoo Kanazawa, End-to-end Recovery of Human Shape and Pose, University of California, Berkeley, Jun. 23, 2018, pp. 1-10. [cited by applicant]
Sven Kreiss, PifPaf: Composite Fields for Human Pose Estimation, EPFL VITA lab CH-1015 Lausanne, Apr. 5, 2019, pp. 1-10. [cited by applicant]
H. W. Kuhn, the Hungarian Method for the Assignment Problem, Bryn Yaw College, pp. 83-97, 1955. [cited by applicant]
Tsung-Yi Lin, Focal Loss for Dense Object Detection, Facebook AI Research (FAIR), Feb. 7, 2018, pp. 1-10. [cited by applicant]
Matthew Loper, SMPL: A Skinned Multi-Person Linear Model, Max Planck Institute for Intelligent Systems, Tubingen, Germany, Industrial Light and Magic, San Francisco, CA, pp. 1-16, 2015. [cited by applicant]
Julieta Martinez, A simple yet effective baseline for 3d human pose estimation, 1University of British Columbia, Vancouver, Canada, Body Labs Inc., New York, NY, Aug. 4, 2017, pp. 1-10. [cited by applicant]
Dushyant Mehta, VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera, To Appear in ACM TOG (SIGGRAPH 2017), Max Planck Institute for Informatics, May 3, 2017, pp. 1-13. [cited by applicant]
Xuecheng Nie, Single-Stage Multi-Person Pose Machines, Department of Electrical and Computer Engineering, Aug. 24, 2019, pp. 1-10. [cited by applicant]
George Papandreou, PersonLab: Person Pose Estimation and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model, Google Research, Los Angeles, USA, Springer Nature Switzerland AG 2018, pp. 282-299. [cited by applicant]
Georgios Pavlakos, Harvesting Multiple Views for Marker-less 3D Human Pose Annotations, University of Pennsylvania, Apr. 16, 2017, pp. 1-10. [cited by applicant]
Alin-Ionut Popa, Deep Multitask Architecture for Integrated 2D and 3D Human Sensing, Department of Mathematics, Faculty of Engineering, Lund University, Jan. 31, 2017, pp. 1-10. [cited by applicant]
Haibo Qiu, Cross View Fusion for 3D Human Pose Estimation, University of Science and Technology of China, Sep. 3, 2019, pp. 1-10. [cited by applicant]
Edoardo Remelli, Lightweight Multi-View 3D Pose Estimation through Camera-Disentangled Representation, CVLab, EPFL, Lausanne, Switzerland, Jun. 20, 2020, pp. 1-16. [cited by applicant]
Xiao Sun, Integral Human Pose Regression, Microsoft Research, Beijing, China, Springer Nature Switzerland AG 2018, pp. 536-553. [cited by applicant]
Ilya Sutskever, Sequence to Sequence Learning with Neural Networks, Dec. 12, 2014, pp. 1-9. [cited by applicant]
Hanyue Tu, VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment, University of Science and Technology of China, Aug. 24, 2020, pp. 1-17. [cited by applicant]
Ashish Vaswani, Attention Is All You Need, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, pp. 1-15. [cited by applicant]
Yuqing Wang, End-to-End Video Instance Segmentation with Transformers, CVF, The University of Adelaide, Australia, 2021, pp. 8741-8750. [cited by applicant]
Felix Wu, Pay Less Attention With Lightweight and Dynamic Convolutions, Cornell University, Feb. 22, 2019, pp. 1-14. [cited by applicant]
Zhanghao Wu, Lite Transformer With Long-Short Range Attention, Massachusetts Institute of Technology, Apr. 24, 2020, pp. 1-13. [cited by applicant]
Bin Xiao, Simple Baselines for Human Pose Estimation and Tracking, Microsoft Research Asia, Beijing, China, Springer Nature Switzerland AG 2018, pp. 472-487. [cited by applicant]
Jianfeng Zhang, Inference Stage Optimization for Cross-scenario 3D Human Pose Estimation, National University of Singapore, Jul. 4, 2020, pp. 1-12. [cited by applicant]
Jianfeng Zhang, Body Meshes as Points, National University of Singapore, CVF, 2021, pp. 546-556. [cited by applicant]
Hengshuang Zhao, Exploring Self-attention for Image Recognition, CVF, 2021, pp. 10076-10085. [cited by applicant]
Xingyi Zhou, Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach, CVF, Shanghai Key Laboratory of Intelligent Information Processing School of Computer Science, Fudan University, pp. 398-407, 2017. [cited by applicant]
Xizhou Zhu, Deformable ConvNets v2: More Deformable, Better Results, University of Science and Technology of China, CVF, pp. 9308-9316, 2019. [cited by applicant]
Diederik P. Kingma, ADAM: a Method for Stochastic Optimization, University of Amsterdam, OpenAI, Jan. 30, 2017, pp. 1-15. [cited by applicant]
Adam Paszke, Automatic differentiation in PyTorch, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, pp. 1-4. [cited by applicant]
Fu Xiong, Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image, Aug. 27, 2019. [cited by applicant]
Wu Liu, Recent Advances in Monocular 2D and 3D Human Pose Estimation: A Deep Learning Perspective, Apr. 23, 2021. [cited by applicant]