IP Library › Granted Patent US 12,614,070
Granted Patent B2
US 12,614,070 · App. 17/764,093 · Granted Apr 28, 2026

Object re-identification using pose part based models

Inventors: Jianguo Li (Beijing, CN); Shuyuan Li (Shanghai, CN); Hanlin Tang (San Francisco, CA)
Assignee: Intel Corporation
G06N3/08G06T7/74G06V10/454G06V10/7715G06V10/776G06V10/806G06V10/82G06V20/53G06T2207/20081G06T2207/20084G06T2207/30196G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,070
App. No.
17/764,093
Granted
Apr 28, 2026
Kind
B2
Abstract

An example apparatus for re-identifying objects includes an image receiver to receive a first image and a second image of an object with an identity. The apparatus also includes a fused model generator to fuse a global representation of the object with local representations of pose parts of the object to generate a fused representation of the object based on the first image. The apparatus further includes an object re-identifier to re-identify the object with the identity in the second image based on the fused representation.

Claims (47)

1 . An apparatus comprising:

interface circuitry;

machine-readable instructions; and

at least one processor circuit to be programmed by the machine-readable instructions to:

generate one or more local representations of pose parts of an object based on a first image of the object and a feature map associated with a global representation of the object, the pose parts associated with a skeleton structure of the object in a pose in the first image, each local representation including local part features of the object associated with a respective pose part, the first image captured by a first camera;

aggregate the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features;

generate a fused representation of the object based on the aggregated local features and the global representation of the object; and

re-identify the object in a second image based on the fused representation, the second image captured using a second camera different than the first camera.

2 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to generate the global representation, the global representation including the feature map.

3 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to estimate pose keypoints in the first image and generate the skeleton structure of the object based on the pose keypoints.

4 . The apparatus of claim 1 , wherein the local representations include star structure models.

5 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to extract the local representations from the global representation using regional average pooling.

6 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to execute a deep neural network trained using a fused-triplet loss function to re-identify the object.

7 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to execute a deep neural network to generate the fused representation and re-identify the object.

8 . The apparatus of claim 1 , wherein an identity of the object is defined by an attribute of the object in the first image.

9 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

output a weight sum vector based on the weighted summation; and

execute a neural network to generate an identity loss vector from the weighted sum vector.

10 . A method comprising:

globally modeling, by at least one processor circuit programmed by at least one instruction, an object based on a first input object image to generate a global representation of the object, the global representation including a feature map, the first input object image captured by a first camera;

estimating, by one or more of the at least one processor circuit, pose keypoints of the object in the first input object image;

generating, by one or more of the at least one processor circuit, a skeleton structure of the object based on the pose keypoints;

modeling, by one or more of the at least one processor circuit, local parts of the object in the first input object image based on the feature map and the skeleton structure to generate one or more local representations of pose parts of the object, the pose parts associated with the skeleton structure of the object in a pose in the first input object image, each local representation including local part features of the object associated with a respective pose part;

aggregating, by one or more of the at least one processor circuit, the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features;

generating, by one or more of the at least one processor circuit, a fused representation of the object based on the aggregated local features and the global representation of the object; and

re-identifying, by one or more of the at least one processor circuit, the object in a second input object image based on the fused representation, the second input object image captured using a second camera different than the first camera.

11 . The method of claim 10 , wherein modeling the local parts includes extracting the local representations from the global representation using regional average pooling.

12 . The method of claim 10 , wherein re-identifying the object includes executing a deep neural network for the second input object image, the deep neural network to output a re-identification of the object.

13 . The method of claim 10 , wherein globally modeling the object includes generating bounding boxes enclosing regions of a first input object image corresponding to different pose parts of the object.

14 . The method of claim 10 , wherein estimating the pose keypoints includes estimating the pose keypoints using a number of pose keypoints based on a category of the object.

15 . The method of claim 14 , further including training one or more deep neural networks to globally model the object, estimate the pose keypoints, model the local parts of the object, and fuse the global representation of the object with the local representations of the object.

16 . The method of claim 10 , wherein fusing the global representation with the local representations includes executing a deep neural network to perform a global transformation on the aggregated local features using a triplet hard loss function.

17 . The method of claim 10 , further including:

outputting a weight sum vector based on the weighted summation; and

executing a neural network to generate an identity loss vector from the weighted sum vector.

18 . A system comprising:

means for generating one or more local representations of pose parts of an object based on a first image of the object and a feature map associated with a global representation of the object, the pose parts associated with a skeleton structure of the object in a pose in the first image, each local representation including local part features of the object associated with a respective pose part, the first image captured by a first camera;

means for fusing to generate a fused representation of the object, the means for fusing to:

aggregate the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features; and

generate a fused representation of the object based on the aggregated local features and the global representation of the object; and

means for re-identifying the object in a second image based on the fused representation, the second image captured using a second camera different than the first camera.

19 . The system of claim 18 , further including means for generating the global representation, the global representation including the feature map.

20 . The system of claim 18 , further including means for estimating pose keypoints in the first image to generate the skeleton structure of the object.

21 . The system of claim 18 , wherein the local representations include star structure models.

22 . The system of claim 18 , wherein the means for fusing is to:

output a weight sum vector based on the weighted summation; and

execute a neural network to generate an identity loss vector from the weighted sum vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: LI, JIANGUO; LI, SHUYUAN; TANG, HANLIN
To: INTEL CORPORATION
Reel/Frame 063119/0004 →
Continuity (1)
Related Publication 20220343639A1 · Oct 27, 2022
References Cited (128)
US 8615105B1 · Cheng et al. · 2013 [cited by applicant]
US 9916508B2 · Pillai et al. · 2018 [cited by applicant]
US 9948902B1 · Trundle · 2018 [cited by examiner]
US 10176405B1 · Zhou et al. · 2019 [cited by applicant]
US 10304191B1 · Mousavian et al. · 2019 [cited by applicant]
US 10321728B1 · Koh · 2019 [cited by applicant]
US 10332264B2 · Schulter et al. · 2019 [cited by applicant]
US 10402983B2 · Schulter et al. · 2019 [cited by applicant]
US 10430966B2 · Varadarajan et al. · 2019 [cited by applicant]
US 10733441B2 · Mousavian et al. · 2020 [cited by applicant]
US 10839543B2 · Cheng · 2020 [cited by applicant]
US 10853970B1 · Akbas et al. · 2020 [cited by applicant]
US 11715213B2 · Leung et al. · 2023 [cited by applicant]
US 12095973B2 · Bylicka et al. · 2024 [cited by applicant]
US 20130322720A1 · Hu · 2013 [cited by examiner]
US 20150095360A1 · Vrcelj et al. · 2015 [cited by applicant]
US 20160267331A1 · Pillai et al. · 2016 [cited by applicant]
US 20170316578A1 · Fua · 2017 [cited by applicant]
US 20180130215A1 · Schulter et al. · 2018 [cited by applicant]
US 20180130216A1 · Schulter et al. · 2018 [cited by applicant]
US 20180173969A1 · Pillai et al. · 2018 [cited by applicant]
US 20180293445A1 · Gao · 2018 [cited by applicant]
US 20180350105A1 · Taylor · 2018 [cited by applicant]
US 20190026917A1 · Liao · 2019 [cited by examiner]
US 20190066326A1 · Tran et al. · 2019 [cited by applicant]
US 20190171909A1 · Mandal et al. · 2019 [cited by applicant]
US 20190220992A1 · Li · 2019 [cited by examiner]
US 20190278983A1 · Iqbal et al. · 2019 [cited by applicant]
US 20190340432A1 · Mousavian et al. · 2019 [cited by applicant]
US 20190371080A1 · Sminchisescu · 2019 [cited by applicant]
US 20200058137A1 · Pujades · 2020 [cited by applicant]
US 20200074678A1 · Ning et al. · 2020 [cited by applicant]
US 20200082180A1 · Wang · 2020 [cited by applicant]
US 20200126297A1 · Tian · 2020 [cited by applicant]
US 20200160102A1 · Bruna et al. · 2020 [cited by applicant]
US 20200193628A1 · Chakraborty · 2020 [cited by applicant]
US 20200327418A1 · Lyons · 2020 [cited by applicant]
US 20200364454A1 · Mousavian et al. · 2020 [cited by applicant]
US 20200372246A1 · Chidananda · 2020 [cited by applicant]
US 20200401793A1 · Leung et al. · 2020 [cited by applicant]
US 20210000404A1 · Wang · 2021 [cited by applicant]
US 20210042520A1 · Molin · 2021 [cited by applicant]
US 20210097718A1 · Fisch · 2021 [cited by applicant]
US 20210097759A1 · Agrawal · 2021 [cited by applicant]
US 20210112238A1 · Bylicka et al. · 2021 [cited by applicant]
US 20210117648A1 · Yang · 2021 [cited by examiner]
US 20210192783A1 · Huelsdunk · 2021 [cited by applicant]
US 20210209797A1 · Lee · 2021 [cited by applicant]
US 20210350555A1 · Fischetti · 2021 [cited by applicant]
US 20210366146A1 · Khamis et al. · 2021 [cited by applicant]
US 20220172429A1 · Tong · 2022 [cited by applicant]
US 20220343639A1 · Li et al. · 2022 [cited by applicant]
US 20220351535A1 · Tao et al. · 2022 [cited by applicant]
US 20230186567A1 · Koh · 2023 [cited by applicant]
US 20230298204A1 · Wang · 2023 [cited by applicant]
CN 108108674 · 2018 [cited by applicant]
CN 108629801A · 2018 [cited by applicant]
CN 108830139A · 2018 [cited by applicant]
CN 108960036A · 2018 [cited by applicant]
CN 108986197A · 2018 [cited by applicant]
CN 109886090A · 2019 [cited by applicant]
CN 109948587A · 2019 [cited by applicant]
CN 110008913 · 2019 [cited by applicant]
CN 110009722A · 2019 [cited by applicant]
CN 110458940A · 2019 [cited by applicant]
CN 110516670A · 2019 [cited by applicant]
KR 20190087258A · 2019 [cited by applicant]
WO 2019025729A1 · 2019 [cited by applicant]
WO 2021109118A1 · 2021 [cited by applicant]
WO 2021120157A1 · 2021 [cited by applicant]
WO 2021258386A1 · 2021 [cited by applicant]
Fu, Xinchuan, et al. “Delving deep into multiscale pedestrian detection via single scale feature maps.” Sensors 18.4 (2018): 1063. https://www.mdpi.com/1424-8220/18/4/1063 (Year: 2018). [cited by examiner]
Chen, Dapeng, et al. “Video Person Re-identification with Competitive Snippet-Similarity Aggregation and Co-attentive Snippet Embedding.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018.h… [cited by examiner]
International Searching Authority, “Search Report,” issued in connection with International Patent Application No. PCT/CN2019/123625, mailed on Sep. 9, 2020, 4 pages. [cited by applicant]
International Searching Authority, “Written opinion,” issued in connection with International Patent Application No. PCT/CN2019/123625, mailed on Sep. 9, 2020, 5 pages. [cited by applicant]
Zajdel et al., “Keeping Track of Humans: Have I Seen This Person Before?” ResearchGate, May 2005, 7 pages. [cited by applicant]
Delannay et al., “Detection and Recognition of Sports(wo)men from Multiple Views,” IEEE International Conference on Distributed Smart Cameras, Aug. 30-Sep. 2, 2009, 7 pages. [cited by applicant]
Felzenszwalb et al. “Object Detection with Discriminatively Trained Part Based Models,” IEEE Trans on PAMI 2010, 20 pages. [cited by applicant]
Andriluka et al., “2D Human Pose Estimation: New Benchmark and State of the Art Analysis,” Computer Vision Foundation, 2014, 8 pages. [cited by applicant]
Zhang et al., “Part-based R-CNNs for Fine-grained Category Detection,” arXiv:1407.3867v1 [cs.CV], Jul. 15, 2014, 16 pages. [cited by applicant]
Zheng et al., “Scalable Person Re-identification: A Benchmark,” Computer Vision Foundation, 2015, 9 pages. [cited by applicant]
Schroff et al., “FaceNet: A unified Embedding for Face Recognition and Clustering,” arXiv: 1503.03832v3 [cs.cv], Jun. 17, 2015, 10 pages. [cited by applicant]
Joo et al., “Panoptic Studio: A Massively Multiview System for Social Interaction Capture,” Dec. 9, 2016, Retrieved from the Internet: <https://arxiv.org/abs/1612.03153> 14 pages. [cited by applicant]
Pavlakos et al., “Harvesting Multiple Views for Marker-less 3D Human Pose Annotations,” Apr. 16, 2017, Retrieved from the Internet: <https://arxiv.org/abs/1704.04793> 10 pages. [cited by applicant]
Zhong et al., “Random Erasing Data Augmentation,” arXiv:1708.04896v2 [cs.CV], Nov. 16, 2017, 10 pages. [cited by applicant]
Hermans et al., “In Defense of the Triplet Loss for Person Re-Identification,” arXiv:1703.07737v4 [cs.CV], Nov. 21, 2017, 17 pages. [cited by applicant]
Hu et al., “Squeeze-and-Excitation Networks,” Computer Vision Foundation, 2018, 10 pages. [cited by applicant]
Sun et al., “Beyond Part Models: Person Retrieval with Refined Part Pooling (and A Strong Convolutional Baseline),” arXiv:1711 .09349v3 [cs.cv], Jan. 9, 2018, 10 pages. [cited by applicant]
Zhang et al., “AlignedReID: Surpassing Human-Level Performance in Person Re-Identification,” arXiv.1711.08184v2 [cs.CV], Jan. 31, 2018, 10 pages. [cited by applicant]
Rhodin et al., “Learning Monocular 3D Human Pose Estimation from Multi-view Images,” Mar. 24, 2018, Retrieved from the Internet: <https://arxiv.org/abs/1803.04775>, 10 pages. [cited by applicant]
Zhong, Z. et al., “Camera Style Adaptation for Person Re-identification”, CVPR (2018), pp. 5157-5166. [cited by applicant]
Wang et al., “Person Re-identification with Cascaded Pairwise Convolutions,” Jun. 18-23, 2018, IEEE/CVF Conference on Computer Vision and Pattern Recognition, Retrieved from the Internet: <https://ieeexplore.ieee.org/do… [cited by applicant]
Ionescu et al., “Human3.6M Dataset,” Retrieved from the Internet: http://vision.imar.ro/human3.6m/description.php, 1 page. [cited by applicant]
Wang et al., “Learning Discriminative Features with Multiple Granularities for Person Re-Identification,” arXiv:1804.01438v3 [cs.CV], Aug. 17, 2018, 9 pages. [cited by applicant]
Wojke et al., “Deep Cosine Metric Learning for Person Re-Identification,” arXiv:1812.00442v1 [cs.CV], Dec. 2, 2018, 9 pages. [cited by applicant]
Intel, “2019 CES: Intel and Alibaba Team on New AI-Powered 3D Athlete Tracking Technology Aimed at the Olympic Games Tokyo 2020,” Retrieved from the Internet: [https://newsroom.intel.com/news/intel-alibaba-team-ai-power… [cited by applicant]
Schwarcz et al., “3D Human Pose Estimation from Deep Multi-View 2D Pose,” Feb. 7, 2019, Retrieved from the Internet: <https://arXiv:1902.02841v1> 6 pages. [cited by applicant]
Sun et al., “Deep High-Resolution Representation Learning for Human pose Estimation,” arXiv:1902.09212v1 [cs.CV], Feb. 25, 2019, 12 pages. [cited by applicant]
Luo et al., Bag of Tricks and a Strong Baseline for Deep Person Re-identification, arXiv: 1903.070710 [cs.cv], Apr. 19, 2019, 9 pages. [cited by applicant]
Iskakov et al., “Learnable Triangulation of Human Pose,” May 14, 2019, Retrieved from the Internet: <https://arXiv:1905.05754v1> 9 pages. [cited by applicant]
Dong, J. et al., “Fast and robust multi-person 3D pose estimation from multiple views”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; pp. 7792-7801, 2019. [cited by applicant]
Arbues-Sanguesa et al., “Multi-Person tracking by multi-scale detection in Basketball scenarios,” arXiv:1907.04637v1, Jul. 10, 2019, 8 pages. [cited by applicant]
Chen et al., “Spatial-Temporal Attention-Aware Learning for Video-Based Person Re-Identification,” IEEE Transactions on Image Processing, vol. 28, Issue 9, Sep. 2019, 14 pages. [cited by applicant]
Min Xin et al., “Motion Capture Research: 3D Human Pose Recovery Based on RGB Video Sequences,” Applied Sciences, vol. 9, No. 17, Sep. 2, 2019, pp. 1-22, XP55885594, DOI:10.3390/app9173613, 22 pages. [cited by applicant]
Qiu et al., “Cross View Fusion for 3D Human Pose Estimation,” Sep. 3, 2019, Retrieved from the Internet: <https://arxiv.org/abs/1909.01203> 10 pages. [cited by applicant]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/CN2019/126906, mailed on Sep. 23, 2020, 3 pages. [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/CN2019/126906, mailed on Sep. 23, 2020, 3 pages. [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/CN2020/098306, mailed on Mar. 25, 2021, 5 pages. [cited by applicant]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/CN2020/098306, mailed on Mar. 25, 2021, 4 pages. [cited by applicant]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/US2021/050609, mailed on Dec. 28, 2021, 5 pages. [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/US2021/050609, mailed on Dec. 28, 2021, 5 pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/CN2019/123625, mailed on Jun. 16, 2022, 6 pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/CN2019/126906, mailed on Jun. 30, 2022, 5 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 16/914,232, dated Jul. 21, 2022, 13 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 16/914,232, dated Nov. 30, 2022, 9 pages. [cited by applicant]
International Bureau, “International Preliminary Report on Patentability,” issued in connection with International Appl. No. PCT/CN2020/098306, dated Dec. 13, 2022, 5 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 18/000,389, dated Jul. 1, 2024, 18 pages. [cited by applicant]
Intel, “Intel True View,” https://www.intel.com/content/www/us/en/sports/technology/true-view.html, last accessed Feb. 24, 2023. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 16/914,232, dated Mar. 15, 2023, 8 pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/US2021/050609, mailed on Jul. 6, 2023, 7 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/131,433, dated Feb. 15, 2024, 11 pages. [cited by applicant]
European Patent Office, “Extended European Search Report,” issued in connection with European patent Application No. 20941827.6, dated Feb. 29, 2024, 9 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/131,433, dated May 16, 2024, 6 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/764,100, dated Jun. 12, 2024, 10 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued in connection with U.S. Appl. No. 18/000,389, dated Oct. 18, 2024, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/764,100, dated Jan. 13, 2025, 9 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 18/000,389, dated Jan. 15, 2025, 10 pages. [cited by applicant]
European Patent Office, “Extended European Search Report,” issued in connection with European Patent Application No. 251179962.3, dated Aug. 26, 2025, 10 pages. [cited by applicant]