IP Library › Granted Patent US 12,406,502
Granted Patent B2
US 12,406,502 · App. 17/873,462 · Granted Sep 2, 2025

Enriched and discriminative convolutional neural network features for pedestrian re-identification and trajectory modeling

Inventors: Kok Yiu Wong (Hong Kong, CN); Jack Chin Pang Cheng (Hong Kong, CN)
Assignee: The Hong Kong University of Science and Technology
G06V20/52G06T7/292G06V10/46G06V10/761G06V10/762
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,502
App. No.
17/873,462
Granted
Sep 2, 2025
Kind
B2
Abstract

A system for pedestrian re-identification comprises multiple cameras and a computing system. The cameras are configured to obtain data. The data comprises multiple images associated with one or more objects. The computing system is configured to extract features for each of the multiple images, determine a descriptor for each of the multiple images, and identify one or more images among the multiple images associated with an object among the one or more objects based on the descriptors of the multiple images. A respective descriptor for a respective image comprises a first set of units representing all the extracted features in the respective image and a second set of units representing a subset of the extracted features in the respective image.

Claims (48)

1. A method for image processing, comprising:

obtaining, by a processing system, data from a plurality of cameras, the data comprising multiple images associated with one or more objects;

extracting, by the processing system, features for each of the multiple images;

determining, by the processing system, a descriptor for each of the multiple images, wherein a respective descriptor for a respective image comprises a first set of units representing all the extracted features in the respective image and a second set of units representing a subset of the extracted features in the respective image, wherein the respective descriptor is represented by a vector comprising a plurality of elements, and wherein each element corresponds to a feature among all the extracted features;

grouping, by the processing system, the descriptors into a number of clusters based on similarities between the descriptors;

determining, by the processing system, an aggregated descriptor for each cluster of the number of clusters, wherein determining the aggregated descriptor for each cluster further comprises determining a mean value for each element in the aggregated descriptor based on corresponding elements in descriptors in a cluster; and

identifying, by the processing system, based on the aggregated descriptors for the number of clusters, one or more images among the multiple images associated with an object among the one or more objects.

2. The method of claim 1 , wherein the first set of units in the respective descriptor is determined based on a first feature map of the respective image corresponding to the respective descriptor, wherein the second set of units in the respective descriptor is determined based on a second feature map associated with the first feature map, and wherein the second feature map is generated by removing a patch of information included in the first feature map, and the patch in the first feature map takes the full height of the first feature map and covers a portion of the full width of the first feature map.

3. The method of claim 1 , wherein the first set of units in the respective descriptor is determined based on a first feature map of the respective image corresponding to the respective descriptor, wherein the second set of units in the respective descriptor is determined based on a second feature map associated with the first feature map, and wherein the second feature map is generated by removing a patch of information included in the first feature map, and the patch in the first feature map takes the full width of the first feature map and covers a portion of the full height of the first feature map.

4. The method of claim 1 , wherein the aggregated descriptor corresponding to a respective cluster is updated in response to the respective cluster being updated.

5. The method of claim 1 , wherein the data from the plurality of cameras comprises timestamp information and camera identity information corresponding to the multiple images.

6. The method of claim 5 , further comprising:

determining a trajectory for each of the one or more objects across the plurality of cameras based on the timestamp information and camera identity information.

7. The method of claim 6 , further comprising:

tracking the trajectories of the one or more objects across the plurality of cameras; and

determining travel preferences and demands of the one or more objects.

8. The method of claim 1 , wherein each of the multiple images is a portion of a raw image, and wherein the portion of the raw image comprises one object among the one or more objects.

9. A system for image processing, comprising:

a plurality of cameras configured to obtain data, the data comprising multiple images associated with one or more objects; and

a computing system configured to:

extract features for each of the multiple images;

determine a descriptor for each of the multiple images, wherein a respective descriptor for a respective image comprises a first set of units representing all the extracted features in the respective image and a second set of units representing a subset of the extracted features in the respective image, wherein the respective descriptor is represented by a vector comprising a plurality of elements, and wherein each element corresponds to a feature among all the extracted features;

group the descriptors into a number of clusters based on similarities between the descriptors;

determine an aggregated descriptor for each cluster of the number of clusters, wherein determining the aggregated descriptor for each cluster further comprises determining a mean value for each element in the aggregated descriptor based on corresponding elements in descriptors in a cluster; and

identify, based on the aggregated descriptors for the number of clusters, one or more images among the multiple images associated with an object among the one or more objects.

10. The system of claim 9 , wherein the first set of units in the respective descriptor is determined based on a first feature map of the respective image corresponding to the respective descriptor, wherein the second set of units in the respective descriptor is determined based on a second feature map associated with the first feature map, and wherein the second feature map is generated by removing a patch of information included in the first feature map, and the patch in the first feature map takes the full height of the first feature map and covers a portion of the full width of the first feature map.

11. The system of claim 9 , wherein the first set of units in the respective descriptor is determined based on a first feature map of the respective image corresponding to the respective descriptor, wherein the second set of units in the respective descriptor is determined based on a second feature map associated with the first feature map, and wherein the second feature map is generated by removing a patch of information included in the first feature map, and the patch in the first feature map takes the full width of the first feature map and covers a portion of the full height of the first feature map.

12. The system of claim 9 , wherein the data from the plurality of cameras comprises timestamp information and camera identity information corresponding to the multiple images.

13. The system of claim 12 , wherein the computing system is further configured to:

determine a trajectory for each of the one or more objects across the plurality of cameras based on the timestamp information and camera identity information.

14. The system of claim 13 , wherein the computing system is further configured to:

track the trajectories of the one or more objects across the plurality of cameras; and

determine travel preferences and demands of the one or more objects.

15. A non-transitory computer-readable medium having processor-executable instructions stored thereon for image processing, wherein the processor-executable instructions, when executed, facilitate:

obtaining data from a plurality of cameras, the data comprising multiple images associated with one or more objects;

extracting features for each of the multiple images;

determining a descriptor for each of the multiple images, wherein a respective descriptor for a respective image comprises a first set of units representing all the extracted features in the respective image and a second set of units representing a subset of the extracted features in the respective image, wherein the respective descriptor is represented by a vector comprising a plurality of elements, and wherein each element corresponds to a feature among all the extracted features;

grouping the descriptors into a number of clusters based on similarities between the descriptors;

determining an aggregated descriptor for each cluster of the number of clusters, wherein determining the aggregated descriptor for each cluster further comprises determining a mean value for each element in the aggregated descriptor based on corresponding elements in descriptors in a cluster; and

identifying based on the aggregated descriptors for the number of clusters, one or more images among the multiple images associated with an object among the one or more objects.

16. The non-transitory computer-readable medium of claim 15 , wherein the data from the plurality of cameras comprises timestamp information and camera identity information corresponding to the multiple images.

17. The non-transitory computer-readable medium of claim 16 , wherein the processor-executable instructions, when executed, further facilitate:

determining a trajectory for each of the one or more objects across the plurality of cameras based on the timestamp information and camera identity information.

18. The non-transitory computer-readable medium of claim 17 , wherein the processor-executable instructions, when executed, further facilitate:

tracking the trajectories of the one or more objects across the plurality of cameras; and

determining travel preferences and demands of the one or more objects.

19. The system of claim 9 , wherein the aggregated descriptor corresponding to a respective cluster is updated in response to the respective cluster being updated.

20. The non-transitory computer-readable medium of claim 15 , wherein the aggregated descriptor corresponding to a respective cluster is updated in response to the respective cluster being updated.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2022
From: WONG, KOK YIU; CHENG, JACK CHIN PENG
To: THE HONG KONG UNIVERSITY OF SCIENCE AND TECHNOLOGY
Reel/Frame 060623/0449 →
Continuity (2)
Provisional Application 63249025 · Sep 28, 2021
Related Publication 20230095533A1 · Mar 30, 2023
References Cited (80)
US 20220415023A1 · Han · 2022 [cited by examiner]
Li, Peng, et al. “State-aware re-identification feature for multi-target multi-camera tracking.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 2019. (Year: 2019). [cited by examiner]
Wong, Peter Kok-Yiu, et al. “Enriched and discriminative human features for person re-identification based on explainable behaviors of convolutional neural networks.” International Conference on Computing in Civil and B… [cited by examiner]
Alam, K, et al., “A dynamic ensemble learning algorithm for neural networks” [cited by applicant]
Cha, Y.J., et al., “Deep Learning-Based Crack Damage Detection Using Convolutional Neural Networks”, [cited by applicant]
Chi, S., et al., A Methodology for Object Identification and Tracking in Construction Based on Spatial Modeling and Image Matching Techniques, [cited by applicant]
Chmielewska, A., et al., “New Approach to Traffic Density Estimation Based on Indoor and Outdoor Scenes from CCTV”, [cited by applicant]
Chu, C.Y., “A Computer Model for Selecting Facility Evacuation Design Using Cellular Automata”, [cited by applicant]
Dai, Z., et al., “Batch DropBlock Network for Person Re-identification and Beyond”, [cited by applicant]
Ewing, R., “Eight Qualities of Pedestrian and Transit-Oriented Design”, [cited by applicant]
Gao, J., et al., “Revisiting Temporal Modeling for Video-based Person ReID”, [cited by applicant]
He, K., et al., “Deep Residual Learning for Image Recognition”, [cited by applicant]
Hermans, A., et al., “In Defense of the Triplet Loss for Person Re-Identification”, [cited by applicant]
Huang, J. et al., “Speed/accuracy trade-offs for modern convolutional objects detectors”, [cited by applicant]
Kelly, C.E., et al., “A comparison of three methods for assessing the walkability of the pedestrian environment”, [cited by applicant]
Li, P., et al., “State-aware Re-identification Feature for Multi-target Multi-camera Tracking”, [cited by applicant]
Li, Shengyuan, et al., “Automatic pixel-level multiple damage detection of concrete structure using fully convolutional network”, [cited by applicant]
Li, Shuang, et al., “Diversity Regularized Spatiotemporal Attention for Video-based Peron Re-identification”, [cited by applicant]
Li, W, et al., DeepReID: Deep Filter Pairing Neural Network for Person Re-Identification, [cited by applicant]
Li, W., et al., “Harmonious Attention Network for Person Re-Identification”, [cited by applicant]
Liang, X., “Image-based post-disaster inspection of reinforced concrete bridge systems using deep learning with Bayesian optimization”, [cited by applicant]
Luo, W., et al., “Multiple Object Tracking: A Literature Review”, [cited by applicant]
Mignon, A., et al., “PCCA: A new approach for distance learning from sparse pairwise constraints”, [cited by applicant]
Minoli, D., et al. “IoT Considerations, Requirements, and Architectures for Smart Buildings—Energy Optimization and Next Generation Building Management Systems”, [cited by applicant]
Nambiar, A, et al., “A multi-camera video dataset for research on high-definition surveillance”, [cited by applicant]
Pathak, A.R., et al., “Application of Deep Learning for Object Detection”, [cited by applicant]
Pedagadi, S., et al., “Local Fisher Discriminant Analysis for Pedestrian Re-identification”, [cited by applicant]
Qian, X., et al., “Multi-scale Deep Learning Architectures for Person Re-identification”, [cited by applicant]
Quispe, R. et al., “Improved Person Re-Identification Based on Saliency and Semantic Parsing with Deep Neural Network Models”, [cited by applicant]
Rabie, T., et al., “Mobile Active-Vision Traffic Surveillance System for Urban Networks”, [cited by applicant]
Ristani, E., et al., “Features for Multi-Target Multi-Camera Tracking and Re-Identification”, [cited by applicant]
Samek, W., et al., “Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models”, [cited by applicant]
Wang, H., et al., “Parameter-Free Spatial Attention Network for Person Re-Identification”, [cited by applicant]
Wojke, N., et al., “Deep Cosine Metric Learning for Person Re-Identification”, [cited by applicant]
Wu, R.T., “Pruning deep convolutional neural networks for efficient edge computing in condition assessment of infrastructures”, [cited by applicant]
Xiao, T., et al., “Learning Deep Feature Representations with Domain Guided Dropout for Person Re-identification”, [cited by applicant]
Yang. W., et al., “Towards Rich Feature Discovery with Class Activation Maps Augmentation for Peron Re-Identification”, [cited by applicant]
Zagoruyko, S., et al., “Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer”, [cited by applicant]
Zhang, Z., et al., “Multi-Target, Multi-Camera Tracking by Hierarchical Clustering; Recent Progress on DukeMTMC Project”, [cited by applicant]
Zheng, L., et al. “Person Re-identification: Past, Present and Future”, [cited by applicant]
Zhong, Z., et al., “Random Erasing Data Augmentation”, [cited by applicant]
Zhou, K., et al., “Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch”, [cited by applicant]
Zhou, K., et al., “Learning Generalisable Omni-Scale Representations for Person Re-Identification”, [cited by applicant]
Transportation for London, “Walking Action Plan: Making London the World's Most Walkable City”, [cited by applicant]
Arabi, S., et al. “A deep-learning-based computer vision solution for construction vehicle detection”, [cited by applicant]
Barney, G., et al., [cited by applicant]
Brunetti, A., et al., “Computer vision and deep learning techniques for pedestrian detection and tracking: A Survey”, [cited by applicant]
Chen, X., et al., “Kinect-Based Pedestrian Detection for Crowded Scenes,” [cited by applicant]
Chi, S., et al., “Automated Object Identification Using Optical Video Cameras on Construction Sites”, [cited by applicant]
Chou, Yu-Sheng, et al., “Dynamic Gallery for Real-time Multi-target Multi-camera Tracking”, [cited by applicant]
Cimellero, G.P., et al., “Integrating a Human Behavior Model within an Agent-Based Approach for Blasting Evacuation”, [cited by applicant]
Gao, Y., et al., “Deep leaf-bootstrapping generative adversarial network for structural image data augmentation”, [cited by applicant]
Han, J., et al., “Data Mining: Concepts and Techniques”, [cited by applicant]
Lee, H.Y., et al., “Laying out the occupant flows in public buildings for operating efficiency”, [cited by applicant]
Li, B., et al., “A Review on Vision-based Pedestrian Detection in Intelligent Transportation Systems”, [cited by applicant]
Liu, L., et al., “Deep Learning for Generic Object Detection: A Survey”, [cited by applicant]
Luo, X., et al., “Combining deep features and activity context to improve recognition of activities of workers in groups”, [cited by applicant]
Malinovskiy, Y., et al., “Model-Free Video Detection and tracking of Pedestrians and Bicyclists”, [cited by applicant]
Martani, C., et al., “Pedestrian monitoring techniques for crowd-flow Prediction”, [cited by applicant]
Rafiei, M.H., et al., “A New Neural Dynamic Classification Algorithm”, [cited by applicant]
Ristani, E., et al., “Performance Measures and a Data Set for Multi-target, Multi-camera Tracking”, [cited by applicant]
Schimpl, M., et al., “Association between Walking Speed and Age in Healthy, Free-Living Individuals Using Mobile Accelerometry—A Cross-Sectional Study”, [cited by applicant]
Sharifi, M.S., et al., “A large-scale controlled experiment on pedestrian walking behavior involving individuals with disabilities”, [cited by applicant]
Shen, J., et al., “A convolutional neural-network-based pedestrian counting model for various crowded scenes”, [cited by applicant]
World Urbanization Prospects, The 2018 Revision, [cited by applicant]
Wang, G. et al., “Learning Discrimination Features with Multiple Granularities for Person Re-Identification,” [cited by applicant]
Wang, M., et al., “A unified convolutional neural network integrated with conditional random field for pipe defect segmentation”, [cited by applicant]
Wang, N., et al., “Autonomous damage segmentation and measurement of glazed tiles in historic buildings via deep learning”, [cited by applicant]
Wu, D., “Deep learning-based methods for person re-identification: A comprehensive review,” [cited by applicant]
Zhang, B., et al., “A methodology for obtaining spatiotemporal information of the vehicles on bridges based on computer vision”, [cited by applicant]
Zheng, J., et al., “Detecting Cycle Failures at Signalized Intersections Using Video Image Processing”, [cited by applicant]
Adeli, H., “Neural Networks in Civil Engineering: 1989-2000”, [cited by applicant]
Wong, P.K., et al., “Enriched and discriminative convolutional neural network features for pedestrian re-identification and trajectory modeling”, [cited by applicant]
Part 1—“Sustainable Communicates in Bronx: Leveraging Regional Rail for Access, Growth and Opportunity” [cited by applicant]
Part 2—Sustainable Communicates in Bronx: Leveraging Regional Rail for Access, Growth and Opportunity [cited by applicant]
Schwartz, W., et al., “Bicycle and Pedestrian Data: Sources, Needs and Gaps”, [cited by applicant]
Seneviratne, P.N., Analysis of factors affecting the choice of route of pedestrians, [cited by applicant]
Thornton, C., et al., “Pathfinder: An Agent-Based Egress Simulator”, In R. D. Peacock, E. D. Kuligowski, & J. D. Averill (Eds.), [cited by applicant]
Muraleetharan, T., et al., Overall Level of Services of Urban Walking Environment and Its Influence on Pedestrian Route Choice Behavior: Analysis of Pedestrian Travel, [cited by applicant]
Wu, Rih-Teng, “Pruning deep convolitional neural networks for efficient edge computing in condition assessment of infrastructures”, [cited by applicant]