IP Library › Granted Patent US 12,080,100
Granted Patent B2
US 12,080,100 · App. 17/519,986 · Granted Sep 3, 2024

Face-aware person re-identification system

Inventors: Yumin Suh (Santa Clara, CA); Xiang Yu (Mountain View, CA); Yi-Hsuan Tsai (Santa Clara, CA); Masoud Faraki (San Jose, CA); Manmohan Chandraker (Santa Clara, CA)
Assignee: NEC Corporation
G06V40/172G06F18/214G06V20/52G06V40/103
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,080,100
App. No.
17/519,986
Granted
Sep 3, 2024
Kind
B2
Abstract

A method for employing facial information in unsupervised person re-identification is presented. The method includes extracting, by a body feature extractor, body features from a first data stream, extracting, by a head feature extractor, head features from a second data stream, outputting a body descriptor vector from the body feature extractor, outputting a head descriptor vector from the head feature extractor, and concatenating the body descriptor vector and the head descriptor vector to enable a model to generate a descriptor vector.

Claims (55)

1. A method for employing facial information in unsupervised person re-identification, the method comprising:

extracting, by a body feature extractor, body features from a first data stream;

extracting, by a head feature extractor, head features from a second data stream;

outputting a body descriptor vector from the body feature extractor;

outputting a head descriptor vector from the head feature extractor;

concatenating the body descriptor vector and the head descriptor vector to enable a model to generate a descriptor vector;

enhancing the first data stream using the head features and the second data stream using the body features by utilizing cross-task consistency as a self-supervision during training;

training the head feature extractor using a face knowledge distillation loss being applied randomly at each iteration with a particular probability;

employing a two-stream network where each stream processes body and head images separately, and utilizing both body and head appearances together;

preserving a relation between images across two modalities using cross-modal consistency loss;

employing additional face recognition loss for training the face stream of the model using pseudo labels for images from a target domain when frontal faces are visible; and

distilling knowledge obtained from a face recognition engine on the target domain to a head sub-network of the model.

2. The method of claim 1 , wherein the model is trained by using face-based pseudo labels, body-based pseudo labels, and head-based pseudo labels.

3. The method of claim 1 , wherein a face-based pseudo label loss is used to train the head feature extractor.

4. The method of claim 1 , wherein a cross-modal consistency loss is employed as a self-supervision during training.

5. The method of claim 1 , wherein a cross-modal consistency loss is employed in tandem with a face knowledge distillation loss.

6. The method of claim 1 , wherein the head feature extractor crops images from the first data stream around a head region and extracts head features by only using cropped regions.

7. The method of claim 1 , wherein temporal weight averaging is used to average network parameters over iterations to extract features for generating pseudo labels.

8. A non-transitory computer-readable storage medium comprising a computer-readable program for employing facial information in unsupervised person re-identification wherein the computer-readable program when executed on a computer causes the computer to perform steps of:

extracting, by a body feature extractor, body features from a first data stream;

extracting, by a head feature extractor, head features from a second data stream;

outputting a body descriptor vector from the body feature extractor;

outputting a head descriptor vector from the head feature extractor;

concatenating the body descriptor vector and the head descriptor vector to enable a model to generate a descriptor vector;

enhancing the first data stream using the head features and the second data stream using the body features by utilizing cross-task consistency as a self-supervision during training;

training the head feature extractor using a face knowledge distillation loss being applied randomly at each iteration with a particular probability;

employing a two-stream network where each stream processes body and head images separately, and utilizing both body and head appearances together;

preserving a relation between images across two modalities using cross-modal consistency loss;

employing additional face recognition loss for training the face stream of the model using pseudo labels for images from a target domain when frontal faces are visible; and

distilling knowledge obtained from a face recognition engine on the target domain to a head sub-network of the model.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the model is trained by using face-based pseudo labels, body-based pseudo labels, and head-based pseudo labels.

10. The non-transitory computer-readable storage medium of claim 8 , wherein a face-based pseudo label loss is used to train the head feature extractor.

11. The non-transitory computer-readable storage medium of claim 8 , wherein a cross-modal consistency loss is employed as a self-supervision during training.

12. The non-transitory computer-readable storage medium of claim 8 , wherein a cross-modal consistency loss is employed in tandem with a face knowledge distillation loss.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the head feature extractor crops images from the first data stream around a head region and extracts head features by only using cropped regions.

14. The non-transitory computer-readable storage medium of claim 8 , wherein temporal weight averaging is used to average network parameters over iterations to extract features for generating pseudo labels.

15. A system for employing facial information in unsupervised person re-identification, the system comprising:

a memory; and

one or more processors in communication with the memory configured to:

extract, by a body feature extractor, body features from a first data stream;

extract, by a head feature extractor, head features from a second data stream;

output a body descriptor vector from the body feature extractor;

output a head descriptor vector from the head feature extractor;

concatenate the body descriptor vector and the head descriptor vector to enable a model to generate a descriptor vector;

enhance the first data stream using the head features and the second data stream using the body features by utilizing cross-task consistency as a self-supervision during training;

train the head feature extractor using a face knowledge distillation loss being applied randomly at each iteration with a particular probability;

employ a two-stream network where each stream processes body and head images separately, and utilizing both body and head appearances together;

preserve a relation between images across two modalities using cross-modal consistency loss;

employ additional face recognition loss for training the face stream of the model using pseudo labels for images from a target domain when frontal faces are visible; and

distill knowledge obtained from a face recognition engine on the target domain to a head sub-network of the model.

16. The system of claim 15 , wherein the model is trained by using face-based pseudo labels, body-based pseudo labels, and head-based pseudo labels.

17. The system of claim 15 , wherein a face-based pseudo label loss is used to train the head feature extractor.

18. The system of claim 15 , wherein a cross-modal consistency loss is employed as a self-supervision during training.

19. The system of claim 15 , wherein a cross-modal consistency loss is employed in tandem with a face knowledge distillation loss.

20. The system of claim 15 , wherein the head feature extractor crops images from the first data stream around a head region and extracts head features by only using cropped regions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 067922/0808 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2021
From: SUH, YUMIN; YU, XIANG; TSAI, YI-HSUAN; FARAKI, MASOUD; CHANDRAKER, MANMOHAN
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 058031/0864 →
Continuity (3)
Provisional Application 63114030 · Nov 16, 2020
Provisional Application 63111809 · Nov 10, 2020
Related Publication 20220147735A1 · May 12, 2022