IP Library Granted Patent US 11,810,388
Granted Patent B1
US 11,810,388 · App. 18/268,943 · Granted Nov 7, 2023

Person re-identification method and apparatus based on deep learning network, device, and medium

Inventors: Li Wang (Shandong, CN); Baoyu Fan (Shandong, CN)
Assignee: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
G06V40/117G06V10/761G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,810,388
App. No.
18/268,943
Granted
Nov 7, 2023
Kind
B1
Abstract

The present application discloses a person re-identification method and apparatus based on a deep learning network, a device, and a medium. The method includes: obtaining an initial person re-identification network; creating a homogeneous training network corresponding to the initial person re-identification network, where the homogeneous training network comprises a plurality of homogeneous branches with a same network structure; training the homogeneous training network by using a target loss function, and determining a final weight parameter of each network layer in the homogeneous training network; and loading the final weight parameter by using the initial person re-identification network to obtain a final person re-identification network, to perform a person re-identification task by using the final person re-identification network.

Claims (482)

1. A person re-identification method based on a deep learning network, comprising:

obtaining an initial person re-identification network;

creating a homogeneous training network corresponding to the initial person re-identification network, wherein the homogeneous training network comprises a plurality of homogeneous branches with a same network structure;

training the homogeneous training network by using a target loss function, and determining a final weight parameter of each network layer in the homogeneous training network; and

loading the final weight parameter by using the initial person re-identification network to obtain a final person re-identification network, to perform a person re-identification task by using the final person re-identification network;

wherein the training the homogeneous training network by using a target loss function and determining a final weight parameter of each network layer in the homogeneous training network comprises:

during training of the homogeneous training network, determining a cross-entropy loss value of a cross-entropy loss function, determining a triplet loss value of a triplet loss function, determining a knowledge synergy for embedding distance loss value of a knowledge synergy for embedding distance loss function, and determining a probabilistic collaboration loss value of a probabilistic collaboration loss function, wherein the knowledge synergy for embedding distance loss function is used for determining the knowledge synergy for embedding distance loss value by using a Euclidean distance between embedding-layer output features of each sample in every two homogeneous branches; and

determining the final weight parameter of each network layer in the homogeneous training network by using a total loss value of the cross-entropy loss value, the triplet loss value, and the knowledge synergy for embedding distance loss value;

wherein a process of determining the probabilistic collaboration loss value of the probabilistic collaboration loss function comprises:

obtaining an image classification probability output by a classification layer of each homogeneous branch;

calculating an argmax value of the image classification probability of each homogeneous branch, and in response to a classification tag of the argmax value being the same as a real classification tag, outputting an embedding-layer output feature of the homogeneous branch, and outputting the argmax value of the homogeneous branch as a predicted probability value; and

determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch.

2. The person re-identification method according to claim 1 , wherein the determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch comprises:

determining a weight value of each homogeneous branch by using the predicted probability value output by each homogeneous branch;

determining a target feature according to a first feature determining rule, wherein the first feature determining rule is:

f

re

=

b

=

1

B

o

b

·

f

e

(

x

n

,

θ

b

)

,

where f re represents a target feature in current iterative training, B represents a total quantity of the plurality of homogeneous branches, b represents a b th homogeneous branch, o b represents a weight value of the b th homogeneous branch, x n represents an n th sample, θ b represents a network parameter of the b th homogeneous branch, and f e (x n ,θ b ) represents an embedding-layer output feature of x n in the b th homogeneous branch; and

determining the probabilistic collaboration loss value by using a first probabilistic collaboration loss function, wherein the first probabilistic collaboration loss function is:

L

vb

=

(

b

=

1

B

"\[LeftBracketingBar]"

f

e

(

x

n

,

θ

b

)

-

f

re

"\[RightBracketingBar]"

2

)

1

2

/

B

,

where L vb represents the probabilistic collaboration loss value.

3. The person re-identification method according to claim 2 , wherein after the determining a target feature according to a first feature determining rule, the method further comprises:

storing the target feature in each iterative training to a first-in first-out cache sequence as a historical feature;

determining a virtual branch feature by using a second feature determining rule, wherein the second feature determining rule is:

f

vb

=

α

·

f

re

+

β

·

(

j

=

1

J

cache

(

j

)

)

)

/

J

,

where f vb represents the virtual branch feature, α represents a first hyperparameter, β represents a second hyperparameter, J represents a quantity of historical features selected from the first-in first-out cache sequence, and cache(j) represents a j th historical feature selected from the first-in first-out cache sequence; and

determining the probabilistic collaboration loss value by using a second probabilistic collaboration loss function, wherein the second probabilistic collaboration loss function is:

L

vb

=

(

b

=

1

B

"\[LeftBracketingBar]"

f

e

(

x

n

,

θ

b

)

-

f

vb

"\[RightBracketingBar]"

2

)

1

2

/

B

.

4. The person re-identification method according to claim 1 , wherein a process of determining the triplet loss value of the triplet loss function comprises:

determining a first loss value of each homogeneous branch according to the embedding-layer output feature of each sample in each homogeneous branch and a first triple loss function; and

selecting a first loss value that is numerically minimum from each homogeneous branch as the triplet loss value;

wherein the first triplet loss function is:

L

TriHard

b

=

-

1

N

a

=

1

N

[

max

y

p

-

y

a

d

(

f

e

a

,

f

e

p

)

-

min

y

q

y

a

d

(

f

e

a

,

f

e

q

)

+

m

]

+

,

where L TriHard b represents a first loss value of a b th homogeneous branch, N represents a total quantity of training samples, a represents an anchor sample, f e a represents an embedding-layer output feature of the anchor sample, y represents a classification tag of the sample, p represents a sample that belongs to a same classification tag as the anchor sample and that is at a maximum intra-class distance from the anchor sample, f e p represents an embedding-layer output feature of the sample p, q represents a sample that belongs to a different classification tag from the anchor sample and that is at a minimum inter-class distance from the anchor sample, f e q represents an embedding-layer output feature of the sample q, m represents a first parameter, d(·,·) is used for calculating a distance, [·] + and max d(·,·) both represent calculation of a maximum distance, min d(·,·) represents calculation of a minimum distance, y a represents a classification tag of the anchor sample, y p represents a classification tag of the sample p, and y q represents a classification tag of the sample q.

5. The person re-identification method according to claim 4 , wherein after the determining a first loss value of each homogeneous branch, the method further comprises:

determining a second loss value of each homogeneous branch by using the first loss value of each homogeneous branch and a second triplet loss function, wherein

the second triplet loss function is:

L

E_TriHard

b

=

L

TriHard

b

+

β

1

N

a

=

1

N

(

d

(

f

e

a

,

f

e

p

)

d

(

f

e

a

,

f

e

q

)

,

where L E_TriHard b represents a second loss value of the b th homogeneous branch, and β represents a second parameter; and

correspondingly, the selecting a first loss value that is numerically minimum from each homogeneous branch as the triplet loss value comprises:

selecting the second loss value that is numerically minimum from each homogeneous branch as the triplet loss value.

6. The person re-identification method according to claim 1 , wherein a process of determining the knowledge synergy for embedding distance loss value of the knowledge synergy for embedding distance loss function comprises:

calculating the knowledge synergy for embedding distance loss value by using the embedding-layer output feature of each sample in each homogeneous branch and the knowledge synergy for embedding distance loss function, wherein the knowledge synergy for embedding distance loss function is:

L

k

s

e

=

1

N

n

=

1

N

u

=

1

B

-

1

v

=

u

+

1

B

(

h

=

1

H

|

f

e

h

(

x

n

,

θ

u

)

-

f

e

h

(

x

n

,

θ

v

)

|

2

)

1

2

,

where L kse represents the knowledge synergy for embedding distance loss value, N represents a total quantity of training samples, B represents a total quantity of the plurality of homogeneous branches, u represents a u th homogeneous branch, v represents a v th homogeneous branch, H represents a dimension of the embedding-layer output feature, x n represents an n th sample, f e h (x n ,θ u ) represents an embedding-layer output feature of x n in an h th dimension in the u th homogeneous branch, f e h (x n ,θ v ) represents an embedding-layer output feature of x n in the h th dimension in the v th homogeneous branch, |·| represents a distance, θ u represents a network parameter of the u th homogeneous branch, and θ v represents a network parameter of the v th homogeneous branch.

7. The person re-identification method according to claim 1 , wherein the creating a homogeneous training network corresponding to the initial person re-identification network comprises:

deriving an auxiliary training branch from an intermediate layer of the initial person re-identification network to generate a homogeneous training network with an asymmetric network structure, or deriving the auxiliary training branch from the intermediate layer of the initial person re-identification network to generate a homogeneous training network with a symmetric network structure.

8. The person re-identification method according to claim 7 , wherein the deriving an auxiliary training branch from an intermediate layer of the initial person re-identification network to generate a homogeneous training network with an asymmetric network structure, or deriving the auxiliary training branch from the intermediate layer of the initial person re-identification network to generate a homogeneous training network with a symmetric network structure comprises:

when a hardware device has high calculation performance, generating the homogeneous training network of the symmetric network structure; or

when the hardware device has average calculation performance, generating the homogeneous training network of the asymmetric network structure.

9. The person re-identification method according to claim 1 , wherein the training of the homogeneous training network comprises:

selecting a derivation position from a backbone network according to a network structure of the initial person Re-ID network, determining an intermediate layer from which an auxiliary training branch is derived, and constructing a homogeneous-network-based auxiliary training branch to obtain the homogeneous training network;

determining a target loss function, and calculating a loss of each homogeneous branch in the homogeneous training network by using the target loss function;

training a network according to the target loss function to converge the network; and

storing a trained weight parameter.

10. The person re-identification method according to claim 9 , wherein the training of the homogeneous training network comprises a first phase and a second phase;

the first phase is a forward propagation phase in which data is propagated from a lower layer to a higher layer, and the second phase is a back propagation phase in which an error is propagated for training from the higher layer to the lower layer when a result obtained by forward propagation is inconsistent with what is expected.

11. The person re-identification method according to claim 10 , wherein the training a network comprises:

initializing a weight of a network layer;

performing forward propagation on input training image data through each network layer, to obtain an output value;

calculating an error between the output value of the network and a target value;

back propagating the error to the network, and sequentially calculating a back propagation error of each network layer;

adjusting, by each network layer, all weight coefficients in the network according to the back propagation error of each layer;

reselecting randomly new training image data, and then performing the step of performing the forward propagation to obtain the output value of the network;

repeating infinitely iteration, and ending the training when an error between the output value of the network and a target value is less than a specific threshold or a quantity of iterations exceeds a specific threshold; and

storing trained network parameters of all layers.

12. The person re-identification method according to claim 11 , wherein the network layers comprise: a convolutional layer, a down-sampling layer, and a fully connected layer.

13. The person re-identification method according to claim 11 , wherein the calculating an error between the output value of the network and a target value comprises:

calculating the output value of the network and obtaining a total loss value based on the target loss function.

14. The person re-identification method according to claim 1 , wherein the loading the final weight parameter by using the initial person re-identification network to obtain a final person re-identification network comprises:

loading the final weight parameter by using the initial person re-identification network without auxiliary training branches to obtain the final person re-identification network.

15. The person re-identification method according to claim 1 , wherein the obtaining an image classification probability output by a classification layer of each homogeneous branch comprises:

obtaining fully connected-layer features; and

normalizing the fully connected-layer features by using a softmax function to obtain the image classification probability of the homogeneous branch.

16. The person re-identification method according to claim 1 , wherein the image classification probability is, when a person is identified, a similarity probability of the person and each object in a database; and

the argmax value is a maximum value in the image classification probability.

17. An electronic device, comprising:

a memory, configured to store a computer program; and

a processor, configured to execute the computer program to implement operations comprising:

obtaining an initial person re-identification network;

creating a homogeneous training network corresponding to the initial person re-identification network, wherein the homogeneous training network comprises a plurality of homogeneous branches with a same network structure;

training the homogeneous training network by using a target loss function, and determining a final weight parameter of each network layer in the homogeneous training network; and

loading the final weight parameter by using the initial person re-identification network to obtain a final person re-identification network, to perform a person re-identification task by using the final person re-identification network;

wherein the operation of training the homogeneous training network by using a target loss function and determining a final weight parameter of each network layer in the homogeneous training network comprises:

during training of the homogeneous training network, determining a cross-entropy loss value of a cross-entropy loss function, determining a triplet loss value of a triplet loss function, determining a knowledge synergy for embedding distance loss value of a knowledge synergy for embedding distance loss function, and determining a probabilistic collaboration loss value of a probabilistic collaboration loss function, wherein the knowledge synergy for embedding distance loss function is used for determining the knowledge synergy for embedding distance loss value by using a Euclidean distance between embedding-layer output features of each sample in every two homogeneous branches; and

determining the final weight parameter of each network layer in the homogeneous training network by using a total loss value of the cross-entropy loss value, the triplet loss value, and the knowledge synergy for embedding distance loss value;

wherein a process of determining the probabilistic collaboration loss value of the probabilistic collaboration loss function comprises:

obtaining an image classification probability output by a classification layer of each homogeneous branch;

calculating an argmax value of the image classification probability of each homogeneous branch, and in response to a classification tag of the argmax value being the same as a real classification tag, outputting an embedding-layer output feature of the homogeneous branch, and outputting the argmax value of the homogeneous branch as a predicted probability value; and

determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch.

18. The electronic device according to claim 17 , wherein the operation of determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch comprises:

determining a weight value of each homogeneous branch by using the predicted probability value output by each homogeneous branch;

determining a target feature according to a first feature determining rule, wherein the first feature determining rule is:

f

re

=

b

=

1

B

o

b

·

f

e

(

x

n

,

θ

b

)

,

where f re represents a target feature in current iterative training, B represents a total quantity of the plurality of homogeneous branches, b represents a b th homogeneous branch, o b represents a weight value of the b th homogeneous branch, x n represents an n th sample, θ b represents a network parameter of the b th homogeneous branch, and f e (x n ,θ b ) represents an embedding-layer output feature of x n in the b th homogeneous branch; and

determining the probabilistic collaboration loss value by using a first probabilistic collaboration loss function, wherein the first probabilistic collaboration loss function is:

L

vb

=

(

b

=

1

B

"\[LeftBracketingBar]"

f

e

(

x

n

,

θ

b

)

-

f

re

"\[RightBracketingBar]"

2

)

1

2

/

B

,

where L vb represents the probabilistic collaboration loss value.

19. A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement operations comprising:

obtaining an initial person re-identification network;

creating a homogeneous training network corresponding to the initial person re-identification network, wherein the homogeneous training network comprises a plurality of homogeneous branches with a same network structure;

training the homogeneous training network by using a target loss function, and determining a final weight parameter of each network layer in the homogeneous training network; and

loading the final weight parameter by using the initial person re-identification network to obtain a final person re-identification network, to perform a person re-identification task by using the final person re-identification network;

wherein the operation of training the homogeneous training network by using a target loss function and determining a final weight parameter of each network layer in the homogeneous training network comprises:

during training of the homogeneous training network, determining a cross-entropy loss value of a cross-entropy loss function, determining a triplet loss value of a triplet loss function, determining a knowledge synergy for embedding distance loss value of a knowledge synergy for embedding distance loss function, and determining a probabilistic collaboration loss value of a probabilistic collaboration loss function, wherein the knowledge synergy for embedding distance loss function is used for determining the knowledge synergy for embedding distance loss value by using a Euclidean distance between embedding-layer output features of each sample in every two homogeneous branches; and

determining the final weight parameter of each network layer in the homogeneous training network by using a total loss value of the cross-entropy loss value, the triplet loss value, and the knowledge synergy for embedding distance loss value;

wherein a process of determining the probabilistic collaboration loss value of the probabilistic collaboration loss function comprises:

obtaining an image classification probability output by a classification layer of each homogeneous branch;

calculating an argmax value of the image classification probability of each homogeneous branch, and in response to a classification tag of the argmax value being the same as a real classification tag, outputting an embedding-layer output feature of the homogeneous branch, and outputting the argmax value of the homogeneous branch as a predicted probability value; and

determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch.

20. The computer-readable storage medium according to claim 19 , wherein the operation of determining the probabilistic collaboration loss value according to the probabilistic collaboration loss function as well as the predicted probability value and the embedding-layer output feature that are output by each homogeneous branch comprises:

determining a weight value of each homogeneous branch by using the predicted probability value output by each homogeneous branch;

determining a target feature according to a first feature determining rule, wherein the first feature determining rule is:

f

re

=

b

=

1

B

o

b

·

f

e

(

x

n

,

θ

b

)

,

where f re represents a target feature in current iterative training, B represents a total quantity of the plurality of homogeneous branches, b represents a b th homogeneous branch, o b represents a weight value of the b th homogeneous branch, x n represents an n th sample, θ b represents a network parameter of the b th homogeneous branch, and f e (x n ,θ b ) represents an embedding-layer output feature of x n in the b th homogeneous branch; and

determining the probabilistic collaboration loss value by using a first probabilistic collaboration loss function, wherein the first probabilistic collaboration loss function is:

L

vb

=

(

b

=

1

B

"\[LeftBracketingBar]"

f

e

(

x

n

,

θ

b

)

-

f

re

"\[RightBracketingBar]"

2

)

1

2

/

B

,

where L vb represents the probabilistic collaboration loss value.

Assignments (2)
LICENSE Recorded Jun 30, 2026
From: IEIT SYSTEMS CO., LTD
To: AIVRES SYSTEMS INC.
Reel/Frame 075857/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2023
From: WANG, LI; FAN, BAOYU
To: INSPUR SUZHOU INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 064020/0825 →
Priority Claims (1)
CN 202110728797.7 · Jun 29, 2021 · national
Cited By (1)
US 12,632,920