Object re-identification using pose part based models
An example apparatus for re-identifying objects includes an image receiver to receive a first image and a second image of an object with an identity. The apparatus also includes a fused model generator to fuse a global representation of the object with local representations of pose parts of the object to generate a fused representation of the object based on the first image. The apparatus further includes an object re-identifier to re-identify the object with the identity in the second image based on the fused representation.
1 . An apparatus comprising:
interface circuitry;
machine-readable instructions; and
at least one processor circuit to be programmed by the machine-readable instructions to:
generate one or more local representations of pose parts of an object based on a first image of the object and a feature map associated with a global representation of the object, the pose parts associated with a skeleton structure of the object in a pose in the first image, each local representation including local part features of the object associated with a respective pose part, the first image captured by a first camera;
aggregate the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features;
generate a fused representation of the object based on the aggregated local features and the global representation of the object; and
re-identify the object in a second image based on the fused representation, the second image captured using a second camera different than the first camera.
2 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to generate the global representation, the global representation including the feature map.
3 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to estimate pose keypoints in the first image and generate the skeleton structure of the object based on the pose keypoints.
4 . The apparatus of claim 1 , wherein the local representations include star structure models.
5 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to extract the local representations from the global representation using regional average pooling.
6 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to execute a deep neural network trained using a fused-triplet loss function to re-identify the object.
7 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to execute a deep neural network to generate the fused representation and re-identify the object.
8 . The apparatus of claim 1 , wherein an identity of the object is defined by an attribute of the object in the first image.
9 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:
output a weight sum vector based on the weighted summation; and
execute a neural network to generate an identity loss vector from the weighted sum vector.
10 . A method comprising:
globally modeling, by at least one processor circuit programmed by at least one instruction, an object based on a first input object image to generate a global representation of the object, the global representation including a feature map, the first input object image captured by a first camera;
estimating, by one or more of the at least one processor circuit, pose keypoints of the object in the first input object image;
generating, by one or more of the at least one processor circuit, a skeleton structure of the object based on the pose keypoints;
modeling, by one or more of the at least one processor circuit, local parts of the object in the first input object image based on the feature map and the skeleton structure to generate one or more local representations of pose parts of the object, the pose parts associated with the skeleton structure of the object in a pose in the first input object image, each local representation including local part features of the object associated with a respective pose part;
aggregating, by one or more of the at least one processor circuit, the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features;
generating, by one or more of the at least one processor circuit, a fused representation of the object based on the aggregated local features and the global representation of the object; and
re-identifying, by one or more of the at least one processor circuit, the object in a second input object image based on the fused representation, the second input object image captured using a second camera different than the first camera.
11 . The method of claim 10 , wherein modeling the local parts includes extracting the local representations from the global representation using regional average pooling.
12 . The method of claim 10 , wherein re-identifying the object includes executing a deep neural network for the second input object image, the deep neural network to output a re-identification of the object.
13 . The method of claim 10 , wherein globally modeling the object includes generating bounding boxes enclosing regions of a first input object image corresponding to different pose parts of the object.
14 . The method of claim 10 , wherein estimating the pose keypoints includes estimating the pose keypoints using a number of pose keypoints based on a category of the object.
15 . The method of claim 14 , further including training one or more deep neural networks to globally model the object, estimate the pose keypoints, model the local parts of the object, and fuse the global representation of the object with the local representations of the object.
16 . The method of claim 10 , wherein fusing the global representation with the local representations includes executing a deep neural network to perform a global transformation on the aggregated local features using a triplet hard loss function.
17 . The method of claim 10 , further including:
outputting a weight sum vector based on the weighted summation; and
executing a neural network to generate an identity loss vector from the weighted sum vector.
18 . A system comprising:
means for generating one or more local representations of pose parts of an object based on a first image of the object and a feature map associated with a global representation of the object, the pose parts associated with a skeleton structure of the object in a pose in the first image, each local representation including local part features of the object associated with a respective pose part, the first image captured by a first camera;
means for fusing to generate a fused representation of the object, the means for fusing to:
aggregate the one or more local representations of the pose parts using a weighted summation of the local part features to generate aggregated local features; and
generate a fused representation of the object based on the aggregated local features and the global representation of the object; and
means for re-identifying the object in a second image based on the fused representation, the second image captured using a second camera different than the first camera.
19 . The system of claim 18 , further including means for generating the global representation, the global representation including the feature map.
20 . The system of claim 18 , further including means for estimating pose keypoints in the first image to generate the skeleton structure of the object.
21 . The system of claim 18 , wherein the local representations include star structure models.
22 . The system of claim 18 , wherein the means for fusing is to:
output a weight sum vector based on the weighted summation; and
execute a neural network to generate an identity loss vector from the weighted sum vector.