Object pose tracking from video images
Apparatuses, systems, and techniques to determined a pose of an object from a plurality of images. In at least one embodiment, the pose of an object is determined from at least two images of a video sequence using one or more neural networks, in which the neural network produces a distribution of pose information that is filtered to determine the current pose.
1 . A computer-implemented method comprising:
generating a first distribution of pose information for an object in an image using a trained network from the image;
generating a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;
determining a cuboid having dimensions centered around the object;
applying a first filter to the first distribution to generate updated pose information corresponding to the object;
applying a second filter to the second distribution to generate relative cuboid dimensions; and
identifying a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.
2 . The computer-implemented method of claim 1 , wherein the current pose or a next pose is a 6-degree of freedom pose.
3 . The computer-implemented method of claim 1 , wherein the image and the prior image are two-dimensional images.
4 . The computer-implemented method of claim 1 , wherein the first filter comprises a Bayesian filter, and the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.
5 . The computer-implemented method of claim 1 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.
6 . The computer-implemented method of claim 1 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.
7 . The computer-implemented method of claim 1 , wherein the cuboid comprises a bounding cuboid for the object.
8 . The computer-implemented method of claim 1 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.
9 . The computer-implemented method of claim 1 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.
10 . A system comprising one or more circuits to:
generate a first distribution of pose information for an object in an image using a trained network from the image;
generate a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;
determine a cuboid having dimensions centered around the object;
apply a first filter to the first distribution to generate updated pose information corresponding to the object;
apply a second filter to the second distribution to generate relative cuboid dimensions; and
identify a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.
11 . The system of claim 10 , wherein the current pose or a next pose is a 6-degree of freedom pose.
12 . The system of claim 10 , wherein the image and the prior image are two-dimensional images.
13 . The system of claim 10 , wherein the first filter comprises a Bayesian filter, and wherein the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.
14 . The system of claim 10 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.
15 . The system of claim 10 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.
16 . The system of claim 10 , wherein the cuboid comprises a bounding cuboid for the object.
17 . The system of claim 10 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.
18 . The system of claim 10 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.
19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
generate a first distribution of pose information for an object in an image using a trained network from the image;
generate a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;
determine a cuboid having dimensions centered around the object;
apply a first filter to the first distribution to generate updated pose information corresponding to the object;
apply a second filter to the second distribution to generate relative cuboid dimensions; and
identify a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.
20 . The non-transitory machine-readable medium of claim 19 , wherein the current pose or a next pose is a 6-degree of freedom pose.
21 . The non-transitory machine-readable medium of claim 19 , wherein the image and the prior image are two-dimensional images.
22 . The non-transitory machine-readable medium of claim 19 , wherein the first filter comprises a Bayesian filter, and wherein the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.
23 . The non-transitory machine-readable medium of claim 19 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.
24 . The non-transitory machine-readable medium of claim 19 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.
25 . The non-transitory machine-readable medium of claim 19 , wherein the cuboid comprises a bounding cuboid for the object.
26 . The non-transitory machine-readable medium of claim 19 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.
27 . The non-transitory machine-readable medium of claim 19 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.