Systems and methods for a shape completion model
Systems and methods for shape completion are provided. In one embodiment, a computer implemented method includes receiving sensor data for a visualized area of an object as at least one point cloud representation. The computer implemented method also includes transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object. The input voxel grid is a volumetric representation. The computer implemented method further includes encoding the input voxel grid into a partial latent vector that lies on a partial latent space. The computer implemented method yet further includes determining a mapping between the partial latent space and a complete latent space based on the sensor data. The computer implemented method includes predicting a complete latent vector based on the complete latent space. The computer implemented method also includes estimating a complete shape of an object based on the complete latent space.
1 . A system for shape completion, comprising:
a processor, and
a memory storing instructions that when executed by the processor cause the processor to:
receive sensor data for a visualized area of an object as at least one point cloud representation;
transform the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;
encode the input voxel grid into a partial latent vector that lies on a partial latent space;
extract visual features from the sensor data;
input the partial latent vector, the visual features, and a Gaussian latent code into a conditional generative model;
generate, by the conditional generative model, a complete latent vector in a complete latent space; and
decode, by a decoder of an autoencoder and using the generated complete latent vector as input to the decoder, a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.
2 . The system of claim 1 , wherein the mapping is further based on visual features extracted from the sensor data.
3 . The system of claim 2 , wherein the system of claim 1 includes an autoencoder having a generator, and wherein visual features of the sensor data are input into the generator as conditional input, the generator being part of the conditional generative model.
4 . The system of claim 1 , wherein the system of claim 1 includes an autoencoder having an encoder to encode the input voxel grid and a decoder to estimate the complete shape, wherein the encoder generates the partial latent vector and the decoder reconstructs the voxel representation of the object.
5 . The system of claim 4 , wherein the autoencoder is optimized by minimizing Jaccard index loss between reconstructed voxel grids and ground truth voxel grids.
6 . The system of claim 1 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.
7 . The system of claim 1 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.
8 . A computer implemented method for shape completion, comprising:
receiving sensor data from an agent for a visualized area of an object as at least one point cloud representation;
transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;
encoding the input voxel grid into a partial latent vector that lies on a partial latent space;
determining a mapping between the partial latent space and a complete latent space based on the sensor data;
generating a complete latent vector representing the object based on the mapping; and
decoding the complete latent vector to reconstruct a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.
9 . The computer implemented method of claim 8 , wherein the mapping is based on visual features extracted from the sensor data.
10 . The computer implemented method of claim 8 , further comprising extracting visual features from the sensor data, wherein the generating the complete latent vector is further based on the extracted visual features as conditional input.
11 . The computer implemented method of claim 8 , further comprising performing optimization by minimizing Jaccard index loss.
12 . The computer implemented method of claim 8 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.
13 . The computer implemented method of claim 8 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.
14 . The computer implemented method of claim 8 , wherein the voxel representation is used for object shape completion in a scenario including one or more of self-occluded object shape completion, in-hand object shape completion, and cluttered object shape completion by the agent.
15 . A non-transitory computer readable storage medium storing instructions that when executed by a computer having a processor to perform a method for shape completion, the method comprising:
receiving sensor data for a visualized area of an object as at least one point cloud representation;
transforming the at least one point cloud representation into an input voxel grid of the visualized area of the object, wherein the input voxel grid is a volumetric representation;
encoding the input voxel grid into a partial latent vector that lies on a partial latent space;
determining a mapping between the partial latent space and a complete latent space based on the sensor data;
generating a complete latent vector representing the object based on the mapping; and
decoding the complete latent vector to reconstruct a voxel representation of the object, wherein the reconstructed object includes the visualized area of the object and a previously occluded portion of the object.
16 . The non-transitory computer readable storage medium of claim 15 , wherein the mapping is based on visual features extracted from the sensor data.
17 . The non-transitory computer readable storage medium of claim 15 , the method further comprising extracting visual features from the sensor data, wherein the generating the complete latent vector is further based on the visual features as conditional input.
18 . The non-transitory computer readable storage medium of claim 15 , the method further comprising performing optimization by minimizing Jaccard index loss.
19 . The non-transitory computer readable storage medium of claim 15 , wherein the input voxel grid is encoded into the partial latent vector using a set of three-dimensional convolutional layers.
20 . The non-transitory computer readable storage medium of claim 15 , wherein the at least one point cloud representation is normalized based on a centroid of the at least one point cloud representation and a farthest distance of the at least one point cloud representation from the centroid.