Complete 3D object reconstruction from an incomplete image
A modeling system accesses a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object. At least one region of the first object is not represented by pixel values of the 2D input image. The modeling system generates, by applying a 3D representation generation model to the 2D input image, a three-dimensional (3D) representation of the first object that depicts an entirety of the first object including the first region. The modeling system displays, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view.
1 . A method performed by one or more computing devices associated with a modeling system, comprising:
accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image;
training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;
combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;
combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and
displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view.
2 . The method of claim 1 , wherein at the first view the at least one region is outside of an area of the 2D input image.
3 . The method of claim 1 , wherein at the first view the at least one region is occluded by a second object depicted in the 2D input image.
4 . The method of claim 1 , further comprising:
generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP); and
generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object.
5 . The method of claim 4 , wherein the multiview surface normals are based on the coarse geometry.
6 . The method of claim 4 , further comprising:
generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and
determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.
7 . The method of claim 4 , further comprising applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.
8 . The method of claim 1 , wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object, and further comprising:
responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,
wherein the 3D representation displayed at the second view depicts the at least one region of the first object.
9 . A system comprising:
a memory component; and
a processing device coupled to the memory component, the processing device configured to perform operations comprising:
accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image;
training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;
combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;
combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and
displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view,
wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object.
10 . The system of claim 9 , the operations further comprising:
responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,
wherein the 3D representation displayed at the second view depicts the at least one region of the first object.
11 . The system of claim 9 , wherein at the first view the at least one region is outside of an area of the 2D input image or is occluded by a second object depicted in the 2D input image.
12 . The system of claim 9 , the operations further comprising:
generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP); and
generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object.
13 . The system of claim 12 , wherein the multiview surface normals are based on the coarse geometry.
14 . The system of claim 12 , the operations further comprising:
generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and
determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.
15 . The system of claim 12 , the operations further comprising applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.
16 . The system of claim 12 , wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object, the operations further comprising:
responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,
wherein the 3D representation displayed at the second view depicts the at least one region of the first object.
17 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:
accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image, wherein the at least one region is outside of an area of the 2D input image or is occluded by a second object depicted in the 2D input image;
combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;
training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;
combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and
displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view,
wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object.
18 . The non-transitory computer-readable medium of claim 17 , the operations further comprising:
generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP);
generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object; and
applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.
19 . The non-transitory computer-readable medium of claim 18 , the operations further comprising:
generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and
determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.
20 . The system of claim 12 , the operations further comprising:
responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,
wherein the 3D representation displayed at the second view depicts the at least one region of the first object.