IP Library Granted Patent US 12,475,604
Granted Patent B2
US 12,475,604 · App. 17/006,702 · Granted Nov 18, 2025

Object image completion

Inventors: Sanja Fidler (Toronto, CA); David Acuna Marrero (Toronto, CA); Seung Wook Kim (Toronto, CA); Karsten Julian Kreis (Vancouver, CA); Huan Ling (Toronto, CA)
Assignee: NVIDIA Corporation
G06T9/002G06F18/214G06N3/08G06T7/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,604
App. No.
17/006,702
Granted
Nov 18, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to generate complete depictions of objects based on a partial depiction of the object. In at least one embodiment, an image of a complete object is generated by one or more neural networks, based on an image of a portion of the object, using an encoder of the one or more neural networks trained using training data generated from output of a decoder of the one or more neural networks.

Claims (50)

1 . One or more processors, comprising:

circuitry to train one or more neural networks to generate an image of an entire object based at least on an image of a portion of the entire object, wherein to train the one or more neural networks the circuitry at least:

causes a decoder portion of the one or more neural networks to generate a first image of a complete first object; and

generates a second image in which the complete first object is occluded by a second object.

2 . The one or more processors of claim 1 , wherein the one or more neural networks comprises a generative model framework to generate a plurality of complete images based on the second image.

3 . The one or more processors of claim 1 , wherein the decoder portion is trained based, at least in part, on a dataset comprising images of complete objects and excluding images of portions of objects.

4 . The one or more processors of claim 1 , wherein parameters of the decoder portion are not adjusted while training an encoder using training data.

5 . The one or more processors of claim 4 , wherein output of the trained encoder causes the decoder portion to generate a plurality of images of a complete object based on input, to the one or more neural networks, of the second image of a portion of the complete first object.

6 . The one or more processors of claim 1 , the circuitry to generate training data by generating the first image of the complete first object and the second image of a complete second object, and combining the first and second images so that the complete first object is at least partially occluded by the complete second object.

7 . The one or more processors of claim 1 , the circuitry to refine a training of an encoder based, at least in part, on a plurality of real-images of portions of objects after training the encoder with generated training data.

8 . The one or more processors of claim 1 , wherein one or more of the one or more neural networks is trained to spatially transform output of the decoder portion.

9 . The one or more processors of claim 1 , wherein a portion of the complete first object comprises a not occluded portion of a partially occluded object.

10 . A system, comprising:

one or more processors to train one or more neural networks to generate an image of an entire object based at least on an image of a portion of the entire object, wherein to train the one or more neural networks the one or more processors at least:

cause a decoder portion of the one or more neural networks to generate a first image of a complete first object; and

generate a second image in which the complete first object is occluded by a second object.

11 . The system of claim 10 , wherein the one or more neural networks comprise a generative model framework to generate a plurality of complete images based on the second image, wherein the generative model framework comprises at least one of a variational autoencoder, a generative adversarial network, or a normalizing flow.

12 . The system of claim 10 , the one or more processors to train the decoder portion, using a dataset comprising images of complete objects, to generate variations of images of complete objects.

13 . The system of claim 10 , wherein parameters of the decoder portion are frozen while training an encoder using training data generated based, at least in part, on output of the decoder portion.

14 . The system of claim 13 , wherein an output of the trained encoder causes the decoder portion to generate a plurality of images of a complete object and a plurality of probabilities corresponding to the plurality of images, based on input, to the one or more neural networks, of the second image of a portion of the complete first object.

15 . The system of claim 10 , the one or more processors to generate training data by generating the first image of the complete first object and the second image of a complete second object, and combining the first and second images so that the complete first object is at least partially occluded by the complete second object.

16 . The system of claim 10 , the one or more processors to train one or more of the one or more neural networks to spatially transform output of the decoder portion.

17 . The system of claim 10 , wherein a portion of the complete first object comprises a not occluded portion of a partially occluded object.

18 . A method, comprising:

causing a decoder portion of one or more neural networks to generate a first image of a complete first object;

generating a second image in which the complete first object is occluded by a second object; and

training the one or more neural networks to generate an image of an entire object from an image of a portion of the entire object, the training based at least on the generated second image.

19 . The method of claim 18 , wherein the one or more neural networks comprise at least one of a variational autoencoder, generative adversarial network, or normalizing flow.

20 . The method of claim 18 further comprising:

training the decoder portion, using a dataset comprising images of complete objects, to generate variations of images of complete objects.

21 . The method of claim 18 , further comprising:

freezing parameters of the decoder portion while training an encoder using training data generated based, at least in part, on output of the decoder portion.

22 . The method of claim 21 , wherein an output of the trained encoder causes the decoder portion to generate a plurality of images of a complete object and a plurality of probabilities corresponding to the plurality of images, based on input, to the one or more neural networks, of the second image of a portion of the complete first object.

23 . The method of claim 18 , further comprising:

generating training data by generating the first image of the complete first object and the second image of a complete second object, and combining the first and second images so that the complete first object is at least partially occluded by the complete second object.

24 . The method of claim 18 , further comprising:

training one or more of the one or more neural networks to spatially transform output of the decoder portion.

25 . The method of claim 18 , wherein a portion of the complete first object comprises a not occluded portion of a partially occluded object.

26 . A non-transitory machine-readable medium having stored thereon instructions, which if performed by one or more processors, cause the one or more processors to at least:

cause a decoder portion of one or more neural networks to generate a first image of a complete first object;

generate a second image in which the complete first object is occluded by a second object; and

train the one or more neural networks to generate an image of an entire object from an image of a portion of the entire object, the training based at least on the generated second image.

27 . The non-transitory machine-readable medium of claim 26 , having stored thereon further instructions, which if performed by one or more processors, cause the one or more processors to at least:

generate a plurality of variations of a complete object, based at least in part on the second image of a portion of the complete first object.

28 . The non-transitory machine-readable medium of claim 26 , wherein parameters of the decoder portion are frozen after training the decoder portion using images of complete objects.

29 . The non-transitory machine-readable medium of claim 26 , wherein training data comprises an image depicting a plurality of objects generated based on an output of the decoder portion.

30 . The non-transitory machine-readable medium of claim 29 , wherein the complete first object overlaps the second object.

31 . The non-transitory machine-readable medium of claim 26 , wherein the second image is altered by replacing a depiction of a portion of the complete first object in the second image with a depiction of a complete object.

32 . The non-transitory machine-readable medium of claim 26 , wherein the second image is altered by removing the second object in the second image that occludes a portion of the complete first object.

33 . The non-transitory machine-readable medium of claim 26 , wherein a portion of the complete first object comprises a not occluded portion of a partially occluded object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2020
From: FIDLER, SANJA; ACUNA MARRERO, DAVID; KIM, SEUNG WOOK; KREIS, KARSTEN JULIAN; LING, HUAN
To: NVIDIA CORPORATION
Reel/Frame 054043/0918 →
Continuity (1)
Related Publication 20220067983A1 · Mar 3, 2022
References Cited (60)
US 10176388B1 · Ghafarianzadeh · 2019 [cited by examiner]
US 20190325273A1 · Kumar · 2019 [cited by examiner]
US 20200045289A1 · Raziel et al. · 2020 [cited by applicant]
US 20200134499A1 · Ryu · 2020 [cited by examiner]
US 20210125380A1 · Lee · 2021 [cited by examiner]
US 20210397966A1 · Sun · 2021 [cited by examiner]
CN 110310242A · 2019 [cited by applicant]
KR 20190049552A · 2019 [cited by applicant]
WO 2020101246A1 · 2020 [cited by applicant]
Acuna et al., “Efficient Interactive Annotation of Segmentation Datasets with Polygon-RNN++,” Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Alemi et al., “Deep Variational Information Bottleneck,” In The International Conference on Learning Representations, 2017, 19 pages. [cited by applicant]
Alemi et al., “Fixing a Broken ELBO,” International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Burgess et al., Understanding disentangling in ß-VAE, Apr. 10, 2018, 11 pages. [cited by applicant]
Chang et al., “ShapeNet: An Information-Rich 3D Model Repository,” Dec. 9, 2015, 11 pages. [cited by applicant]
Chen et al., “Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFS,” IEEE transactions on pattern analysis and machine intelligence, 40(4): May 12, 2017, 14 pag… [cited by applicant]
Cordts et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding,” IEEE Conference on ComputerVision and Pattern Recognition, 2016, 11 pages. [cited by applicant]
Ehsani et al., “SeGAN: Segmenting and Generating the Invisible,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Fei-Fei et al., “What do we perceive in a glance of a real-world scene?” Journal of Vision, 7(1): 2007, 29 pages. [cited by applicant]
Geiger et al., “Vision Meets Robotics: The KITTI Dataset,” The International Journal of Robotics Research, 32(11): 2013, 6 pages. [cited by applicant]
He et al., “Mask R-CNN,” ICCV, 2017, 9 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” Nov. 8, 2017, 38 pages. [cited by applicant]
Higgins et al., “Beta-vae: Learning Basic Visual Concepts with a Constrained Variational Framework,” ICLR, 2(5):6, 2017, 13 pages. [cited by applicant]
Hu et al., “SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation—A Synthetic Dataset and Baselines,” CVPR, 2019. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Jaderberg et al., “Spatial Transformer Networks,” In Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Kar et al., “Amodal Completion and Size Constancy in Natural Scenes,” Proceedings of the IEEE International Conference on Computer Vision, 2015, 9 pages. [cited by applicant]
Kimia et al., “Euler Spiral for Shape Completion,” International Journal of Computer Vision 54, 2003, 24 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” International Conference on Learning Representations, Apr. 14, 2014, 14 pages. [cited by applicant]
Kingma et al., “Semi-Supervised Learning with Deep Generative Models,” In Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Li et al., “Amodal Instance Segmentation,” European Conference on Computer Vision, 2016, 23 pages. [cited by applicant]
Lin et al., “A Computational Model of Topological and Geometric Recovery for Visual Curve Completion,” Computational Visual Media, 2(4): Dec. 2016, 14 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Liu et al., “Image Inpainting for Irregular Holes using Partial Convolutions,” Proceedings of the European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]
Park et al., “Semantic Image Synthesis with Spatially Adaptive Normalization,” CVPR, 2019, 10 pages. [cited by applicant]
Qi et al., “Amodal Instance Segmentation with KINS Dataset,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” In International Conference on Machine Learning, 2014, 9 pages. [cited by applicant]
Silberman et al., “A Contour Completion Model for Augmenting Surface Reconstructions,” European Conference on Computer Vision, 2014, 15 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Sohn et al., “Learning Structured Output Representation using Deep Conditional Generative Models,” Advances in Neural Information Processing Systems 28, 2015, 9pages. [cited by applicant]
Stutz et al., “Learning 3D Shape Completion under Weak Supervision,” International Journal of Computer Vision, 2018, 20 pages. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” CVPR, 2016, 9 pages. [cited by applicant]
Takikawa et al., “Gated-SCNN: Gated Shape CNNs for Semantic Segmentation,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Tsai et al., “Deep Image Harmonization,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Xiang et al., “Beyond PASCAL: A Benchmark for 3D Object Detection in the Wild,” IEEE Winter Conference on Applications of Computer Vision, 2014, 8 pages. [cited by applicant]
Yan et al., “Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Yi et al., “GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Zhan et al., “Self-Supervised Scene De-occlusion,” Apr. 6, 2020, 11 pages. [cited by applicant]
Zhao et al., “InfoVAE: Information Maximizing Variational Autoencoders,” Nov. 17, 2017, 9 pages. [cited by applicant]
Zhou et al., “Scene Parsing through ADE20K Dataset,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Zhu et al., “Semantic Amodal Segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Zhu et al., “Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning,” IEEE International Conference on Robotics and Automation, 2016, 8pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Patent Application No. 2112265.0 dated Jan. 24, 2022, 3 pages. [cited by applicant]
United Kingdom Search Report for Application No. GB2300959.0, mailed Aug. 23, 2023, 4 pages. [cited by applicant]
Yu et al., “Generative Image Inpainting with Contextual Attention,” CVPR, 2018, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 202111006843.9, mailed Aug. 21, 2024, 20 pages. [cited by applicant]
Office Action for Chinese Application No. 202111006843.9, mailed Jun. 4, 2025, 17 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202111006843.9, mailed Jul. 31, 2025, 22 pages. [cited by applicant]
Liu et al., “Deep Practice OCR: Text Recognition Based on Deep Learning,” China Machine Press, Apr. 2020, 6 pages. [cited by applicant]
Cited By (1)
US 12,718,595