IP Library › Granted Patent US 12,682,568
Granted Patent B2
US 12,682,568 · App. 18/339,936 · Granted Jul 14, 2026

Generating complete three-dimensional scene geometries using machine learning

Inventors: Dongsu Zhang (Seoul, KR); Amlan Kar (Toronto, CA); Francis Williams (Brooklyn, NY); Zan Gojcic (Zurich, CH); Karsten Kreis (Vancouver, CA); Sanja Fidler (Toronto, CA)
Assignee: NVIDIA CORPORATION
G06T17/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,568
App. No.
18/339,936
Filed
Jun 22, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2616
USPC
345/418
Abstract

In various examples, a technique for performing three-dimensional (3D) scene completion includes determining an initial representation of a first 3D scene. The technique also includes executing a machine learning model to generate a first update to the initial representation at a previous time step and a second update to the initial representation at a current time step, wherein the second update is generated based at least on a threshold applied to a set of predictions corresponding to the first update. The technique also includes generating a 3D model of the 3D scene based at least on the second update to the initial representation.

Claims (68)

1 . A method, comprising:

determining an initial representation of a three-dimensional (3D) scene including a first set of cells;

computing a transition kernel for a second set of cells within a neighborhood of at least one cell included in the first set of cells;

generating, using a machine learning model, a first update to the initial representation at a previous time step and a second update to the initial representation at a current time step, the second update generated based on sampling the transition kernel and on at least on a set of predictions associated with the first set of cells; and

generating a 3D model of the 3D scene based at least on the second update to the initial representation.

2 . The method of claim 1 , wherein the machine learning model was trained using a loss that is computed based at least on a set of visible cells included in one or more training 3D scenes.

3 . The method of claim 1 , wherein the machine learning model was trained to generate a target geometry derived from a first set of scans of a training 3D scene based at least on an input geometry derived from a subset of the first set of scans of the training 3D scene.

4 . The method of claim 1 , wherein the machine learning model was trained to generate a series of intermediate states between a first representation of a training 3D scene and a second representation of the training 3D scene.

5 . The method of claim 1 , wherein generating the first update and the second update comprises:

applying a threshold to the set of predictions to determine the first set of cells corresponding to the first update; and

generating, via the machine learning model based at least on the first set of cells, the second update that includes a second set of predictions for the second set of cells located within a neighborhood of the first set of cells.

6 . The method of claim 5 , wherein generating the second set of predictions for the second set of cells comprises aggregating a plurality of probabilities associated with a cell included in the second set of cells into an overall probability associated with the cell.

7 . The method of claim 1 , wherein determining the initial representation comprises converting one or more scans of the 3D scene into the initial representation.

8 . The method of claim 1 , wherein the initial representation comprises a set of voxel occupancies.

9 . The method of claim 1 , wherein the machine learning model comprises a sparse convolutional neural network.

10 . A processor comprising:

one or more circuits to perform operations comprising:

determining an initial representation of a three-dimensional (3D) scene including a first set of cells;

computing a transition kernel for a second set of cells within a neighborhood of at least one cell included in the first set of cells;

generating, using a machine learning model, a first update to the initial representation at a previous time step and a second update to the initial representation at a current time step, the second update generated based on sampling the transition kernel and on at least on a set of predictions associated with the first set of cells; and

generating a 3D model of the 3D scene based at least on the second update to the initial representation.

11 . The processor of claim 10 , wherein the machine learning model was trained to generate a target geometry derived from a first set of scans of a second 3D shape based at least on an input geometry derived from a second set of scans of the second 3D shape, and wherein the second set of scans corresponds to a subset of the first set of scans.

12 . The processor of claim 10 , wherein the machine learning model was further trained using at least one of:

a binary cross-entropy loss associated with a set of visible cells included in one or more scans of a set of training 3D shapes; or

one or more sets of simulated scans of a set of synthetic 3D shapes.

13 . The processor of claim 10 , wherein the operations further comprise generating a simulation based at least on the model.

14 . The processor of claim 13 , wherein the simulation comprises at least one of a collision simulation, a driving simulation, or a structural load simulation.

15 . The processor of claim 10 , wherein generating the model comprises:

determining one or more 3D cells associated with the second update that are located between a sensor used to capture the initial representation and a set of 3D cells included in the initial representation; and

omitting the one or more 3D cells from the model.

16 . The processor of claim 10 , wherein the processor is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content;

a system implemented using a robot;

a system for performing conversational Al operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

17 . A system comprising:

one or more processing units to execute operations comprising:

generating an input geometry based at least on a first set of scans of a shape;

generating a target geometry based at least on a second set of scans of the shape, wherein the first set of scans corresponds to a subset of the second set of scans;

executing a machine learning model to generate a first update to the input geometry at a previous time step and a second update to the input geometry at a current time step, including, for at least one cell included in the first set of cells, computing a transition kernel for a third set of cells within a neighborhood of the cell, wherein the second update is generated based at least on a set of predictions associated with a first set of cells and corresponding to the first update and a set of probabilities indicating occupancy within a 3D scene for a second set of cells located within a neighborhood of the first set of set of cells, and including sampling from the transition kernel to generate the second update; and

updating the machine learning model using one or more losses that are computed based at least on a set of visible cells associated with at least one of the first update or the second update.

18 . The system of claim 17 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content;

a system implemented using a robot;

a system for performing conversational Al operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: ZHANG, DONGSU; KAR, AMLAN; WILLIAMS, FRANCIS; GOJCIC, ZAN; KREIS, KARSTEN; FIDLER, SANJA
To: NVIDIA CORPORATION
Reel/Frame 064048/0348 →
Continuity (2)
Provisional Application 63429275 · Dec 1, 2022
Related Publication 20240185523A1 · Jun 6, 2024
References Cited (64)
US 9235920B2 · Girdzijauskas · 2016 [cited by examiner]
US 11741643B2 · Kim · 2023 [cited by examiner]
US 20200033880A1 · Kehl · 2020 [cited by examiner]
US 20230281913A1 · Rematas · 2023 [cited by examiner]
Dai, Angela, et al. “Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. (Year: 2018). [cited by examiner]
Chen, Yukang, et al. “Focal sparse convolutional networks for 3d object detection.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Wu, Shun-Cheng, et al. “Scfusion: Real-time incremental scene reconstruction with semantic completion.” 2020 International Conference on 3D Vision (3DV). IEEE, 2020. (Year: 2020). [cited by examiner]
Arora et al., “Multimodal Shape Completion via IMLE”, arXiv.2106.16237, Jul. 7, 2021, 16 pages. [cited by applicant]
Bautista et al., “GAUDI: A Neural Architect for Immersive 3D Scene Generation”, arXiv.2207.13751, Jul. 27, 2022, 21 pages. [cited by applicant]
Behley et al., “SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), DOI: 10.1109/ICCV.2019.00939, 2019, pp. 9… [cited by applicant]
Bordes et al., “Learning to Generate Samples from Noise through Infusion Training”, In ICLR, 2017, 4 pages. [cited by applicant]
Brock et al., “Large Scale Gan Training for High Fidelity Natural Image Synthesis”, arXiv.1809.11096, Sep. 28, 2018, 29 pages. [cited by applicant]
Cai et al., “Learning Gradient Fields for Shape Generation”, arXiv:2008.06520, Aug. 18, 2020, 33 pages. [cited by applicant]
Chabra et al., “Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction”, arXiv:2003.10983, Aug. 21, 2020, 26 pages. [cited by applicant]
Chang et al., “ShapeNet: An Information-Rich 3D Model Repository”, arXiv:1512.03012, Dec. 9, 2015, 11 pages. [cited by applicant]
Chen et al., “Learning Implicit Fields for Generative Shape Modeling”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR.2019.00609, 2019, pp. 5932-5941. [cited by applicant]
Cheng et al., “S3CNet: A Sparse Semantic Scene Completion Network for LiDAR Point Clouds”, In 4th Conference on Robot Learning (CoRL 2020), 2020, 14 pages. [cited by applicant]
Dai et al., “SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR4… [cited by applicant]
Dai et al., “Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, DOI: 10.1109/CVPR.2017.693, 2017, pp. 6545-6554. [cited by applicant]
Dai et al., “ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, DOI: 10.1109/CVPR.2018.00481, 2018, pp. … [cited by applicant]
Gao et al., “GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images”, In 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2209.11163, Sep. 22, 2022, 39 pages. [cited by applicant]
Graham et al., “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, DOI: 10.1109/CVPR.2018.00961, 2018, pp. 922… [cited by applicant]
Gu et al., “Weakly-supervised 3D Shape Completion in the Wild”, arXiv:2008.09110, Aug. 20, 2020, 28 pages. [cited by applicant]
Henzler et al., “Escaping Plato's Cave: 3D Shape From Adversarial Rendering”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), DOI: 10.1109/ICCV.2019.01008, 2019, pp. 9983-9992. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, In 34th Conference on Neural Information Processing Systems (NeurIPS 2020), arXiv:2006.11239, Dec. 16, 2020, 25 pages. [cited by applicant]
Jiang et al., “Local Implicit Grid Representations for 3D Scenes”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR42600.2020.00604, 2020, pp. 6000-6009. [cited by applicant]
Kar et al., “Meta-Sim: Learning to Generate Synthetic Datasets”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), DOI: 10.1109/ICCV.2019.00465, 2019, pp. 4550-4559. [cited by applicant]
Kim et al., “DriveGAN: Towards a Controllable High-Quality Neural Simulation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR46437.2021.00576, 2021, pp. 58… [cited by applicant]
Manivasagam et al., “LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World”, arXiv:2006.09348, Jun. 16, 2020, 11 pages. [cited by applicant]
Mescheder et al., “Occupancy Networks: Learning 3D Reconstruction in Function Space”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, DOI: 10.1109/CVPR.2019.00459, 2019, pp. 4455-44… [cited by applicant]
Michalkiewicz et al., “Deep Level Sets: Implicit Surface Representations for 3D Shape Inference” arXiv:1901.06802, Jan. 21, 2019, 10 pages. [cited by applicant]
Mordvintsev et al., “Growing Neural Cellular Automata”, Distill, DOI: 10.23915/distill.00023, Feb. 11, 2020, 23 pages. [cited by applicant]
Newcombe et al., “KinectFusion: Real-Time Dense Surface Mapping and Tracking”, In Proceedings of the 10th IEEE International Symposium on Mixed and Augmented Reality, DOI: 10.1109/ISMAR.2011.6092378, Oct. 26-29, 2011, p… [cited by applicant]
Park et al., “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR.2019.00025, … [cited by applicant]
Rist et al., “Semantic Scene Completion using Local Deep Implicit Functions on LiDAR Data”, arXiv:2011.09141, Apr. 12, 2021, 19 pages. [cited by applicant]
Ruiz et al., “Learning to Simulate”, In ICLR, arXiv:1810.02513, Oct. 5, 2018, 12 pages. [cited by applicant]
Sohl-Dickstein et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics”, In Proceedings of the 32nd International Conference on International Conference on Machine Learning, vol. 37, arXiv:1503.03585, Ju… [cited by applicant]
Song et al., “Semantic Scene Completion from a Single Depth Image”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR.2017.28, 2017, pp. 190-198. [cited by applicant]
Peng et al., “Convolutional Occupancy Networks”, arXiv:2003.04618, Aug. 1, 2020, 17 pages. [cited by applicant]
Sun et al., “Scalability in Perception for Autonomous Driving: Waymo Open Dataset”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR42600.2020.00252, 2020, p… [cited by applicant]
Tancik et al., “Block-NeRF: Scalable Large Scene Neural View Synthesis”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR52688.2022.00807, 2022, pp. 8248-825… [cited by applicant]
Vizzo et al., “Make it Dense: Self-Supervised Geometric Scan Completion of Sparse 3D LiDAR Scans in Large Outdoor Environments”, In IEEE Robotics and Automation Letters, vol. 7, No. 3, Jul. 2022, pp. 8534-8541. [cited by applicant]
Neumann, John Von, “Theory of Self-Reproducing Automata”, 1966, 403 pages. [cited by applicant]
Wu et al., “Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling”, In 29th Conference on Neural Information Processing Systems (NIPS 2016), arXiv:1610.07584, Oct. 24, 2016, 11 pa… [cited by applicant]
Wu et al., “Multimodal Shape Completion via Conditional Generative Adversarial Networks”, arXiv:2003.07717, Jul. 8, 2020, 22 pages. [cited by applicant]
Yan et al., “Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion”, In the Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), 2021, pp. 3101-3109. [cited by applicant]
Yang et al., “PointFlow: 3D Point Cloud Generation with Continuous Normalizing Flows”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), DOI: 10.1109/ICCV.2019.00464, 2019, pp. 4540-4549. [cited by applicant]
Yang et al., “SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR42600.2020.01113, 2020, … [cited by applicant]
Yuan et al., “PCN: Point Completion Network”, In International Conference on 3D Vision (3DV), DOI: 10.1109/3DV.2018.00088, 2018, pp. 728-737. [cited by applicant]
Zeng et al., “LION: Latent Point Diffusion Models for 3D Shape Generation”, In 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2210.06978, Oct. 12, 2022, 63 pages. [cited by applicant]
Zhang et al., “Learning to Generate 3D Shapes With Generative Cellular Automata”, in ICLR 2021, arXiv:2103.04130, Mar. 6, 2021, 22 pages. [cited by applicant]
Zhang et al., “Probabilistic Implicit Scene Completion”, In ICLR 2022, arXiv:2204.01264, Apr. 4, 2022, 32 pages. [cited by applicant]
Zhou et al., “3D Shape Generation and Completion through Point-Voxel Diffusion”, In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), DOI: 10.1109/ICCV48922.2021.00577, 2021, pp. 5806-5815. [cited by applicant]
Choy et al., “4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109/CVPR.2019.00319, 2019, pp. 3… [cited by applicant]
Cuturi, Marco, “Sinkhorn Distances: Lightspeed Computation of Optimal Transport”, In Proceedings of the 26th International Conference on Neural Information Processing Systems, 2013, 9 pages. [cited by applicant]
Fan et al., “A Point Set Generation Network for 3D Object Reconstruction from a Single Image”, arXiv:1612.00603, Dec. 7, 2016, 12 pages. [cited by applicant]
Insafutdinov et al., “Unsupervised Learning of Shape and Pose with Differentiable Point Clouds”, In 32nd Conference on Neural Information Processing Systems (NIPS 2018), arXiv:1810.09381, Oct. 22, 2018, 16 pages. [cited by applicant]
Kingma et al., “Adam: A Method For Stochastic Optimization”, In ICLR 2015, arXiv:1412.6980, Jul. 23, 2015, 15 pages. [cited by applicant]
Liu et al., “Morphing and Sampling Network for Dense Point Cloud Completion”, arXiv:1912.00280, Nov. 30, 2019, 9 pages. [cited by applicant]
Lorensen et al., “Marching Cubes: A High Resolution 3D Surface Construction Algorithm”, ACM SIGGRAPH Computer Graphics, vol. 21, No. 4, DOI: 10.1145/37402.37422, Jul. 1987, pp. 163-169. [cited by applicant]
Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, In 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), arXiv:1912.01703, Dec. 3, 2019, 12 pages. [cited by applicant]
Qi et al., “Frustum PointNets for 3D Object Detection from RGB-D Data”, arXiv:1711.08488, Nov. 22, 2017, 15 pages. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv:1505.04597, May 18, 2015, 8 pages. [cited by applicant]
Sun et al., “Scalability in Perception for Autonomous Driving: An Open Dataset Benchmark”, arXiv:1912.04838, Dec. 10, 2019, 9 pages. [cited by applicant]