IP Library Granted Patent US 12,299,962
Granted Patent B2
US 12,299,962 · App. 17/959,915 · Granted May 13, 2025

Diffusion-based generative modeling for synthetic data generation systems and applications

Inventors: Karsten Kreis (Vancouver, CA); Tim Dockhorn (Waterloo, CA); Arash Vahdat (Mountain View, CA)
Assignee: Nvidia Corporation
G06V10/774G06N3/045G06N3/047G06T7/277G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,962
App. No.
17/959,915
Granted
May 13, 2025
Kind
B2
Abstract

Systems and methods described relate to the synthesis of content using generative models. In at least one embodiment, a score-based generative model can use a stochastic differential equation with critically-damped Langevin diffusion to learn to synthesize content. During a forward diffusion process, noise can be introduced into a set of auxiliary (e.g., “velocity”) values for an input image to learn a score function. This score function can be used with the stochastic differential equation during a reverse diffusion denoising process to remove noise from the image to generate a reconstructed version of the input image. A score matching objective for the critically-damped Langevin diffusion process can require only the conditional distribution learned from the velocity data. A stochastic differential equation based integrator can then allow for efficient sampling from these critically-damped Langevin diffusion models.

Claims (78)

1. A computer-implemented method, comprising:

providing an input image to a generative neural network, the input image including a first representation of an object;

determining a set of velocity values coupled to a set of pixel values of the input image;

introducing noise values to the set of velocity values for the image to obtain a noise image, the noise values being introduced iteratively during a forward diffusion process;

removing one or more of the noise values from the noise image to obtain a reconstructed image including a second representation of the object, the noise values being removed iteratively during a reverse denoising diffusion process; and

adjusting network parameters for the generative neural network based at least on one or more differences between at least the input image and the reconstructed image.

2. The computer-implemented method of claim 1 , further comprising:

determining the one or more noise values to remove from the noise image according to a score function learned during the forward diffusion process.

3. The computer-implemented method of claim 2 , further comprising:

determining the one or more noise values to remove according to a stochastic differential equation including a term corresponding to the score function for the set of velocity values.

4. The computer-implemented method of claim 3 , further comprising:

determining the one or more noise values to remove using a numerical solver to simulate the stochastic differential equation.

5. The computer-implemented method of claim 3 , further comprising:

performing hybrid score matching (HSM) to determine the score function, the hybrid score matching including diffusion score matching and denoising score matching.

6. The computer-implemented method of claim 1 , further comprising:

mapping one or more pixel values of the input image to a multi-dimensional space including both the one or more pixel values and one or more velocity values of the set of velocity values corresponding to the input image.

7. The computer-implemented method of claim 1 , further comprising:

introducing the noise values at least by perturbing the input image during the forward diffusion process into a tractable distribution.

8. The computer-implemented method of claim 1 , further comprising:

learning a score function of a conditional distribution of the set of velocity values during the forward diffusion process using critically-damped Langevin diffusion.

9. The computer-implemented method of claim 1 , wherein the set of velocity values is coupled to the set of pixel values using Hamiltonian dynamics.

10. The computer-implemented method of claim 1 , further comprising:

calculating the set of velocity values as first-order time derivatives of the set of pixel values.

11. A processor, comprising:

one or more circuits to cause the processor to perform operations comprising:

providing input to a generative neural network;

determining a set of auxiliary values corresponding to a set of data values of the input;

introducing noise values to the set of auxiliary values corresponding to the input to obtain noise data, the one or more noise values being introduced iteratively during a forward diffusion process;

removing the one or more noise values from the auxiliary values to obtain a reconstructed input, the one or more noise values being removed iteratively during a reverse denoising diffusion process; and

adjusting network parameters for the generative neural network based at least on differences between at least the input and the reconstructed input.

12. The processor of claim 11 , wherein the one or more circuits are to perform operations further comprising:

determining the one or more noise values to remove from the auxiliary values according to a stochastic differential equation including a term corresponding to a score function for the set of auxiliary values.

13. The processor of claim 12 , wherein the one or more circuits are to perform operations further comprising:

determining the one or more noise values to remove using a numerical solver to simulate the stochastic differential equation.

14. The processor of claim 11 , wherein the one or more circuits are to perform operations further comprising:

mapping one or more data values corresponding to the input to a multi-dimensional space including the one or more data values and one or more auxiliary values of the set of auxiliary values corresponding to the input.

15. The processor of claim 11 , wherein the one or more auxiliary values include one or more velocity values, and wherein one or more circuits are to perform operations further comprising:

Learning a score function of a conditional distribution corresponding to the one or more velocity values during the diffusion process using critically-damped Langevin diffusion.

16. The processor of claim 11 , wherein the processor is comprised in at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

17. A system, comprising:

one or more processing units to train a score-based generative neural network at least by adding noise during a forward diffusion process to a set of velocity variables corresponding to a set of pixel values of an input image, and removing the noise from the set of velocity variables during a reverse diffusion denoising process to generate a reconstruction of the input image.

18. The system of claim 17 , wherein the one or more processing units are further to:

determine noise values corresponding to the noise to remove according to a stochastic differential equation including a term corresponding to a score function for the set of velocity variables.

19. The system of claim 18 , wherein the one or more processing units are further to:

learn the score function of a conditional distribution of the set of velocity variables during the forward diffusion process using critically-damped Langevin diffusion.

20. The system of claim 17 , wherein the system comprises at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: KREIS, KARSTEN; DOCKHORN, TIM; VAHDAT, ARASH
To: NVIDIA CORPORATION
Reel/Frame 061662/0316 →
Continuity (2)
Provisional Application 63252301 · Oct 5, 2021
Related Publication 20230109379A1 · Apr 6, 2023
References Cited (24)
US 7127127B2 · Jojic · 2006 [cited by examiner]
US 11347973B2 · Rhee · 2022 [cited by examiner]
US 20170372155A1 · Odry · 2017 [cited by examiner]
US 20170372193A1 · Mailhe · 2017 [cited by examiner]
US 20200020098A1 · Odry · 2020 [cited by examiner]
US 20200065626A1 · Kaufhold · 2020 [cited by examiner]
US 20200234402A1 · Schwartz · 2020 [cited by examiner]
US 20200273167A1 · Wilson · 2020 [cited by examiner]
US 20210073959A1 · Elmalem · 2021 [cited by examiner]
US 20210360199A1 · Oz · 2021 [cited by examiner]
US 20210392296A1 · Rabinovich · 2021 [cited by examiner]
US 20220051412A1 · Gronau · 2022 [cited by examiner]
US 20220107378A1 · Dey · 2022 [cited by examiner]
US 20220189133A1 · Fuchs · 2022 [cited by examiner]
US 20220198612A1 · Weinmann · 2022 [cited by examiner]
US 20220215510A1 · Weinmann · 2022 [cited by examiner]
US 20220222781A1 · Jacob · 2022 [cited by examiner]
US 20230067841A1 · Saharia · 2023 [cited by examiner]
Dietmar Uttenweiler et al., “Spatiotemporal anisotropic diffusion filtering to improve signal-to-noise ratios and object restoration in fluorescence microscopic image sequences,” Sep. 12, 2002, Journal of Biomedical Opt… [cited by examiner]
Yueqin Yin et al.,“DiffGAR: Model-Agnostic Restoration from Generative Artifacts Using Image-to-Image Diffusion Models,” Mar. 30, 2023 ,CSAI '22: Proceedings of the 2022 6th International Conference on Computer Science … [cited by examiner]
Jonathan Ho et al.,“Denoising Diffusion Probabilistic Models,” Dec. 16, 2020, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, pp. 1-8. [cited by examiner]
V B Surya Prasath et al.,“Analysis of adaptive forward-backward diffusion flows with applications in image processing,” Sep. 24, 2015,Inverse Problems 31,pp. 6-25. [cited by examiner]
Siwei Yu et al.,“Deep learning for denoising,” Oct. 9, 2019, Geophysics, vol. 84, No. 6 (Nov.-Dec. 2019), pp. V333-V342. [cited by examiner]
Kevin De Haan et al.,“Deep-Learning-Based Image Reconstruction and Enhancement in Optical Microscopy,” Dec. 26, 2019, Proceedings of the IEEE | vol. 108, No. 1, Jan. 2020,pp. 30-46. [cited by examiner]
Cited By (1)
US 12,555,200