IP Library › Granted Patent US 12,340,457
Granted Patent B2
US 12,340,457 · App. 18/207,923 · Granted Jun 24, 2025

Radiance field gradient scaling for unbiased near-camera training

Inventors: Julien Philip (London, GB); Valentin Deschaintre (London, GB)
Assignee: Adobe Inc.
G06T15/06G06T3/40G06T7/70G06T7/90G06V10/761G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,457
App. No.
18/207,923
Granted
Jun 24, 2025
Kind
B2
Abstract

Methods and systems disclosed herein relate generally to radiance field gradient scaling for unbiased near-camera training. In a method, a processing device accesses an input image of a three-dimensional environment comprising a plurality of pixels, each pixel comprising a pixel color. The processing device determines a camera location based on the input image and a ray from the camera location in a direction of a pixel. The processing device integrates sampled information from a volumetric representation along the ray from the camera location to obtain an integrated color. The processing device trains a machine learning model configured to predict a density and a color, comprising minimizing a loss function using a scaling factor that is determined based on a distance between the camera location and a point along the ray. The processing device outputs the trained ML model for use in rendering an output image.

Claims (63)

1. A method comprising one or more computing devices performing operations comprising:

accessing an input image of a three-dimensional (3D) environment, the input image comprising a plurality of pixels, wherein each pixel of the plurality of pixels comprises a pixel color;

determining a camera location based on the input image of the 3D environment;

determining a ray from the camera location in a direction of a pixel of the plurality of pixels;

integrating sampled information from a volumetric representation of the 3D environment along the ray from the camera location to obtain an integrated color corresponding to the pixel;

training a machine learning (ML) model configured to predict a density and a color of the 3D environment, the training comprising minimizing a loss function using a scaling factor that is determined based on a distance between the camera location and a point along the ray, the loss function defined based on a difference between the integrated color and the pixel color of the pixel; and

outputting the trained ML model for use in rendering an output image of the 3D environment.

2. The method of claim 1 , wherein the scaling factor is determined further based on a scene scale factor, wherein the scene scale factor is based on the distance between the camera location and a feature of the 3D environment.

3. The method of claim 1 , wherein the sampled information from the volumetric representation of the 3D environment along the ray from the camera location is generated by:

determining a plurality of ray locations along the ray; and

for each ray location of the plurality of ray locations along the ray:

determining a sampled density and a sampled color from the volumetric representation associated with the ray location.

4. The method of claim 3 , wherein minimizing the loss function using the scaling factor that is determined based on the distance between the camera location and the point along the ray comprises:

for each ray location of the plurality of ray locations along the ray:

determining a point-wise gradient of the ML model, based on the difference between the integrated color and the pixel color of the pixel; and

scaling the point-wise gradient using the scaling factor based on the distance between the camera location and the point along the ray to obtain a scaled point-wise gradient;

aggregating the scaled point-wise gradients determined for each ray location of the plurality of ray locations to obtain an aggregated scaled gradient; and

applying the aggregated scaled gradient to one or more parameters of the ML model.

5. The method of claim 4 , wherein the scaling factor is the minimum of 1 and a distance term based on the distance between the camera location and the point along the ray.

6. The method of claim 1 , wherein the ML model is a multilayer perceptron (MLP).

7. The method of claim 1 , wherein the volumetric representation comprises one of an MLP, a voxel hash-grid, a tensor decomposition, or voxels.

8. A system, comprising:

a memory component; and

one or more processing devices coupled to the memory component configured to perform operations comprising:

accessing a trained machine learning (ML) model, wherein the trained ML model is trained to predict a density and a color of a 3D environment by minimizing a loss function based on a plurality of scaling factors, wherein each of the plurality of scaling factors are determined based on a distance between a first camera location and a point along a first ray projected from the first camera location towards a pixel in an input image comprising a pixel color;

receiving a second camera location and a camera direction;

using the trained ML model, determining a plurality of densities and colors of the 3D environment from a perspective of the second camera location at a plurality of respective points sampled along a second ray projected from the second camera location in the direction of the camera direction; and

aggregating the plurality of densities and colors of the 3D environment to generate an output pixel comprising an integrated color that represents the 3D environment.

9. The system of claim 8 , wherein each scaling factor is determined further based on a scene scale factor, wherein the scene scale factor is based on a distance scale that characterizes the 3D environment.

10. The system of claim 8 , wherein predicting the density and the color of a 3D environment comprises:

determining a plurality of ray locations along the first ray; and

for each ray location of the plurality of ray locations along the first ray:

determining a sampled density and a sampled color from a volumetric representation associated with the ray location.

11. The system of claim 10 , wherein minimizing the loss function based on the plurality of scaling factors, wherein each of the plurality of scaling factors are determined based on the distance between the first camera location and the point along the first ray projected from the first camera location towards the pixel in the input image comprising the pixel color comprises:

for each ray location of the plurality of ray locations along the first ray:

determining a point-wise gradient of the ML model, based on a difference between the integrated color and the pixel color of the pixel; and

scaling the point-wise gradient using the scaling factor based on the distance between the first camera location and the point along the first ray to obtain a scaled point-wise gradient;

aggregating the scaled point-wise gradients determined for each ray location of the plurality of ray locations to obtain an aggregated scaled gradient; and

applying the aggregated scaled gradient to one or more parameters of the ML model.

12. The system of claim 8 , wherein the ML model is a fully-connected neural network.

13. The system of claim 8 , wherein determining the plurality of densities and colors of the 3D environment from the perspective of the second camera location at the plurality of respective points sampled along the second ray projected from the second camera location is based on a volumetric representation, wherein the volumetric representation comprises one of a multilayer perceptron (MLP), a voxel hash-grid, a tensor decomposition, or voxels.

14. A computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more processing devices to perform actions including:

accessing a machine learning (ML) model trained via a training process, wherein the ML model is configured to predict a color of a 3D environment, wherein the training process comprises:

a step for determining a plurality of densities and colors of the 3D environment from a training perspective of a first camera location at a plurality of respective points sampled along a first ray projected from the first camera location;

a step for aggregating the plurality of densities and colors of the 3D environment to generate an integrated pixel color; and

a step for minimizing a loss function of the ML model using a scaling factor that is determined based on a distance between the first camera location and a point along the first ray projected from the first camera location towards the 3D environment; and

receiving a second camera location and a camera direction;

using the trained ML model, determining a plurality of densities and colors of the 3D environment from a perspective of the second camera location at a plurality of respective points sampled along a second ray projected from the second camera location in the direction of the camera direction; and

aggregating the plurality of densities and colors of the 3D environment to generate an output pixel comprising an integrated color that represents the 3D environment.

15. The computer program product of claim 14 , wherein the scaling factor is determined further based on a scene scale factor, wherein the scene scale factor is based on the distance between the first camera location and a feature of the 3D environment.

16. The computer program product of claim 14 , wherein the step for determining the plurality of densities and colors of the 3D environment from the training perspective of the first camera location at the plurality of respective points sampled along the first ray projected from the first camera location comprises:

determining a plurality of ray locations along the first ray; and

for each ray location of the plurality of ray locations along the first ray:

determining a sampled density and a sampled color based on a volumetric representation associated with the ray location.

17. The computer program product of claim 16 , wherein minimizing the loss function using the scaling factor that is determined based on the distance between the first camera location and the point along the first ray comprises:

for each ray location of the plurality of ray locations along the first ray:

determining a point-wise gradient of the ML model, based on a difference between the integrated color and the pixel color of the pixel; and

scaling the point-wise gradient using the scaling factor based on the distance between the first camera location and the point along the first ray to obtain a scaled point-wise gradient;

aggregating the scaled point-wise gradients determined for each ray location of the plurality of ray locations to obtain an aggregated scaled gradient; and

applying the aggregated scaled gradient to one or more parameters of the ML model.

18. The computer program product of claim 17 , wherein the scaling factor is the minimum of 1 and the square of the distance between the first camera location and the point along the first ray.

19. The computer program product of claim 14 , wherein the ML model is a multilayer perceptron (MLP).

20. The computer program product of claim 14 , wherein determining the plurality of densities and colors of the 3D environment from the perspective of the second camera location at the plurality of respective points sampled along the second ray projected from the second camera location is based on a volumetric representation, wherein the volumetric representation comprises one of an MLP, a voxel hash-grid, a tensor decomposition, or voxels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2023
From: PHILIP, JULIEN; DESCHAINTRE, VALENTIN
To: ADOBE INC.
Reel/Frame 063909/0672 →
Continuity (1)
Related Publication 20240412444A1 · Dec 12, 2024
References Cited (39)
US 11869132B2 · Kim · 2024 [cited by examiner]
US 20210174539A1 · Duong · 2021 [cited by examiner]
US 20240070972A1 · Kosiorek · 2024 [cited by examiner]
US 20250037244A1 · Mildenhall · 2025 [cited by examiner]
Mildenhall et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (Year: 2022). [cited by examiner]
Deng et al., Depth-supervised NeRF: Fewer Views and Faster Training for Free (Year: 2022). [cited by examiner]
Kendall et al., PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization (Year: 2016). [cited by examiner]
Zhang et al. NeRFactor: Neural Factorization of Shape and Reflectance Under an Unknown Illumination (Year: 2021). [cited by examiner]
Aliev et al., Neural Point-Based Graphics, Available online at: arXiv:1906.08240v3.2, Apr. 5, 2020, pp. 1-16. [cited by applicant]
Barron et al., Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Mar. 25, 2022, pp. 5470-5479. [cited by applicant]
Barron et al., Mip-Nerf: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields, Available online at: arXiv:2103.13415.2,4, 2021, pp. 5855-5864. [cited by applicant]
Chen et al., TensoRF: Tensorial Radiance Fields, European Conference on Computer Vision (ECCV), Nov. 29, 2022, pp. 1-33. [cited by applicant]
Hedman et al., Baking Neural Radiance Fields for Real-Time View Synthesis, International Conference on Computer Vision, 2021, pp. 5875-5884. [cited by applicant]
Hedman et al., Deep Blending for Free-Viewpoint Image-Based Rendering, ACM Transactions on Graphics, vol. 37, No. 6, Nov. 2018, pp. 257:1-257:15. [cited by applicant]
Hedman et al., Scalable Inside-Out Image-Based Rendering, ACM Transactions on Graphics, vol. 35, No. 6, Nov. 2016, pp. 231:1-231:11. [cited by applicant]
Jambon et al., Nerfshop: Interactive Editing of Neural Radiance Fields, Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 6, No. 1, May 2023, pp. 1-21. [cited by applicant]
Kopanas et al., Neural Point Catacaustics for Novel-view Synthesis of Reflections, ACM Transactions on Graphics, vol. 41, No. 6, Dec. 2022, pp. 1-15. [cited by applicant]
Kopanas et al., Point-Based Neural Rendering with Per-View Optimization, Computer Graphics Forum, vol. 40, No. 4, Sep. 8, 2021, 15 pages. [cited by applicant]
Mildenhall et al., MultiNeRF: A Code Release for Mip-NeRF 360, Ref-NeRF, and RawNeRF, Available Online at: https://github.com/google-research/multinerf, Oct. 13, 2022, 8 pages. [cited by applicant]
Mildenhall et al., NeRF in the Dark: High Dynamic Range View Synthesis from Noisy Raw Images, Conference on Computer Vision and Pattern Recognition, Nov. 26, 2021, 18 pages. [cited by applicant]
Mildenhall et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, European Conference on Computer Vision, Aug. 3, 2020, pp. 1-25. [cited by applicant]
Muller, et al., Instant Neural Graphics Primitives with a Multiresolution Hash Eencoding, ACM Trans. Graph., vol. 41, No. 4, Article 102, Jul. 2022, 15 pages. [cited by applicant]
Nimier-David et al., Unbiased Inverse Volume Rendering with Differential Trackers, ACM Transactions on Graphics, vol. 41, No. 4, Available Online at: https://dl.acm.org/doi/pdf/10.1145/3528223.3530073, Jul. 2022, pp. 1-… [cited by applicant]
Ortiz-Cayon et al., A Bayesian Approach for Selective Image-Based Rendering using Superpixels, International Conference on 3D Vision (3DV), Available Online at: http://www-sop.inria.fr/reves/Basilic/2015/ODD15, 2015, 9 … [cited by applicant]
Philip et al., Free-viewpoint Indoor Neural Relighting from Multi-View Stereo, ACM Transactions on Graphics, vol. 1, No. 1, Article 1, Available Online at: http://www-sop.inria.fr/reves/Basilic/2021/PMGD21, Jan. 2021, 1… [cited by applicant]
Rakhimov et al., NPBG++: Accelerating Neural Point-Based Graphics, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, 11 pages. [cited by applicant]
Reiser et al., KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs, Available Online at: https://arxiv.org/pdf/2103.13744.pdf, Aug. 2, 2021, 11 pages. [cited by applicant]
Riegler et al., Free View Synthesis, In European Conference on Computer Vision, Aug. 12, 2020, 17 pages. [cited by applicant]
Riegler et al., Stable View Synthesis, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, May 2, 2021, 13 pages. [cited by applicant]
Roessle et al., Dense Depth Priors for Neural Radiance Fields from Sparse Input Views, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, 12 pages. [cited by applicant]
Sun et al., Direct Voxel Grid Optimization: Super-Fast Convergence for Radiance Fields Reconstruction, Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 3, 2022, 25 pages. [cited by applicant]
Sun et al., Improved Direct Voxel Grid Optimization for Radiance Fields Reconstruction, Available Online at: https://arxiv.org/pdf/2206.05085.pdf, Jul. 2, 2022, 5 pages. [cited by applicant]
Tewari et al., Advances in Neural Rendering, Computer Graphics Forum, State of The Art Report, vol. 41, No. 2, Available Online at: https://onlinelibrary.wiley.com/doi/abs/10.1111/cgf.14507,arXiv:https://onlinelibrary.w… [cited by applicant]
Verbin et al., Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, In Proceeding IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Dec. 7, 2021, 14 pages. [cited by applicant]
Yang et al., FreeNeRF: Improving Few-Shot Neural Rendering with free Frequency Regularization, In Proceeding IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Mar. 13, 2023, 16 pages. [cited by applicant]
Yu et al., PlenOctrees for Real-Time Rendering of Neural Radiance Fields, Computer Vision and Pattern Recognition, Aug. 2021, 18 pages. [cited by applicant]
Yu et al., Plenoxels: Radiance Fields without Neural Networks, In Proceeding IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Dec. 9, 2021, 21 pages. [cited by applicant]
Zhang et al., NeRF++: Analyzing and Improving Neural Radiance Fields, Available online at https://arxiv.org/pdf/2010.07492.pdf, Oct. 21, 2020, pp. 1-9. [cited by applicant]
Zhang et al., The Unreasonable Effectiveness of Deep Features as a Perceptual Metric, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available Online at: https://openaccess.thecvf.com/… [cited by applicant]