IP Library › Granted Patent US 12,348,700
Granted Patent B2
US 12,348,700 · App. 17/739,572 · Granted Jul 1, 2025

Real-time novel view synthesis with forward warping and depth

Inventors: Justin Johnson (Ann Arbor, MI); Ang Cao (Ann Arbor, MI); Chris Rockwell (Ann Arbor, MI)
Assignee: The Regents of the University of Michigan
H04N13/275G06T3/18G06V10/7715H04N13/257H04N2013/0088H04N13/282
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,348,700
App. No.
17/739,572
Granted
Jul 1, 2025
Kind
B2
Abstract

A fast and generalizable novel view synthesis method with sparse inputs is disclosed. The method may comprise: accessing at least a first input image with a first view of a subject in the first input image, and a second input image with a second view of the subject in the second input image using a computer system; estimating depths for pixels in the at least first and second input images; constructing a point cloud of image features from the estimated depths; and synthesizing a novel view by forward warping by using a point cloud rendering of the constructed point cloud.

Claims (29)

1. A method for novel view synthesis, the method comprising:

accessing at least a first input image with a first view of a subject in the first input image, and a second input image with a second view of the subject in the second input image using a computer system;

estimating depths for pixels in the at least first and second input images;

constructing a point cloud of image features from the estimated depths;

synthesizing a novel view by forward warping by using a point cloud rendering of the constructed point cloud, wherein synthesizing the novel view includes generating fused data by fusing the at least first input image and the second input image, wherein generating fused data includes using a fusion Transformer T, and rendering a set of feature maps {{tilde over (F)} i } from the point cloud and fused into a feature map.

2. The method of claim 1 , further comprising modeling view-dependent effects from the synthesized novel view.

3. The method of claim 2 , wherein view-dependent effects include missing pixel data in the synthesized novel view.

4. The method of claim 1 , wherein the feature map is decoded into an RGB image by a refinement module.

5. The method of claim 1 , wherein the fusion Transformer T extracts feature vectors from {{tilde over (F)}i} as inputs and output a fused one at each pixel.

6. The method of claim 1 , further comprising generating output pixels for the synthesized novel view by inpainting missing pixel data based on the fused data.

7. The method of claim 1 , wherein the synthesized novel view includes a viewpoint of the subject different from that at least first input and second input image.

8. The method of claim 1 , wherein the computer system is further configured to access a plurality of input images.

9. The method of claim 1 , wherein estimating the depth for pixels in the at least first and second input images includes inputting the at least first and second images to a multi-view stereo algorithm (MVS) module and receiving, from the MVS module, an initial depth estimation.

10. The method of claim 9 , wherein estimating the depth further includes providing the initial depth estimation to a U-Net and receiving, from the U-Net a refined depth, wherein the refined depth is used to build the point cloud.

11. The method of claim 1 , further including depth information from one or more sensors, wherein the estimated depths are based at least in part on the depth information.

12. A system for novel view synthesis, the system comprising:

a computer system configured to:

i) access at least a first input image with a first view of a subject in the first input image, and a second input image with a second view of the subject in the second input image;

ii) estimate depths for pixels in the at least first and second input images;

iii) construct a point cloud of image features from the estimated depths; and

iv) synthesize a novel view by forward warping by using a point cloud rendering of the constructed point cloud, wherein synthesizing the novel view includes generating fused data using a fusion Transformer T, and rendering a set of feature maps {{tilde over (F)}i} from the point cloud and fused into a feature map.

13. The system of claim 12 , wherein the computer system is further configured to model view-dependent effects from the synthesized novel view.

14. The system of claim 13 , wherein view-dependent effects include missing pixel data in the synthesized novel view.

15. The system of claim 13 , wherein the generating the fused data includes fusing the at least first input image and the second input image.

16. The system of claim 15 , wherein the computer system is further configured to generate output pixels for the synthesized novel view by inpainting missing pixel data based on the fused data.

17. The system of claim 12 , wherein the computer system is further configured to decode the feature map into an RGB image using a refinement module.

18. The system of claim 12 , wherein the computer system is further configured to extract feature vectors from {{tilde over (F)}i} as inputs and output a fused one at each pixel using the fusion Transformer T.

19. The system of claim 12 , wherein the synthesized novel view includes a viewpoint of the subject different from that at least first input and second input image.

20. The system of claim 12 , wherein the computer system is further configured to access a plurality of input images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2025
From: JOHNSON, JUSTIN; CAO, ANG; ROCKWELL, CHRIS
To: THE REGENTS OF THE UNIVERSITY OF MICHIGAN
Reel/Frame 071175/0969 →
Continuity (1)
Related Publication 20230362347A1 · Nov 9, 2023
References Cited (96)
US 20130215220A1 · Wang · 2013 [cited by examiner]
US 20170293810A1 · Allen · 2017 [cited by examiner]
US 20200226816A1 · Kar · 2020 [cited by examiner]
Zhu, Shiping, “An Improved Depth Image Based Virtual View Synthesis Method for Interactive 3D Video”, Aug. 2019, IEEE Access, vol. 7, pp. 115171-115180 (Year: 2019). [cited by examiner]
Le, Hoang-An, “Novel View Synthesis from Single Images via Point Cloud Transformation”, Computer Vision and Pattern Recognition, Sep. 2020, arXiv, pp. 1-19 (Year: 2020). [cited by examiner]
Song, Zhembo, “Deep Novel View Synthesis from Colored 3D Point Clouds”. Computer Vision—ECCV 2020, Nov. 2020, SpringerLink, pp. 1-17 (Year: 2020). [cited by examiner]
Aanaes, H. et al., Large-Scale Data for Multiple-View Stereopsis, International Journal of Computer Vision, 2016, 120:153-168. [cited by applicant]
Brock, A. et al., Large Scale GAN Training for High Fidelity Natural Image Synthesis, arXiv:1809.11096, 2018, pp. 1-29. [cited by applicant]
Carion, N. et al., End-to-End Object Detection with Transformers, In European Conference on Computer Vision, 2020, pp. 213-229. [cited by applicant]
Chang, A. et al., ShapeNet: An Information-Rich 3D Model Repository, arXiv:1512.03012, 2015, pp. 1-11. [cited by applicant]
Chang, A. et al., Matterport3D: Learning from RGB-D Data in Indoor Environments, arXiv:1709.06158, 2017, 25 pages. [cited by applicant]
Chaurasia, G. et al., Depth Synthesis and Local Warps for Plausible Image-Based Navigation, ACM Transactions on Graphics (TOG), 2013, 32(3):1-12. [cited by applicant]
Chen, W. et al., Single-Image Depth Perception in the Wild, 30th Conference on Neural Information Processing Systems, NIPS, 2016, pp. 1-9. [cited by applicant]
Chen, W. et al., Learning to Predict 3D Objects with an Interpolation-Based Differentiable Renderer, 33rd Conference on Neural Information Processing Systems (NeurIPS), 2019, pp. 1-11. [cited by applicant]
Chen, A. et al., MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View Stereo, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14124-14133. [cited by applicant]
Choi, I. et al., Extreme View Synthesis, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 7781-7790. [cited by applicant]
Dai, A. et al., ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5828-5839. [cited by applicant]
Dai, Y. et al., MVS2: Deep Unsupervised Multi-View Stereo with Multi-View Symmetry, arXiv:1908. 11526, 2019, 10 pages. [cited by applicant]
Debevec, P. et al., Modeling and Rendering Architecture from Photographs: A Hybrid Geometry- and Image-Based Approach, In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, … [cited by applicant]
Deng, K. et al., Depth-Supervised NeRF: Fewer Views and Faster Training for Free, arXiv:2107.02791, 2021, pp. 1-13. [cited by applicant]
Dosovitskiy, A. et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv:2010.11929, 2021, pp. 1-22. [cited by applicant]
Garbin, S. et al., FastNeRF: High-Fidelity Neural Rendering at 200FPS, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14346-14355. [cited by applicant]
Genova, K. et al., Local Deep Implicit Functions for 3D Shape, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 4857-4866. [cited by applicant]
Gortler, S. et al., The Lumigraph, In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, pp. 43-54. [cited by applicant]
Guo, P. et al., Fast and Explicit Neural View Synthesis, arXiv:2107.05775, 2021, pp. 1-21. [cited by applicant]
Hani, N. et al., Continuous Object Representation Networks: Novel View Synthesis Without Target View Supervision, arXiv:2007.15627, 2020, 22 pages. [cited by applicant]
He, K. et al., Deep Residual Learning for Image Recognition, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Hedman, P. et al., Scalable Inside-Out Image-Based Rendering, ACM Transactions on Graphics (TOG), 2016, 35 (6):1-11. [cited by applicant]
Hedman, P. et al., Deep Blending for Free-Viewpoint Image-Based Rendering, ACM Transactions on Graphics (TOG), 2018, 37(6):1-15. [cited by applicant]
Hedman, P. et al., Instant 3D Photography, ACM Transactions on Graphics (TOG), 2018, 37(4):1-12. [cited by applicant]
Hedman, P. et al., Baking Neural Radiance Fields for Real-Time View Synthesis, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5875-5884. [cited by applicant]
Hu, R. et al., Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single Image, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12528-12537. [cited by applicant]
Huang, P. et al., DeepMVS: Learning Multi-View Stereopsis, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2821-2830. [cited by applicant]
Huang, R. et al., An LSTM Approach to Temporal 3D Object Detection in Lidar Point Clouds, arXiv:2007.12392, 2020, pp. 1-18. [cited by applicant]
Jain, A. et al., Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5885-5894. [cited by applicant]
Jatavallabhula, K. et al., Kaolin: A Pytorch Library for Accelerating 3D Deep Learning Research, arXiv:1911.05063, 2019, pp. 1-7. [cited by applicant]
Jensen, R. et al., Large Scale Multi-View Stereopsis Evaluation, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 406-413. [cited by applicant]
Jiang, C. et al., Local Implicit Grid Representations for 3D Scenes, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6001-6010. [cited by applicant]
Jiang, Y. et al., SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 1251-1261. [cited by applicant]
Karras, T. et al., Analyzing and Improving the Image Quality of StyleGAN, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 8110-8119. [cited by applicant]
Kingma, D. et al., Adam: A Method for Stochastic Optimization, arXiv:1412.6980, 2015, pp. 1-13. [cited by applicant]
Le, H. et al., Novel View Synthesis from Single Images via Point Cloud Transformation, arXiv:2009.08321, 2020, pp. 1-19. [cited by applicant]
Ledig, C. et al., Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4681-4690. [cited by applicant]
Levoy, M. et al., Light Field Rendering, In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, pp. 31-42. [cited by applicant]
Liu, S. et al., Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction, arXiv:1901.05567, 2019, 10 pages. [cited by applicant]
Liu, S. et al., DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2019-2028. [cited by applicant]
Liu, A. et al., Infinite Nature: Perpetual View Generation of Natural Scenes from a Single Image, arXiv:2012.09855, 2021, 17 pages. [cited by applicant]
Liu, L. et al., Neural Sparse Voxel Fields, arXiv:2007.11571, 2021, pp. 1-22. [cited by applicant]
Lombardi, S. et al., Neural vols. Learning Dynamic Renderable vols. from Images, arXiv:1906.07751, 2019, pp. 1-14. [cited by applicant]
Luo, X. et al., Consistent Video Depth Estimation, ACM Transactions on Graphics (TOG), 2020, vol. 39, No. 4, Article 71, pp. 1-13. [cited by applicant]
Martin-Brualla, R. et al., NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7210-7219. [cited by applicant]
Mescheder, L. et al., Occupancy Networks: Learning 3D Reconstruction in Function Space, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4460-4470. [cited by applicant]
Mildenhall, B. et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Communications of the ACM, 2022, 65(1):99-106. [cited by applicant]
Najibi, M. et al., DOPS: Learning to Detect 3D Objects and Predict their 3D Shapes, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11913-11922. [cited by applicant]
Neff, T. et al., DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle Networks, Eurographics Symposium on Rendering, 2021, 40(4):45-59. [cited by applicant]
Niemeyer, M. et al., Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3504… [cited by applicant]
Novotny, D. et al., PerspectiveNet: A Scene-Consistent Image Generator for New View Synthesis in Real Indoor Environments, Advances in Neural Information Processing Systems, 2019, 32:7601-7612. [cited by applicant]
Park, J. et al., DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 165-174. [cited by applicant]
Park, K. et al., Nerfies: Deformable Neural Radiance Fields, arXiv:2011.12948, 2021, pp. 1-18. [cited by applicant]
Penner, E. et al., Soft 3D Reconstruction for View Synthesis, ACM Transactions on Graphics (TOG), 2017, vol. 36, No. 6, Article 235, pp. 1-11. [cited by applicant]
Ravi, N. et al., Accelerating 3D Deep Learning with PyTorch3D, arXiv:2007.08501, 2020, pp. 1-18. [cited by applicant]
Riegler, G. et al., Free View Synthesis, arXiv:2008.05511, 2020, pp. 1-17. [cited by applicant]
Riegler, G. et al., Stable View Synthesis, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12216-12225. [cited by applicant]
Rockwell, C. et al., PixelSynth: Generating a 3D-Consistent Experience from a Single Image, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14104-14113. [cited by applicant]
Rombach, R. et al., Geometry-Free View Synthesis: Transformers and No. 3D Priors, In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14356-14366. [cited by applicant]
Rosu, R. et al., NeuralMVS: Bridging Multi-View Stereo and Novel View Synthesis, arXiv:2108:03880, 2021, pp. 1-9. [cited by applicant]
Schonberger, J. et al., Pixelwise View Selection for Unstructured Multi-View Stereo, In Computer Vision—ECCV 2016: 14th European Conference, 2016, pp. 501-518. [cited by applicant]
Shih, M. et al., 3D Photography Using Context-Aware Layered Depth Inpainting, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 8028-8038. [cited by applicant]
Silberman, N. et al., Indoor Segmentation and Support Inference from RGBD Images, In Computer Vision—ECCV 2012: 12th European Conference on Computer Vision, 2012, pp. 746-760. [cited by applicant]
Sitzmann, V. et al., DeepVoxels: Learning Persistent 3D Feature Embeddings, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2437-2446. [cited by applicant]
Sitzmann, V. et al., Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, 33rd Conference on Neural Information Processing Systems, 2019, pp. 1-12. [cited by applicant]
Song, S. et al., SUN RGB-D: A RGB-D Scene Understanding Benchmark Suite, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 567-576. [cited by applicant]
Song, Z. et al., Deep Novel View Synthesis from Colored 3D Point Clouds, ECCV 2020, LNCS 12369, pp. 1-17. [cited by applicant]
Srinivasan, P. et al., Pushing the Boundaries of View Extrapolation with Multiplane Images, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 175-184. [cited by applicant]
Stelzner, K. et al., Decomposing 3D Scenes into Objects via Unsupervised vol. Segmentation, arXiv:2104.01148, 2021, pp. 1-15. [cited by applicant]
Tatarchenko, M. et al., Multi-View 3D Models from Single Images with a Convolutional Network, ECCV 2016, Part VII, LNCS 9911, 2016, pp. 322-337. [cited by applicant]
Trevithick, A. et al., GRF: Learning a General Radiance Field for 3D Scene Representation and Rendering, arXiv:2010.04595, 2020, pp. 1-23. [cited by applicant]
Vaswani, A. et al., Attention Is All You Need, 31st Conference on Neural Information Processing Systems, 2017, pp. 1-11. [cited by applicant]
Wang, Z. et al., Image Quality Assessment: From Error Visibility to Structural Similarity, IEEE Transactions on Image Processing, 2004, 13(4):600-612. [cited by applicant]
Wang, T. et al., High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8798-8807. [cited by applicant]
Wang, F. et al., PatchmatchNet: Learned Multi-View Patchmatch Stereo, arXiv:2012.01411, 2020, pp. 1-16. [cited by applicant]
Wang, Q. et al., IBRNet: Learning Multi-View Image-Based Rendering, arXiv:2102.13090, 2021, pp. 1-10. [cited by applicant]
Wang, Z. et al., NeRF—: Neural Radiance Fields Without Known Camera Parameters, arXiv:2102.07064, 2022, pp. 1-17. [cited by applicant]
Wiles, O. et al., SynSin: End-to-End View Synthesis from a Single Image, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 7467-7477. [cited by applicant]
Wizadwongsa, S. et al., NeX: Real-time View Synthesis with Neural Basis Expansion, arXiv:2103.05606, 2021, pp. 1-14. [cited by applicant]
Xian, W. et al., Space-Time Neural Irradiance Fields for Free-Viewpoint Video, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9421-9431. [cited by applicant]
Yang, C. et al., High-Resolution Image Inpainting Using Multi-Scale Neural Patch Synthesis, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6721-6729. [cited by applicant]
Yao, Y. et al., MVSNet: Depth Inference for Unstructured Multi-View Stereo, In Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 767-783. [cited by applicant]
Yu, A. et al., pixelNeRF: Neural Radiance Fields from One or Few Images, arXiv:2012.02190, 2020, pp. 1-20. [cited by applicant]
Yu, A., et al., PlenOctrees for Real-time Rendering of Neural Radiance Fields, arXiv:2103.14024, 2021, pp. 1-18. [cited by applicant]
Yu, J. et al., Generative Image Inpainting with Contextual Attention, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5505-5514. [cited by applicant]
Zhang, R. et al., The Unreasonable Effectiveness of Deep Features as a Perceptual Metric, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586-595. [cited by applicant]
Zhang, H. et al., Self-Attention Generative Adversarial Networks, In International Conference on Machine Learning, PMLR, 2019, pp. 7354-7363. [cited by applicant]
Zhang, K. et al., NeRF++: Analyzing and Improving Neural Radiance Fields, arXiv:2010.07492, 2020, pp. 1-9. [cited by applicant]
Zhou, T. et al., Stereo Magnification: Learning View Synthesis Using Multiplane Images, ACM Trans. Graph, 2018, vol. 37, No. 4, Article 65, pp. 1-12. [cited by applicant]
Zhu, J. et al., Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks, In Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 2223-2232. [cited by applicant]