IP Library Granted Patent US 12,243,273
Granted Patent B2
US 12,243,273 · App. 17/571,285 · Granted Mar 4, 2025

Neural 3D video synthesis

Inventors: Zhaoyang Lv (Redmond, WA); Miroslava Slavcheva (Seattle, WA); Tianye Li (Los Angeles, CA); Michael Zollhoefer (Pittsburgh, PA); Simon Gareth Green (Deptford, GB); Tanner Schmidt (Seattle, WA); Michael Goesele (Woodinville, WA); Steven John Lovegrove (Woodinville, WA); Christoph Lassner (San Francisco, CA); Changil Kim (Seattle, WA)
Assignee: META PLATFORMS TECHNOLOGIES, LLC
G06T7/97G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,273
App. No.
17/571,285
Granted
Mar 4, 2025
Kind
B2
Abstract

In one embodiment, a method includes initializing latent codes respectively associated with times associated with frames in a training video of a scene captured by a camera. For each of the frames, a system (1) generates rendered pixel values for a set of pixels in the frame by querying NeRF using the latent code associated with the frame, a camera viewpoint associated with the frame, and ray directions associated with the set of pixels, and (2) updates the latent code associated with the frame and the NeRF based on comparisons between the rendered pixel values and original pixel values for the set of pixels. Once trained, the system renders output frames for an output video of the scene, wherein each output frame is rendered by querying the updated NeRF using one of the updated latent codes corresponding to a desired time associated with the output frame.

Claims (48)

1. A method comprising:

rendering output frames for an output video of a scene, wherein each output frame is rendered by querying an updated neural radiance field (NeRF) using at least one updated latent code respectively associated with a desired time associated with the output frame, a desired viewpoint for the output frame, and ray directions associated with pixels in the output frame, wherein the updated NeRF and the at least one updated latent code are based on:

a set of pixels selected from a plurality of pixels of at least two frames in a training video for the scene captured by a camera, the set of pixels selected based on temporal variances of the plurality of pixels,

rendered pixel values for the set of pixels that were identified by querying a pre-trained NeRF using:

ray directions associated with the set of pixels,

initialized latent codes respectively associated with times associated with the at least two frames in the training video, and

a first camera viewpoint associated with the at least two frames; and

a comparison between the rendered pixel values and original pixel values for the set of pixels.

2. The method of claim 1 , wherein the updated NeRF and the at least one updated latent code are further based on:

a second training video of the scene captured by a second camera having a second camera viewpoint different from the first camera viewpoint, wherein the second training video and the training video are captured concurrently.

3. The method of claim 2 , wherein a first frame of the at least two frames in the training video and a second frame in the second training video are both associated with a particular time and used for updating the latent code respectively associated with the particular time.

4. The method of claim 1 , wherein the desired viewpoint for the output frame is different from any camera viewpoint associated with any frame of the at least two frames in the training video.

5. The method of claim 1 , wherein each of the initialized latent codes consists of a predetermined number of values.

6. The method of claim 1 , further comprising:

rendering, for the output video, an additional output frame associated with an additional desired time that is temporally between two adjacent frames of the frames in the training video, wherein the additional output frame is rendered by querying the updated NeRF using an interpolated latent code generated by interpolating updated latent codes respectively associated with the two adjacent frames.

7. The method of claim 1 , wherein the temporal variances are used to determine probabilities of the corresponding pixels in the plurality of pixels being selected into the set of pixels used for updating the NeRF and the at least one updated latent code.

8. The method of claim 1 , wherein the at least two frames of the training video used for the updated NeRF and the at least one updated latent code are keyframes within a larger set of frames of the training video, and the keyframes were selected from the larger set of frames based on positions of the keyframes in the larger set of frames.

9. The method of claim 8 , wherein the positions of the keyframes in the larger set of frames are equally spaced by a predetermined number of frames.

10. The method of claim 8 ,

wherein the updated NeRF and the updated at least one latent code are further based on using additional frames in the larger set of frames in between the keyframes.

11. One or more computer-readable non-transitory storage media storing instructions, which, when executed by a system that includes an apparatus and one or more processors, causes the one or more processors to perform a set of operations, including:

rendering output frames for an output video of a scene, wherein each output frame is rendered by querying an updated neural radiance field (NeRF) using at least one updated latent code respectively associated with a desired time associated with the output frame, a desired viewpoint for the output frame, and ray directions associated with pixels in the output frame, wherein the updated NeRF and the at least one updated latent code are based on:

a set of pixels selected from a plurality of pixels of at least two frames in a training video for the scene captured by a camera, the set of pixels selected based on temporal variances of the plurality of pixels,

rendered pixel values for the set of pixels that were identified by querying a pre-trained NeRF using:

ray directions associated with the set of pixels,

initialized latent codes respectively associated with times associated with the at least two frames in the training video, and

a first camera viewpoint associated with the at least two frames; and

a comparison between the rendered pixel values and original pixel values for the set of pixels.

12. The one or more computer-readable non-transitory storage media of claim 11 , further storing instructions for:

rendering, for the output video, an additional output frame associated with an additional desired time that is temporally between two adjacent frames of the frames in the training video, wherein the additional output frame is rendered by querying the updated NeRF using an interpolated latent code generated by interpolating updated latent codes respectively associated with the two adjacent frames.

13. The one or more computer-readable non-transitory storage media of claim 11 , wherein the at least two frames of the training video used for the updated NeRF and the at least one updated latent code are keyframes within a larger set of frames of the training video, and

the keyframes were selected from the larger set of frames based on positions of the keyframes in the larger set of frames.

14. The one or more computer-readable non-transitory storage media of claim 13 ,

wherein the updated NeRF and the at least one updated latent code are further based on using additional frames in the larger set of frames in between the keyframes.

15. A system comprising:

one or more processors; and

one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

render output frames for an output video of a scene, wherein each output frame is rendered by querying an updated neural radiance field (NeRF) using at least one of the updated latent code respectively associated with a desired time associated with the output frame, a desired viewpoint for the output frame, and ray directions associated with pixels in the output frame, wherein the updated NeRF and the at least one updated latent code are based on:

a set of pixels selected from a plurality of pixels of at least two frames in a training video for the scene captured by a camera, the set of pixels selected based on temporal variances of the plurality of pixels,

rendered pixel values for the set of pixels that were identified by querying a pre-trained NeRF using:

ray directions associated with the set of pixels,

initialized latent codes respectively associated with times associated with the at least two frames in the training video, and

a first camera viewpoint associated with the at least two frames; and

a comparison between the rendered pixel values and original pixel values for the set of pixels.

16. The system of claim 15 , wherein one or more of the processors are further operable when executing the instructions to:

render, for the output video, an additional output frame associated with an additional desired time that is temporally between two adjacent frames of the frames in the training video, wherein the additional output frame is rendered by querying the updated NeRF using an interpolated latent code generated by interpolating updated latent codes respectively associated with the two adjacent frames.

17. The system of claim 15 , wherein the at least two frames of the training video used for the updated NeRF and the at least one updated latent code are keyframes within a larger set of frames of the training video, and

the keyframes were selected from the larger set of frames based on positions of the keyframes in the larger set of frames.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2025
From: LV, ZHAOYANG; SLAVCHEVA, MIROSLAVA; LI, TIANYE; ZOLLHOEFER, MICHAEL; GREEN, SIMON GARETH; SCHMIDT, TANNER; GOESELE, MICHAEL; LOVEGROVE, STEVEN JOHN; LASSNER, CHRISTOPH; KIM, CHANGIL
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 069949/0095 →
CHANGE OF NAME Recorded Jul 6, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060591/0848 →
Continuity (2)
Provisional Application 63142234 · Jan 27, 2021
Related Publication 20220239844A1 · Jul 28, 2022
References Cited (138)
US 7151545B2 · Spicer · 2006 [cited by applicant]
US 11546568B1 · Yoon et al. · 2023 [cited by applicant]
US 20180189667A1 · Tsou et al. · 2018 [cited by applicant]
US 20220189104A1 · Wetzstein · 2022 [cited by examiner]
US 20220198731A1 · Lombardi · 2022 [cited by examiner]
US 20220198738A1 · Xu · 2022 [cited by examiner]
US 20220222897A1 · Yang et al. · 2022 [cited by applicant]
US 20220245910A1 · Lombardi · 2022 [cited by examiner]
US 20220301241A1 · Kim et al. · 2022 [cited by applicant]
US 20220301257A1 · Garbin et al. · 2022 [cited by applicant]
US 20220343522A1 · Bi et al. · 2022 [cited by applicant]
US 20230116250A1 · Kowalski · 2023 [cited by examiner]
US 20230230275A1 · Lin · 2023 [cited by examiner]
US 20230306655A1 · Duckworth · 2023 [cited by examiner]
US 20230360372A1 · Zhao · 2023 [cited by examiner]
US 20240005590A1 · Martin Brualla · 2024 [cited by examiner]
WO WO2020242170A1 · 2020 [cited by examiner]
Wang Z., et al., “Image Quality Assessment: From Error Visibility to Structural Similarity,” IEEE Transactions on Image Processing, Apr. 4, 2004, vol. 13 (4), pp. 1-14. [cited by applicant]
Wiegand T., et al., “Overview of the H.264/AVC Video Coding Standard,” IEEE Transactions on Circuits and Systems for Video Technology, 2003, vol. 13, No. 7, pp. 560-576. [cited by applicant]
Wiles O., et al., “SynSin: End-to-End View Synthesis From a Single Image,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7467-7477. [cited by applicant]
Wood D.N., et al., “Surface Light Fields for 3D Photography,” Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, 2000, pp. 287-296. [cited by applicant]
Xian W., et al., “Space-time Neural Irradiance Fields for Free-Viewpoint Video,” arXiv:2011.12950, 2020, 10 pages. [cited by applicant]
Xiao L., et al., “Neural Supersampling for Real-time Rendering,” ACM Trans. Graph., vol. 39, No. 4, Jul. 8, 2020, pp. 142:1-142:12. [cited by applicant]
Xu R., et al., “Deep Flow-Guided Video Inpainting,” Computer Vision and Pattern Recognition (CVPR), 2019, pp. 3723-3732. [cited by applicant]
Yao Y., et al., “MVSNet: Depth Inference for Unstructured Multi-view Stereo,” European Conference on Computer Vision (ECCV), 2018, 17 Pages. [cited by applicant]
Yao Y., et al., “MVSNet: Depth Inference for Unstructured Multiview Stereo,” European Conference on Computer Vision (ECCV), 2018, pp. 767-783. [cited by applicant]
Yariv L., et al., “Multiview Neural Surface Reconstruction by Disentangling Geometry and Appearance,” Neural Information Processing Systems (NeurIPS), 2020, 11 pages. [cited by applicant]
Yoon J. S., et al., “Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular Camera,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 5336-5345. [cited by applicant]
Yu A., et al., “pixelNeRF: Neural Radiance Fields from One or Few Images,” arXiv, 2020, 20 pages, https://arxiv.org/abs/2012.02190. [cited by applicant]
Yuan W., et al., “STaR: Self-supervised Tracking and Reconstruction of Rigid Objects in Motion with Neural Rendering,” arXiv:2101.01602, 2020, 12 pages. [cited by applicant]
Zhang K., et al., “NeRF++: Analyzing and Improving Neural Radiance Fields,” arXiv preprint arXiv:2010.07492, 2020, 9 Pages. [cited by applicant]
Zhang R., et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Zhou T., et al., “Stereo Magnification: Learning View Synthesis using Multiplane Images,” ACM Transactions Graph, Aug. 2018, vol. 37 (4), Article 65, pp. 65:1-65:12. [cited by applicant]
Zhou T., et al., “View Synthesis by Appearance Flow,” European Conference on Computer Vision (ECCV), 2016, pp. 286-301. [cited by applicant]
Zitnick C.L., et al., “High-Quality Video View Interpolation Using a Layered Representation,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2004, vol. 23, No. 3, pp. 600-608. [cited by applicant]
Zollhofer M., et al., “State of the Art on 3D Reconstruction with RGB-D Cameras,” Computer Graphics Forum, 2018, vol. 37, No. 2, pp. 625-652. [cited by applicant]
Gafni G., et al., “Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction,” 2020, 11 pages, Retrieved from Internet: URL: https://arxiv.org/abs/2012.03065. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/013888, mailed Aug. 30, 2022, 10 pages. [cited by applicant]
Li H., et al., “Temporally Coherent Completion of Dynamic Shapes,” ACM Transactions on Graphics (TOG), 2012, vol. 31, No. 1, pp. 1-11. [cited by applicant]
Li Z., et al., “Crowdsampling the Plenoptic Function,” European Conference on Computer Vision (ECCV), 2020, pp. 178-196. [cited by applicant]
Li Z., et al., “Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,” arXiv:2011.13084, 2020, 11 pages. [cited by applicant]
Lindell D.B., et al., “AutoInt: Automatic Integration for Fast Neural Volume Rendering,” arXiv:2012.01714, 2020, 15 pages. [cited by applicant]
Liu C., et al., “Neural RGB(r)D Sensing: Depth and Uncertainty From a Video Camera,” Computer Vision and Pattern Recognition (CVPR), 2019, pp. 10986-10995. [cited by applicant]
Liu L., et al., “Neural Sparse Voxel Fields,” Neural Information Processing Systems (NeurIPS), 2020, 20 pages. [cited by applicant]
Lombardi S., et al., “Neural Volumes: Learning Dynamic Renderable vols. from Images,” ACM Transactions Graph, Jun. 18, 2019, vol. 38 (4), Article 65, pp. 1-14, XP081383263. [cited by applicant]
Luo X., et al., “Consistent Video Depth Estimation,” ACM Transactions on Graphics (ACM), 2020, vol. 39, No. 4, pp. 71:1-71:13. [cited by applicant]
Martin-Brualla R., et al., “NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections,” Computer Vision and Pattern Recognition (CVPR), 2020, arXiv: 2008.02268v2 [cs.CV], 14 Pages. [cited by applicant]
Marwah K., et al., “Compressive Light Field Photography using Overcomplete Dictionaries and Optimized Projections,” ACM Transactions on Graphics (TOG), 2013, vol. 32, No. 4, pp. 1-12. [cited by applicant]
Meka A., et al., “Deep Reflectance Fields: High-Quality Facial Reflectance Field Inference from Color Gradient Illumination,” ACM Transactions on Graphics (TOG), 2019, vol. 38, No. 4, pp. 1-12. [cited by applicant]
Mescheder L., et al., “Occupancy Networks: Learning 3D Reconstruction in Function Space,” Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4460-4470. [cited by applicant]
Meshry M., et al., “Neural Rerendering in the Wild,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6878-6887. [cited by applicant]
Michalkiewicz M., “Implicit Surface Representations as Layers in Neural Networks,” International Conference on Computer Vision (ICCV), 2019, pp. 4743-4752. [cited by applicant]
Mildenhall B., et al., “Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines,” ACM Transactions on Graphics (TOG), Jul. 12, 2019, vol. 38 (4), pp. 1-14. [cited by applicant]
Mildenhall B., et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” European Conference on Computer Vision (ECCV), Aug. 3, 2020, 25 pages. [cited by applicant]
Muller T., et al., “Neural Importance Sampling,” ACM Transactions on Graphics (TOG), 2019, vol. 38, No. 5, pp. 1-19. [cited by applicant]
Newcombe R.A., et al., “DynamicFusion: Reconstruction and Tracking of Non-Rigid Scenes in Real-Time,” In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, Jun. 7-12, 2015, pp. 343-… [cited by applicant]
Newcombe R.A., et al., “KinectFusion: Real-Time Dense Surface Mapping and Tracking,” International Symposium on Mixed and Augmented Reality (ISMAR), 2011, 10 Pages. [cited by applicant]
Niebner M., et al., “Real-time 3D Reconstruction at Scale Using Voxel Hashing,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2013, vol. 32, No. 6, 11 Pages. [cited by applicant]
Niemeyer M., et al., “Differentiable Volumetric Rendering: Learning Implicit 3D Representations Without 3D Supervision,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 3504-3515. [cited by applicant]
Niemeyer M., et al., “Occupancy Flow: 4D Reconstruction by Learning Particle Dynamics,” International Conference on Computer Vision (ICCV), 2019, pp. 5379-5389. [cited by applicant]
Niklaus S., et al., “3D Ken Burns Effect from a Single Image,” ACM Transactions on Graphics (Proc. SIGGRAPH Asia), 2019, vol. 38, No. 6, pp. 1-15. [cited by applicant]
Oechsle M., et al., “Texture Fields: Learning Texture Representations in Function Space,” International Conference on Computer Vision (ICCV), 2019, pp. 4531-4540. [cited by applicant]
Orts-Escolano S., et al., “Holoportation: Virtual 3D Teleportation in Real-time,” In Proceedings of the 29th Annual Symposium on User Interface Software and Technology, 2016, pp. 741-754. [cited by applicant]
Park J. J., et al., “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation,” Computer Vision and Pattern Recognition (CVPR), 2019, pp. 165-174. [cited by applicant]
Park K., et al., “Deformable Neural Radiance Fields,” arXiv preprint arXiv:2011.12948, 2020, 12 pages. [cited by applicant]
Penner E., et al., “Soft 3D Reconstruction for View Synthesis,” ACM Transactions on Graphics, vol. 36 (6), Nov. 2017, Article 235, pp. 1-11. [cited by applicant]
Pumarola A., et al., “D-NeRF: Neural Radiance Fields for Dynamic Scenes,” arXiv:2011.13961, 2020, 10 pages. [cited by applicant]
Rebain D., et al., “DeRF: Decomposed Radiance Fields,” arXiv:2011.12490, 2020, 14 pages. [cited by applicant]
Riegler G., et al., “Free View Synthesis,” European Conference on Computer Vision (ECCV), 2020, pp. 623-640. [cited by applicant]
Saito S., et al., “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization,” International Conference on Computer Vision (ICCV), 2019, pp. 2304-2314. [cited by applicant]
Saito S, et al., “PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Apr. 1, 2020, pp… [cited by applicant]
Sara U., et al., “Image Quality Assessment through FSIM, SSIM, MSE and PSNR—A Comparative Study,” Journal of Computer and Communications, 2019, vol. 7, No. 3, pp. 8-18. [cited by applicant]
Schonberger J. L., et al., “Pixelwise View Selection for Unstructured Multi-View Stereo,” European Conference on Computer Vision (ECCV), Jul. 27, 2016, 18 pages. [cited by applicant]
Schonberger J. L., et al., “Structure-from-Motion Revisited,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104-4113. [cited by applicant]
Schwarz K., et al., “GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis,” Advances in Neural Information Processing Systems (NeurIPS), 2020, vol. 33, 13 pages. [cited by applicant]
Shih M-L., et al., “3D Photography using Context-Aware Layered Depth Inpainting,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8028-8038. [cited by applicant]
Sitzmann V., et al., “Deep Voxels: Learning Persistent 3D Feature Embeddings,” Computer Vision and Pattern Recognition, Apr. 11, 2019, 10 pages. [cited by applicant]
Srinivasan P.P., et al., “Pushing the Boundaries of View Extrapolation with Multiplane Images,” Computer Vision and Pattern Recognition (CVPR), 2019, pp. 175-184. [cited by applicant]
Starck J., et al., “Surface Capture for Performance-Based Animation,” IEEE Computer Graphics and Applications, 2007, vol. 27 (3), pp. 21-31. [cited by applicant]
Tancik M., et al., “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains,” arXiv:2006.10739, 2020, 24 pages. [cited by applicant]
Tancik M., et al., “Learned Initializations for Optimizing Coordinate-Based Neural Representations,” arXiv:2012.02189, 2020, 13 pages. [cited by applicant]
Teed Z., et al., “DeepV2D: Video to Depth with Differentiable Structure from Motion,” International Conference on Learning Representations (ICLR), 2020, 20 Pages. [cited by applicant]
Tewari A., et al., “State of the Art on Neural Rendering,” State of the Art Report (STAR), 2020, vol. 39, No. 2, 27 Pages. [cited by applicant]
Tretschk E., et al., “Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Deforming Scene from Monocular Video,” arXiv:2012.12247, 2020, 9 pages. [cited by applicant]
Trevithick A., et al., “GRF: Learning a General Radiance Field for 3D Scene Representation and Rendering,” arXiv:2010.04595, 2020, 28 pages, https://arxiv.org/abs/2010.04595. [cited by applicant]
Tucker R., et al., “Single-View View Synthesis with Multiplane Images,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 551-560. [cited by applicant]
Upchurch P., et al., “From A to Z: Supervised Transfer of Style and Content using Deep Neural Network Generators,” arXiv:1603.02003, 2016, 11 pages. [cited by applicant]
Waechter M., et al., “Let there be Color! Largescale Texturing of 3D Reconstructions,” European Conference on Computer Vision (ECCV), 2014, pp. 836-850. [cited by applicant]
Adelson E. H., et al., “The Plenoptic Function and the Elements of Early Vision,” Computational Models of Visual Processing, MIT Press, 1991, pp. 3-20. [cited by applicant]
Andersson P., et al., “Flip: A Difference Evaluator for Alternating Images,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, 2020, vol. 3, No. 2, pp. 1-23. [cited by applicant]
Anonymous: “JaxNeRF,” Google, 2020, 5 pages, Retrieved from the Internet: https://github.com/google-research/google-research/tree/master/jaxnerf. [cited by applicant]
Attal B., et al., “MatryODShka: Real-Time 6DoF Video View Synthesis using Multi-Sphere Images,” European Conference on Computer Vision (ECCV), 2020, 19 Pages. [cited by applicant]
Atzmon M., et al., “SAL: Sign Agnostic Learning of Shapes from Raw Data,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2565-2574. [cited by applicant]
Ballan L., et al., “Unstructured Video-Based Rendering: Interactive Exploration of Casually Captured Videos,” ACM Transactions on Graphics (Proc. SIGGRAPH), Jul. 2010, vol. 29 (4), Article 87, 10 pages. [cited by applicant]
Bansal A., et al., “4D Visualization of Dynamic Events from Unconstrained Multi-View Videos,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 5366-5375. [cited by applicant]
Bemana M., et al., “X-Fields: Implicit Neural View-, Light- and Time-Image Interpolation,” ACM Transactions on Graphics (TOG), 2020, vol. 39, No. 6, pp. 1-15. [cited by applicant]
Bi S., et al., “Deep Reflectance vols. Relightable Reconstructions from Multi-View Photometric Images,” arXiv:2007.09892, 2020, 21 pages. [cited by applicant]
Boss M., et al., “NeRD: Neural Reflectance Decomposition from Image Collections,” arXiv:2012.03918, 2020, 15 pages. [cited by applicant]
Broxton M., et al., “Immersive Light Field Video with a Layered Mesh Representation,” ACM Transactions on Graphics, Article 86, Jul. 2020, vol. 39 (4), 15 pages. [cited by applicant]
Butler D.J., et al., “A Naturalistic Open Source Movie for Optical Flow Evaluation,” European Conference on Computer Vision (ECCV), 2012, pp. 611-625. [cited by applicant]
Carranza J., et al., “Free-Viewpoint Video of Human Actors,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2003, vol. 22, No. 3, pp. 569-577. [cited by applicant]
Chai J-X., et al., “Plenoptic Sampling,” Proceedings of the 27th annual conference on Computer graphics and Interactive techniques, 2000, pp. 307-318. [cited by applicant]
Chen S.E., et al., “View Interpolation for Image Synthesis,” In ACM SIGGRAPH Conference Proceedings, 1993, pp. 279-288. [cited by applicant]
Choi I., et al., “Extreme View Synthesis,” International Conference on Computer Vision (ICCV), 2019, pp. 7781-7790. [cited by applicant]
Collet A., et al., “High-Quality Streamable Free-Viewpoint Video,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2015, vol. 34, No. 4, pp. 1-13. [cited by applicant]
Curless B., et al., “A Volumetric Method for Building Complex Models from Range Images,” Special Interest Group on Computer Graphics, 1996, pp. 303-312. [cited by applicant]
Dabala L., et al., “Efficient Multi-Image Correspondences for On-line Light Field Video Processing,” In Computer Graphics Forum, 2016, vol. 35, No. 7, pp. 401-410. [cited by applicant]
Davis A., et al., “Unstructured Light Fields,” Computer Graphics Forum, vol. 31 (2), 2012, pp. 305-314. [cited by applicant]
Debevec P., et al., “Acquiring the Reflectance Field of a Human Face,” Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, 2000, pp. 145-156. [cited by applicant]
Debevec P.E., et al., “Modeling and Rendering Architecture from Photographs: A Hybrid Geometry-and Image-Based Approach,” Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, … [cited by applicant]
Dou M., et al., “Fusion 4D: Real-time Performance Capture of Challenging Scenes,” ACM Transactions on Graphics, Jul. 2016, vol. 35 (4), pp. 114:1-114:13. [cited by applicant]
Du Y., et al., “Neural Radiance Flow for 4D View Synthesis and Video Processing,” arXiv:2012.09790, 2020, 14 pages. [cited by applicant]
Esteban C.H., et al., “Silhouette and Stereo Fusion for 3D Object Modeling,” Computer Vision and Image Understanding, 2004, vol. 96, No. 3, pp. 367-392. [cited by applicant]
Flynn J., et al., “DeepView: View Synthesis with Learned Gradient Descent,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 2367-2376. [cited by applicant]
Furukawa Y., et al., “Accurate, Dense, and Robust Multiview Stereopsis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2009, vol. 32, No. 8, pp. 1362-1376. [cited by applicant]
Furukawa Y., et al., “Multi-View Stereo: A Tutorial,” Foundations and Trends® in Computer Graphics and Vision, 2015, vol. 9, No. 1-2, pp. 1-148. [cited by applicant]
Gao C., et al., “Flow-Edge Guided Video Completion,” European Conference on Computer Vision (ECCV), 2020, 17 Pages. [cited by applicant]
Geman S., et al., “Bayesian Image Analysis: An Application to Single Photon Emission Tomography,” American Statistical Association, 1985, pp. 12-18. [cited by applicant]
Gortler S. J., et al., “The Lumigraph,” Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive techniques, 1996, pp. 43-54. [cited by applicant]
Gu X., et al., “Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching,” Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2495-2504. [cited by applicant]
Guo K., et al., “The Relightables: Volumetric Performance Capture of Humans with Realistic Relighting,” ACM Transactions on Graphics, Article 217, vol. 38(6), Nov. 2019, pp. 1-19. [cited by applicant]
Habermann M., et al., “LiveCap: Real-Time Human Performance Capture from Monocular Video,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2019, vol. 38, No. 2, pp. 1-17. [cited by applicant]
Hedman P., et al., “Deep Blending for Free-Viewpoint Image-Based Rendering,” ACM Transactions on Graphics, Nov. 2018, vol. 37 (6), Article 257, pp. 1-15. [cited by applicant]
Huang J-B., et al., “Temporally Coherent Completion of Dynamic Video,” ACM Transactions on Graphics (Proc. SIGGRAPH Asia), 2016, vol. 35, No. 6, pp. 1-11. [cited by applicant]
Huang P-H., et al., “DeepMVS: Learning Multi-View Stereopsis,” Computer Vision and Pattern Recognition (CVPR), 2018, pp. 2821-2830. [cited by applicant]
Huang Z., et al., “Deep Volumetric Video from Very Sparse Multi-View Performance Capture,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 336-354. [cited by applicant]
Ilan S., et al., “A Survey on Data-Driven Video Completion,” Computer Graphics Forum, 2015, vol. 34, No. 6, pp. 60-85. [cited by applicant]
Innmann M., et al., “VolumeDeform: Real-Time Volumetric Non-Rigid Reconstruction,” European Conference on Computer Vision (ECCV), Jul. 30, 2016, 17 pages. [cited by applicant]
Izadi S., et al., “KinectFusion: Real-Time 3D Reconstruction and Interaction Using a Moving Depth Camera,” In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology, 2011, pp. 559-568. [cited by applicant]
Kalantari N.K., et al., “Learning-Based View Synthesis for Light Field Cameras,” ACM Transactions on Graphics (Proc. SIGGRAPH), Nov. 2016, vol. 35 (6), Article 193, 193:1-193:10, 10 pages. [cited by applicant]
Kanade T., et al., “Virtualized Reality: Constructing Virtual Worlds from Real Scenes,” IEEE Multimedia, 1997, vol. 4, No. 1, pp. 34-47. [cited by applicant]
Kaplanyan A. S., et al., “DeepFovea: Neural Reconstruction for Foveated Rendering and Video Compression Using Learned Statistics of Natural Videos,” ACM Transactions on Graphics (TOG), Nov. 8, 2019, vol. 38 (6), pp. 1-1… [cited by applicant]
Kar A., et al., “Learning a Multi-View Stereo Machine,” Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 365-376. [cited by applicant]
Kingma D.P., et al., “Adam: A Method for Stochastic Optimization,” International Conference on Learning Representations (ICLR 2015), arXiv:1412.6980v9 [cs.LG], Jan. 30, 2017, 15 pages. [cited by applicant]
Kopf J., et al., “One Shot 3D Photography,” ACM Transactions on Graphics (Proc. SIGGRAPH), 2020, vol. 39, No. 4, pp. 76:1-76:13. [cited by applicant]
Lassner C., et al., “Pulsar: Efficient Sphere-Based Neural Rendering,” arXiv:2004.07484, 2020, 13 pages. [cited by applicant]
Laurentini A., “The Visual Hull Concept for Silhouette-Based Image Understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 1994, vol. 16, No. 2, pp. 150-162. [cited by applicant]
Levoy M., et al., “Light Field Rendering,” Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, pp. 31-42. [cited by applicant]
Cited By (4)
US 12,488,483 US 12,505,512 US 12,608,882 US 12,651,406