IP Library › Granted Patent US 12,450,823
Granted Patent B2
US 12,450,823 · App. 18/515,024 · Granted Oct 21, 2025

Neural dynamic image-based rendering

Inventors: Keith Noah Snavely (New York, NY); Zhengqi Li (Jersey City, NJ); Forrester H. Cole (Cambridge, MA); Richard Tucker (New York, NY); Qianqian Wang (Berkeley, CA)
Assignee: Google LLC
G06T15/20G06T7/215G06T11/001G06T15/06G06V10/44G06V10/56H04N13/117G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,823
App. No.
18/515,024
Filed
Nov 20, 2023
Granted
Oct 21, 2025
Kind
B2
Art Unit
2619
USPC
345/419
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for rendering a new image that depicts a scene from a perspective of a camera at a new camera viewpoint at a given time point in a video.

Claims (65)

1. A method performed by one or more computers, the method comprising:

receiving a video of a scene comprising a plurality of images at respective time points;

receiving a query specifying a particular time point and a new camera viewpoint; and

generating, using a view synthesis machine learning model and the video of the scene, a new image of the scene that appears to be taken from the new camera viewpoint at the particular time point, comprising:

generating, based on the particular time point, a set of source images that comprises one or more images from the video;

generating respective features for each of the source images;

for each of a plurality of pixels of the new image:

sampling a plurality of three-dimensional points along a ray corresponding to the pixel;

for each sampled point, generating, using a first neural network within the view synthesis machine learning model, data defining a motion trajectory of the sampled point around the particular time point;

generating, from the respective features for the source images, respective features for each of the sampled points using the data defining the motion trajectory of the sampled point; and

generating, from the respective features of each of the sampled points, a final color of the pixel in the new image.

2. The method of claim 1 , further comprising:

providing the new image for presentation.

3. The method of claim 1 , wherein generating the set of source images comprises:

identifying, as source images, each image that is at a respective time point that is within a temporal radius of the particular time point.

4. The method of claim 1 , wherein generating respective features for each of the source images comprises, for each source image, processing the source image using an encoder neural network.

5. The method of claim 1 , wherein generating, using a first neural network within the view synthesis machine learning model, data defining a motion trajectory of the sampled point around the particular time point comprises:

processing an input comprising a positional encoding of the sampled point and a positional encoding of the particular time point using the first neural network to generate respective coefficients for each of a plurality of basis functions that define the motion trajectory.

6. The method of claim 5 , wherein the motion trajectory is defined by the respective coefficients for the basis functions and a respective global coefficient for each of the basis functions that is independent of the particular time point and the sampled point and that is learned during the training of the view synthesis machine learning model.

7. The method of claim 1 , wherein generating, from the respective features for the source images, respective features for each of the sampled points using the data defining the motion trajectory of the sampled point comprises:

for each source image, identifying using the motion trajectory of the sampled point, a corresponding pixel location in the source image; and

for each source image, generating the respective features for the sampled point from a color of the corresponding pixel location in the source image and a feature vector corresponding to the corresponding pixel location in the respective features for the source image.

8. The method of claim 7 , wherein generating, from the respective features of the sampled points, a final color of the pixel in the new image further comprises:

generating an aggregated feature for the sampled point from the respective features for the sampled point for each of the source images.

9. The method of claim 8 , wherein generating an aggregated feature for the sampled point from the respective features for the sampled point for each of the source image comprises:

for each source image, processing the respective features for the sampled point for the source image using a second neural network within the view synthesis machine learning model to generate an output; and

applying a pooling operation to the outputs for the source images to generate the aggregated features.

10. The method of claim 1 , wherein generating, from the respective features of each of the sampled points, a final color of the pixel in the new image comprises:

processing the respective features for the sampled points using a third neural network within the view synthesis machine learning model to generate a respective color and volumetric density for each sampled point.

11. The method of claim 10 , wherein the third neural network is a Transformer neural network.

12. The method of claim 10 , wherein generating, from the respective features of each of the sampled points, a final color of the pixel in the new image further comprises:

generating a first color for the pixel from the respective colors and volumetric densities for the sampled points.

13. The method of claim 12 , wherein generating, from the respective features of each of the sampled points, a final color of the pixel in the new image further comprises:

generating a second color for the pixel using a time-invariant model; and

combining the first and second colors to generate the final color of the pixel.

14. The method of claim 12 , wherein generating a first color for the pixel from the respective colors and volumetric densities for the sampled points comprises:

applying volume rendering to the respective colors and volumetric densities for the sampled points to generate the final color.

15. The method of claim 1 , further comprising:

prior to generating, using the view synthesis machine learning model and the video of the scene, the new image of the scene that appears to be taken from the new camera viewpoint at the particular time point:

training the view synthesis machine learning model on the images in the video of the scene to minimize a loss function.

16. The method of claim 15 , wherein the loss function comprises a photometric consistency loss.

17. The method of claim 16 , wherein the photometric consistency loss measures, for a given target image in the video and a given source image in the video and for each pixel in the given target image, an error between an actual color at the pixel in the given target image and a color generated at the pixel by cross-time rendering using the given source image.

18. The method of claim 15 , wherein the loss function comprises a segmentation mask loss that uses a motion mask derived through image-based motion segmentation.

19. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

receiving a video of a scene comprising a plurality of images at respective time points;

receiving a query specifying a particular time point and a new camera viewpoint; and

generating, using a view synthesis machine learning model and the video of the scene, a new image of the scene that appears to be taken from the new camera viewpoint at the particular time point, comprising:

generating, based on the particular time point, a set of source images that comprises one or more images from the video;

generating respective features for each of the source images;

for each of a plurality of pixels of the new image:

sampling a plurality of three-dimensional points along a ray corresponding to the pixel;

for each sampled point, generating, using a first neural network within the view synthesis machine learning model, data defining a motion trajectory of the sampled point around the particular time point;

generating, from the respective features for the source images, respective features for each of the sampled points using the data defining the motion trajectory of the sampled point; and

generating, from the respective features of each of the sampled points, a final color of the pixel in the new image.

20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a video of a scene comprising a plurality of images at respective time points;

receiving a query specifying a particular time point and a new camera viewpoint; and

generating, using a view synthesis machine learning model and the video of the scene, a new image of the scene that appears to be taken from the new camera viewpoint at the particular time point, comprising:

generating, based on the particular time point, a set of source images that comprises one or more images from the video;

generating respective features for each of the source images;

for each of a plurality of pixels of the new image:

sampling a plurality of three-dimensional points along a ray corresponding to the pixel;

for each sampled point, generating, using a first neural network within the view synthesis machine learning model, data defining a motion trajectory of the sampled point around the particular time point;

generating, from the respective features for the source images, respective features for each of the sampled points using the data defining the motion trajectory of the sampled point; and

generating, from the respective features of each of the sampled points, a final color of the pixel in the new image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2023
From: SNAVELY, KEITH NOAH; LI, ZHENGQI; COLE, FORRESTER H.; TUCKER, RICHARD; WANG, QIANQIAN
To: GOOGLE LLC
Reel/Frame 065750/0272 →
Continuity (2)
Provisional Application 63598044 · Nov 10, 2023
Related Publication 20250157133A1 · May 15, 2025
References Cited (100)
US 10614581B2 · Senthamil · 2020 [cited by examiner]
US 11210859B1 · Fredericks · 2021 [cited by examiner]
US 20220113795A1 · Eder · 2022 [cited by examiner]
US 20240007585A1 · Assouline · 2024 [cited by examiner]
US 20240046516A1 · Anciukevicius · 2024 [cited by examiner]
US 20250085108A1 · Coimbra De Andrade · 2025 [cited by examiner]
Aguiar et al., “Performance capture from sparse multi-view video,” ACM SIGGRAPH 2008 papers, Aug. 2008, pp. 1-10. [cited by applicant]
Bansal et al., “4D visualization of dynamic events from unconstrained multi-view videos,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 5366-5375. [cited by applicant]
Barron et al., “Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 5470-5479. [cited by applicant]
Bemana et al., “X-fields: Implicit neural view-, light- and time-image interpolation,” ACM Transactions on Graphics (TOG), Dec. 2020, 39(6):1-15. [cited by applicant]
Bi et al., “Neural reflectance fields for appearance acquisition,” CoRR, submitted on Aug. 16, 2020, arXiv:2008.03824v2, 11 pages. [cited by applicant]
Bozic et al., “Deepdeform: Learning non-rigid RGB-D reconstruction with semi-supervised data,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 7002-7012. [cited by applicant]
Broxton et al., “Immersive light field video with a layered mesh representation,” ACM Transactions on Graphics (TOG), Jul. 2020, 39(4):1-15. [cited by applicant]
Buehler et al., “Unstructured lumigraph rendering,” Proceedings of the 28th annual conference on Computer graphics and interactive techniques, Aug. 2001, pp. 425-432. [cited by applicant]
Carranza et al., “Free-viewpoint video of human actors,” ACM Transactions on Graphics (TOG), Jul. 2003, 22(3):569-577. [cited by applicant]
Chai et al., “Plenoptic sampling,” Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), Jul. 2000, pp. 307-318. [cited by applicant]
Charbonnier et al., “Two deterministic half-quadratic regularization algorithms for computed imaging,” Proceedings of 1st International Conference on Image Processing, Nov. 1994, pp. 168-172. [cited by applicant]
Chen et al., “MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, pp. 14124-14133. [cited by applicant]
Choi et al., “Extreme view synthesis,” Proceeding of International Conference on Computer Vision (ICCV), Oct.-Nov. 2019, pp. 7781-7790. [cited by applicant]
Debevec et al., “Modeling and rendering architecture from photographs: A hybrid geometry- and image-based approach,” Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, Aug. 1996, … [cited by applicant]
Dou et al., “Fusion4D: real-time performance capture of challenging scenes,” ACM Transaction on Graphics (TOG), Jul. 2016, 35(4):114:1-114:13. [cited by applicant]
Du et al., “Neural radiance flow for 4d view synthesis and video processing,” Proceeding of 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, Oct. 10-17, 2021, pp. 14304-14314. [cited by applicant]
Flynn et al., “DeepStereo: Learning to predict new views from the world's imagery,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 5515-5524. [cited by applicant]
Flynn et al., “DeepView: View synthesis with learned gradient descent,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 2367-2376. [cited by applicant]
Fridovich-Keil et al., “Plenoxels: Radiance fields without neural networks,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 19-24, 2022, pp. 5501-5510. [cited by applicant]
Gao et al., “Dynamic View Synthesis from Dynamic Monocular Video,” CoRR, submitted on May 13, 2021, arXiv:2105.06468v1, 11 pages. [cited by applicant]
Gao et al., “Dynamic view synthesis from dynamic monocular video,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 11-17, 2021, pp. 5712-5721. [cited by applicant]
Gao et al., “Monocular dynamic view synthesis: A reality check,” CoRR, submitted on Oct. 24, 2022, arXiv:2210.13445c1, 13 pages. [cited by applicant]
Gortler et al., “The lumigraph,” Proceedings of the 23rd annual conference on Computer graphics and interactive techniques (SIGGRAPH 96), Aug. 4-9, 1996, pp. 43-54. [cited by applicant]
Guo et al., “The relightables: Volumetric performance capture of humans with realistic relighting,” ACM Transactions on Graphics (ToG), Nov. 2019, 38(6):1-19. [cited by applicant]
He et al., “Deep residual learning for image recognition,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 27-30, 2016, pp. 770-778. [cited by applicant]
Hedman et al., “Deep blending for free-viewpoint image-based rendering,” ACM Transaction on Graphics (TOG), Nov. 2018, 37(6):1-15. [cited by applicant]
Hedman et al., “Scalable inside-out image-based rendering,” ACM Transactions on Graphics (TOG), Nov. 2016, 35(6):1-11. [cited by applicant]
Innmann et al., “VolumeDeform: Real-time volumetric non-rigid reconstruction,” Proceeding of European Conference on Computer Vision (ECCV), Oct. 11-14, 2016, pp. 1-17. [cited by applicant]
Jia, “Plücker coordinates for lines in the space,” Problem Solver Techniques for Applied Computer Science, Com-S-477/577 Course Handout, Aug. 30, 2020, 11 pages. [cited by applicant]
Kalantari et al., “Learning-based view synthesis for light field cameras,” ACM Transaction on Graphics, Nov. 2016, 35(6):1-10. [cited by applicant]
Kanade et al., “Virtualized reality: Constructing virtual worlds from real scenes,” IEEE multimedia, Jan. 1997, 4(1):34-47. [cited by applicant]
Kasten et al., “Layered neural atlases for consistent video editing,” ACM Transactions on Graphics (TOG), Dec. 2021, 40(6):1-12. [cited by applicant]
Ke et al., “Mask transfiner for high-quality instance segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 19-24, 2022, pp. 4412-4421. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, submitted on Dec. 22, 2014, arXiv:1412.6980v1, 9 pages. [cited by applicant]
Kopf et al., “First-person hyper-lapse videos,” ACM Transactions on Graphics (TOG), Jul. 27, 2014, 33(4):1-10. [cited by applicant]
Kopf et al., “Robust consistent video depth estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, pp. 1611-1621. [cited by applicant]
Levoy et al., “Light field rendering,” Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, Aug. 1, 1996, pp. 31-42. [cited by applicant]
Li et al., “Infinitenature-zero: Learning perpetual view generation of natural scenes from single images,” Proceeding of European Conference on Computer Vision, Oct. 23-27, 2022, pp. 515-534. [cited by applicant]
Li et al., “Learning the depths of moving people by watching frozen people,” In Proc. Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 4521-4530. [cited by applicant]
Li et al., “Neural 3D video synthesis from multi-view video,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, pp. 5521-5531. [cited by applicant]
Li et al., “Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes,” CoRR, submitted on Apr. 21, 2021, arXiv:2011.13084v3, 11 pages. [cited by applicant]
Li et al., “Neural scene flow fields for space-time view synthesis of dynamic scenes,” Proceedings of Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 6498-6508. [cited by applicant]
Lin et al., “Deep 3D mask volume for view synthesis of dynamic scenes,” Proceeding of International Conference on Computer Vision (ICCV), Oct. 10-17, 2021, pp. 1749-1758. [cited by applicant]
Liu et al., “Neural rays for occlusion-aware image-based rendering,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 19-24, 2022, pp. 7824-7833. [cited by applicant]
Liu et al., “Neural sparse voxel fields,” Advances in Neural Information Processing Systems, Jul. 22, 2020, 33:15651-15663. [cited by applicant]
Lombardi et al., “Neural volumes: Learning dynamic renderable volumes from images,” ACM Transaction on Graphics (TOG), Jul. 2019, 38(4):65:1-65:14. [cited by applicant]
Loper et al., “SMPL: A skinned multiperson linear model,” ACM transactions on graphics (TOG), Nov. 2015, 34(6):248.1-248.16. [cited by applicant]
Lu et al., “Layered neural rendering for retiming people in video,” ACM Transaction on Graphics (TOG), Dec. 2020, 39(6):256:1-256:14. [cited by applicant]
Lu et al., “Omnimatte: Associating objects and their effects in video,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 4507-4515. [cited by applicant]
Luo et al., “Consistent video depth estimation,” ACM Transactions on Graphics (ToG), Jul. 2020, 39(4):71-1. [cited by applicant]
Martin-Brualla et al., “NeRF in the wild: Neural radiance fields for unconstrained photo collections,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 7210-7219. [cited by applicant]
Mildenhall et al., “NeRF: Representing scenes as neural radiance fields for view synthesis,” Proceeding of European Conf. on Computer Vision (ECCV), Nov. 3, 2020, pp. 405-421. [cited by applicant]
Muller et al., “Instant neural graphics primitives with a multiresolution hash encoding,” CoRR, submitted on May 4, 2022, arXiv:2201.05989v2, 15 pages. [cited by applicant]
Newcombe et al., “DynamicFusion: Reconstruction and tracking of non-rigid scenes in real-time,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2015, pp. 343-352. [cited by applicant]
Niemeyer et al., “Differentiable volumetric rendering: Learning implicit 3D representations without 3D supervision,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13-19, 2020, p… [cited by applicant]
Park et al., “HyperNeRF: A higherdimensional representation for topologically varying neural radiance fields,” CoRR, submitted on Sep. 10, 2021, arXiv:2106.13228v2, 16 pages. [cited by applicant]
Park et al., “Nerfies: Deformable neural radiance fields,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2021, pp. 5865-5874. [cited by applicant]
Peng et al., “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun… [cited by applicant]
Penner et al., “Soft 3d reconstruction for view synthesis,” ACM Transactions on Graphics (TOG), Nov. 2017, 36(6):235.1-235.11. [cited by applicant]
Poole et al., “Dreamfusion: Text-to-3d using 2d diffusion,” CoRR, submitted on Sep. 29, 2022, arXiv:2209.14988v1, 18 pages. [cited by applicant]
Pumarola et al., “D-NeRF: Neural radiance fields for dynamic scenes,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, pp. 10318-10327. [cited by applicant]
Ranftl et al., “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Mar. 2022, 44(3):1623-1637. [cited by applicant]
Riegler et al., “Free view synthesis,” Proceeding of European Conf. on Computer Vision (ECCV), Nov. 13, 2020, pp. 623-640. [cited by applicant]
Riegler et al., “Stable view synthesis,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, pp. 12216-12225. [cited by applicant]
Saito et al., “PIFu: Pixel-aligned implicit function for high-resolution clothed human digitization,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 27, 2019-Nov. 2, 2019, pp. 2304-2314. [cited by applicant]
Schonberger et al., “Structurefrom-motion revisited,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 4104-4113. [cited by applicant]
Shum et al., “A Review of image-based rendering techniques,” Visual Communications and Image Processing 2000, May 30, 2000, pp. 2-13. [cited by applicant]
Sitzmann et al., “Deepvoxels: Learning persistent 3D feature embeddings,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 2437-2446. [cited by applicant]
Sitzmann et al., “Implicit neural representations with periodic activation functions,” Advances in Neural Information Processing Systems, Dec. 6, 2020, 33(626):7462-7473. [cited by applicant]
Sitzmann et al., “Scene representation networks: Continuous 3D-structureaware neural scene representations,” In Neural Information Processing Systems, Dec. 2019, pp. 1119-1130. [cited by applicant]
Song et al., “NeRFPlayer: A streamable dynamic scene representation with decomposed neural radiance fields,” CoRR, submitted on Oct. 28, 2022, arXiv:2210.15947v1, 15 pages. [cited by applicant]
Srinivasan et al., “Pushing the boundaries of view extrapolation with multiplane images,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 175-184. [cited by applicant]
Stich et al., “View and time interpolation in image space,” Computer Graphics Forum, Oct. 2008, 27(7):1781-1787. [cited by applicant]
Suhail et al., “Light field neural rendering,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2022, 8269-8279. [cited by applicant]
Tancik et al., “Fourier features let networks learn high frequency functions in low dimensional domains,” Advances in Neural Information Processing Systems, Apr. 2020, 33:7537-7547. [cited by applicant]
Teed et al., “RAFT: Recurrent all-pairs field transforms for optical flow,” Proceeding of European Conference on Computer Vision (ECCV), Nov. 3, 2020, 12347(1):402-419. [cited by applicant]
Tretschk et al., “Nonrigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 10-17, 2… [cited by applicant]
Ulyanov et al., “Instance normalization: The missing ingredient for fast stylization,” CoRR, submitted on Sep. 20, 2016, arXiv:1607.08022v2, 6 pages. [cited by applicant]
Wang et al., “Fourier plenoctrees for dynamic radiance field rendering in real-time,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, pp. 13524-13534. [cited by applicant]
Wang et al., “Ibrnet: Learning multi-view image-based rendering,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 19-25, 2021, pp. 4690-4699. [cited by applicant]
Wang et al., “Neural prior for trajectory estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2022, pp. 6532-6542. [cited by applicant]
Wang et al., “Neural trajectory fields for dynamic novel view synthesis,” CoRR, submitted on May 12, 2021, arXiv:2105.05994v1, 12 pages. [cited by applicant]
Weng et al., “HumanNeRF: Free-viewpoint rendering of moving people from monocular video,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2022, pp. 16210-16220. [cited by applicant]
Wizadwongsa et al., “NeX: Real-time view synthesis with neural basis expansion,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 8534-8543. [cited by applicant]
Wu et al., “D2NeRF: Self-supervised decoupling of dynamic and static objects from a monocular video,” CoRR, Submitted on May 31, 2022, arXiv:2205.15838v1, 20 pages. [cited by applicant]
Xian et al., “Space-time neural irradiance fields for free-viewpoint video,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, pp. 9421-9431. [cited by applicant]
Yoon et al., “Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 5336-5345. [cited by applicant]
Zhang et al., “Consistent depth of moving objects in video,” ACM Transactions on Graphics (TOG), Aug. 2021, 40(4):148:1-148:12. [cited by applicant]
Zhang et al., “Editable free-viewpoint video using a layered neural representation,” ACM Transactions on Graphics (TOG), Aug. 2021, 40(4):149:1-149:18. [cited by applicant]
Zhang et al., “Structure and motion from casual videos,” European Conference on Computer Vision, Oct. 23, 2022, pp. 20-37. [cited by applicant]
Zhang et al., “The unreasonable effectiveness of deep features as a perceptual metric,” Proceeding of Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 586-595. [cited by applicant]
Zhou et al., “Stereo magnification: learning view synthesis using multiplane images,” ACM Transactions on Graphics (TOG), Aug. 2018, 37(4):65:1-65:12. [cited by applicant]
Zitnick et al., “High-quality video view interpolation using a layered representation,” ACM Transaction on Graphics (TOG), Aug. 2004, 23(3):600-608. [cited by applicant]
Zollhöfer et al., “Real-time non-rigid reconstruction using an RGB-D camera,” ACM Transactions on Graphics (ToG), Jul. 27, 2014, 33(4):1-12. [cited by applicant]