IP Library › Granted Patent US 12,494,013
Granted Patent B2
US 12,494,013 · App. 18/211,149 · Granted Dec 9, 2025

Autodecoding latent 3D diffusion models

Inventors: Evangelos Ntavelis (Los Angeles, CA); Kyle Olszewski (Los Angeles, CA); Aliaksandr Siarohin (Los Angeles, CA); Sergey Tulyakov (Santa Monica, CA)
Assignee: Snap Inc.
G06T15/08G06N3/0455G06N3/08G06T9/001G06T15/20G06T19/20G06T2210/21G06T2219/2016G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,013
App. No.
18/211,149
Granted
Dec 9, 2025
Kind
B2
Abstract

Systems and methods for generating static and articulated 3D assets are provided that include a 3D autodecoder at their core. The 3D autodecoder framework embeds properties learned from the target dataset in the latent space, which can then be decoded into a volumetric representation for rendering view-consistent appearance and geometry. The appropriate intermediate volumetric latent space is then identified and robust normalization and de-normalization operations are implemented to learn a 3D diffusion from 2D images or monocular videos of rigid or articulated objects. The methods are flexible enough to use either existing camera supervision or no camera information at all—instead efficiently learning the camera information during training. The generated results are shown to outperform state-of-the-art alternatives on various benchmark datasets and metrics, including multi-view image datasets of synthetic objects, real in-the-wild videos of moving people, and a large-scale, real video dataset of static objects.

Claims (108)

1 . A method of training a three-dimensional (3D) diffusion model to embed properties from two-dimensional (2D) images learned from a target dataset in a latent space using an autodecoder, comprising:

processing embedding vectors of an autodecoder (G) comprising a library of embedding vectors corresponding to objects in a training dataset to generate a latent 3D feature volume;

decoding, by the autodecoder, the latent 3D feature volume into a 3D voxel grid for density and radiance representative of an object's shape and appearance;

splitting the autodecoder into a first part G 1 and a second part G 2 ; t,?

normalizing features before using features F from the latent 3D feature volume for diffusion by the 3D diffusion model, where median m is a center of distribution of the latent 3D feature volume and a Normalized InterQuartile Range (IQR) approximates a scale of the latent 3D feature volume:

training, using the autodecoder, the 3D diffusion model operating in a 3D latent space obtained from the first part G 1 using volumetric rendering of the 3D voxel grid with two-dimensional (2D) reconstruction supervision from training images in a training dataset to extract structure and appearance properties from the training dataset; and

generating, using the second part G 2 and the structure and appearance properties extracted from the training dataset, a 3D representation of the object.

2 . The method of claim 1 , further comprising progressively upsampling the latent 3D feature volume before decoding the upsampled latent 3D feature volume into the 3D voxel grid.

3 . The method of claim 1 , further comprising, during inference, denormalizing the features F from the structure and appearance properties extracted from the training dataset by the second part G 2 as F×IQR+m prior to generating the 3D representation of the object.

4 . The method of claim 1 , further comprising learning the embedding vectors by the autodecoder.

5 . The method of claim 1 , wherein decoding by the autodecoder comprises providing at least four residual blocks at each resolution in the autodecoder and using self-attention layers in a second level of resolution 8 3 and in a third level of resolution 16 3 of the autodecoder.

6 . The method of claim 1 , wherein the object is in a canonical pose and training the 3D voxel grid comprises training the 3D voxel grid using ground truth poses, poses estimated using structure from motion, or poses learned from the training dataset during training.

7 . The method of claim 6 , wherein the canonical pose comprises a canonical voxel representation of a density grid that is a discrete representation of a density field and a canonical representation of a red, green, blue (RGB) radiance field, further comprising tri-linearly interpolating density values and RGB values from the 3D voxel grid after decoding.

8 . The method of claim 1 , further comprising removing a background of the training images in the training dataset prior to training the 3D diffusion model.

9 . The method of claim 1 , wherein the object is an articulated non-rigid object, further comprising modeling a shape of the object and local motion from dynamic poses as well as a corresponding non-rigid deformation of a local region.

10 . The method of claim 9 , further comprising estimating, using a differentiable Perspective-n-Point algorithm, camera poses for each component of the non-rigid object and progressively refining estimated camera poses during training using a combination of learned 3D keypoints for each component of the non-rigid object and corresponding 2D projections predicted in each image, and combining the components with plausible deformations using a learned volumetric linear blend skinning (LBS) algorithm having skinning weights for each component of the non-rigid object that are estimated during training of the 3D diffusion model.

11 . The method of claim 1 , further comprising representing each object in the training dataset by an embedding vector comprising a concatenation of smaller embedding vectors, wherein representing each object comprises using a deterministic mapping from each training object index to its corresponding concatenated embedding vector using a hashing function where for object index k, the corresponding embedding index is:

m

⁡

(

k

)

=

[

(

a

·

k

)

⁢

mod

⁢

2

w

]

≫

(

w

-

r

)

,

for a table having 2 ′ entries where w and a are heuristic hashing parameters used to reduce a number of collisions while maintaining an appropriate table size.

12 . The method of claim 1 , further comprising decomposing a target non-rigid object into regions, where each region contains 3D keypoints and corresponding 2D projections per image, that are shared across all non-rigid objects and aligning the non-rigid objects in a learned canonical space to allow for motion transfer between the non-rigid objects.

13 . The method of claim 1 , wherein the training comprises extracting a text description of an object in the training dataset by providing a hint and a first view of the object along with a question requesting a description of a shape and color of the object for use in an inference stage to identify the object.

14 . A system that embeds properties learned from a target dataset in a latent space into a volumetric representation of an object for rendering, comprising:

a volumetric autodecoder (G) that learns embedding vectors of a library of embedding vectors corresponding to objects in a training dataset to generate a latent 3D feature volume and that decodes the latent 3D feature volume into a 3D voxel grid for density and radiance representative of an object's shape and appearance, the autodecoder comprising a first part G 1 and a second part G 2 ;

and

a 3 D diffusion model that is trained on a latent representation by the volumetric autodecoder, the 3D diffusion model operating in a 3D latent space obtained from the first part G 1 using volumetric rendering of the voxel grid with two-dimensional (2D) reconstruction supervision from training images in a training dataset to extract structure and appearance properties from the training dataset,

wherein the volumetric autodecoder normalizes features

F

^

=

(

F

-

m

)

IQR

before using features F from the latent 3D feature volume for diffusion by the 3D diffusion model, where median m is a center of distribution of the latent 3D feature volume and a Normalized InterQuartile Range (IQR) approximates a scale of the latent 3D feature volume, and

wherein the second part G 2 of the volumetric autodecoder generates a 3D representation of the object from the structure and appearance properties extracted from the training dataset.

15 . The system of claim 14 , wherein the volumetric autodecoder progressively upsamples the latent 3D feature volume and performs robust normalization on the upsampled latent 3D feature volume before training the 3D diffusion model, and wherein, during inference, the features F are denormalized from the structure and appearance properties extracted from the training dataset by the second part G 2 as {circumflex over (F)}×IQR+m generating the 3D representation of the object.

16 . The system of claim 14 , wherein the volumetric autodecoder provides at least four residual blocks at each resolution in the volumetric autodecoder and includes self-attention layers in a second level of resolution 8 3 and in a third level of resolution 16 3 .

17 . The system of claim 14 , wherein the object is an articulated non-rigid object, wherein the volumetric autodecoder models a shape of the object and local motion from dynamic poses as well as a corresponding non-rigid deformation of a local region, and wherein the volumetric autodecoder comprises a differentiable Perspective-n-Point algorithm that estimates camera poses for each component of the non-rigid object and progressively refines the estimated camera poses during training using a combination of learned 3D keypoints for each component of the non-rigid object and corresponding 2D projections predicted in each image, and a learned volumetric linear blend skinning (LBS) algorithm that combines the components with plausible deformations using skinning weights for each component of the non-rigid object that are estimated during training of the 3D diffusion model.

18 . The system of claim 14 , wherein the 3D diffusion model represents each object in the training dataset by an embedding vector comprising a concatenation of smaller embedding vectors, further comprising an encoder that encodes each object using a deterministic mapping from each training object index to its corresponding concatenated embedding vector using a hashing function where for object index k, the corresponding embedding index is:

m

⁡

(

k

)

=

[

(

a

·

k

)

⁢

mod

⁢

2

w

]

≫

(

w

-

r

)

,

for a table having 2 r entries where w and a are heuristic hashing parameters used to reduce a number of collisions while maintaining an appropriate table size.

19 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a processor cause the processor to implement a method of training a three-dimensional (3D) diffusion model to embed properties from two-dimensional (2D) images learned from a target dataset in a latent space using an autodecoder, by performing operations comprising:

processing embedding vectors of an autodecoder (G) comprising a library of embedding vectors corresponding to objects in a training dataset to generate a latent 3 D feature volume;

decoding, by the autodecoder, the latent 3D feature volume into a 3D voxel grid for density and radiance representative of an object's shape and appearance;

splitting the autodecoder into a first part G 1 and a second part G 2 ;

normalizing features

F

^

=

(

F

-

m

)

IQR

before using features F from the latent 3D feature volume for diffusion by the 3D diffusion model, where median m is a center of distribution of the latent 3D feature volume and a Normalized InterQuartile Range (IQR) approximates a scale of the latent 3D feature volume:

training, using the autodecoder, the 3D diffusion model operating in a 3D latent space obtained from the first part G 1 using volumetric rendering of the voxel grid with two-dimensional (2D) reconstruction supervision from training images in a training dataset to extract structure and appearance properties from the training dataset; and

generating, using the second part G 2 and the structure and appearance properties extracted from the training dataset, a 3D representation of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2024
From: NTAVELIS, EVANGELOS; OLSZEWSKI, KYLE; SIAROHIN, ALIAKSANDR; TULYAKOV, SERGEY
To: SNAP INC.
Reel/Frame 067586/0469 →
Continuity (1)
Related Publication 20240420407A1 · Dec 19, 2024
References Cited (108)
US 20190370965A1 · Lay · 2019 [cited by examiner]
US 20200058137A1 · Pujades · 2020 [cited by examiner]
US 20230130281A1 · Brown · 2023 [cited by examiner]
US 20240005590A1 · Martin Brualla · 2024 [cited by examiner]
US 20240013497A1 · Sun · 2024 [cited by examiner]
US 20240096001A1 · Sajjadi · 2024 [cited by examiner]
US 20240221258A1 · Chai · 2024 [cited by examiner]
US 20240371081A1 · Matthews · 2024 [cited by examiner]
US 20240371096A1 · Khamis · 2024 [cited by examiner]
Gupta, Anchit, et al. “3dgen: Triplane latent diffusion for textured mesh generation.” arXiv preprint arXiv:2303.05371 (2023). (Year: 2023). [cited by examiner]
Devadas, Daskalakis. âLecture 5: Hashing I: Chaining, Hash Functions. â MIT, 2009, courses.csail.mit.edu/6.006/fall09/lecture_notes/lecture05.pdf. (Year: 2009). [cited by examiner]
Zhu, K. P., et al. “3D CAD model search: a regularized manifold learning approach.” 2009 IEEE International Conference on Robotics and Biomimetics (ROBIO). IEEE, 2009. (Year: 2009). [cited by examiner]
Nam, Gimin, et al. “3d-Idm: Neural implicit 3d shape generation with latent diffusion models.” arXiv preprint arXiv:2212.00842 (2022). (Year: 2022). [cited by examiner]
Wang, Tengfei, et al. “Rodin: a Generative Model for Sculpting 3D Digital Avatars Using Diffusion.” arXiv preprint arXiv:2212.06135 (2022). (Year: 2022). [cited by examiner]
Ntavelis, Evangelos, et al. “Autodecoding latent 3d diffusion models.” Advances in Neural Information Processing Systems 36 (2023): 67021-67047. (Year: 2023). [cited by examiner]
Achlioptas et al.: Learning Representations and Generative Models for 3D Point Clouds. In Proceedings of the International Conference on Machine Learning, 2018. [cited by applicant]
Binkowski et al.: Demystifying MMD GANs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Bojanowski et al.: Optimizing the latent space of generative networks. In arXiv, 2017. [cited by applicant]
Brock et al.: Large scale gan training for high fidelity natural image synthesis. In arXiv, 2018. [cited by applicant]
Chan et al.: Efficient Geometry-aware 3D Generative Adversarial Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Chan et al: pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Chang et al.: ShapeNet: an Information-Rich 3D Model Repository. In arXiv, 2015. [cited by applicant]
Chen et al.: Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation. In arXiv, 2023. [cited by applicant]
Chen et al.: WaveGrad: Estimating Gradients for Waveform Generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Cheng et al.: SDFusion: Multimodal 3d shape completion, reconstruction, and generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Collins et al.: ABO: Dataset and Benchmarks for Real-World 3D Object Understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Deitke et al.: Objaverse: a Universe of Annotated 3D Objects. In arXiv, 2022. [cited by applicant]
Deng et al.: Imagenet: a large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009. [cited by applicant]
Devadas et al. MIT 6.006, Lecture 5: Hashing I: Chaining, Hash Functions, 2009. [cited by applicant]
Dhariwal et al.: Diffusion Models Beat Gans on Image Synthesis. In Proceedings of the Neural Information Processing Systems Conference, 2021. [cited by applicant]
Dockhorn et al.: Score-based generative modeling with critically-damped langevin diffusion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Falcon et al. PyTorch Lightning. GitHub. Note: https://github.com/PyTorchLightning/pytorch-lightning, 3, 2019. [cited by applicant]
Forsgren et al.: Riffusion—Stable diffusion for real-time music generation, 2022. URL https://riffusion.com/about. [cited by applicant]
Goodfellow et al.: Generative adversarial nets. In Proceedings of the Neural Information Processing Systems Conference, 2014. [cited by applicant]
Harvey et al.: Flexible Diffusion Modeling of Long Videos. In Proceedings of the Neural Information Processing Systems Conference, 2022. [cited by applicant]
He et al.: Latent video diffusion models for high-fidelity long video generation. In arXiv, 2023. [cited by applicant]
Heusel et al.: GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Proceedings of the Neural Information Processing Systems Conference, 2017. [cited by applicant]
Ho et al.: Classifier-free diffusion guidance. In arXiv, 2022. [cited by applicant]
Ho et al.: Denoising diffusion probabilistic models. In Proceedings of the Neural Information Processing Systems Conference, 2020. [cited by applicant]
Ho et al.: Video Diffusion Models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Hore et al.: Image quality metrics: Psnr vs. ssim. In Proceedings of the International Conference on Pattern Recognition, 2010. [cited by applicant]
International Search Report and Written Opinion for PCT/U2024/032242 dated Sep. 19, 2024 (Sep. 19, 2024), 9 pages. [cited by applicant]
Johnson et al.: Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In Proceedings of the European Conference on Computer Vision, 2016. [cited by applicant]
Karras et al.: Analyzing and Improving the Image Quality of StyleGAN. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020. [cited by applicant]
Karras et al.: Elucidating the Design Space of Diffusion-Based Generative Models. In Proceedings of the Neural Information Processing Systems Conference, 2022. [cited by applicant]
Karras et al.: A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019. [cited by applicant]
Karras et al.: Progressive growing of GANs for improved quality, stability, and variation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Karras et al.: Training generative adversarial networks with limited data. In arXiv, 2020. [cited by applicant]
Kingma et al.: Adam: a Method for Stochastic Optimization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015. [cited by applicant]
Kingma et al.: Auto-encoding variational bayes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014. [cited by applicant]
Kirillov et al.: Segment Anything. In arXiv, 2023. [cited by applicant]
Lepetit et al.: EPnP: an Accurate O(n) Solution to the PnP Problem. In International Journal of Computer Vision, 2009. [cited by applicant]
Lewis et al.: Pose Space Deformation: a Unified Approach to Shape Interpolation and Skeleton-Driven Deformation. In ACM Transactions on Graphics, 2000. [cited by applicant]
Lin et al.: BARF: Bundle-Adjusting Neural Radiance Fields. In Proceedings of the IEEE International Conference on Computer Vision, 2021. [cited by applicant]
Lin et al.: Robust High-Resolution Video Matting with Temporal Guidance. In Proceedings of the Winter Conference on Applications of Computer Vision, 2022. [cited by applicant]
Lin et al.: Magic3D: High-Resolution Text-to-3D Content Creation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Liu et al.: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. In arXiv, 2023. [cited by applicant]
Lorensen et al.: Marching Cubes: a High Resolution 3D Surface Construction Algorithm. In ACM Transactions on Graphics, 1987. [cited by applicant]
Mei et al.: VIDM: Video Implicit Diffusion Models. In Association for the Advancement of Artificial Intelligence Conference, 2023. [cited by applicant]
Mildenhall et al.: NeRF: Representing scenes as Neural Radiance Fields for View Synthesis. In Proceedings of the European Conference on Computer Vision, 2020. [cited by applicant]
Müller et al.: DiffRF: Rendering-Guided 3D Radiance Field Diffusion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Nagrani et al.: VoxCeleb: Large-scale speaker verification in the wild. Computer Science and Language, 2019. [cited by applicant]
Nam et al: “3D-LDM: Neural Implicit 3D Shape Generation with Latent Diffusion Models”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 15, 2022 (Dec. 15, 2022), XP091384… [cited by applicant]
Nguyen-Phuoc et al.: HoloGAN: Unsupervised Learning of 3D Representations From Natural Images. In Proceedings of the IEEE International Conference on Computer Vision, 2019. [cited by applicant]
Nguyen-Phuoc et al.; Blockgan: Learning 3d object-aware scene representations from unlabelled images. In arXiv, 2020. [cited by applicant]
Nichol et al.: Improved denoising diffusion probabilistic models. In ICML, 2021. [cited by applicant]
Niemeyer et al.: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Ntavelis et al.: StyleGenes: Discrete and Efficient Latent Distributions for GANs. In arXiv, 2023. [cited by applicant]
Ntavelis, Evangelos et al. “Autodecoding Latent 3D Diffusion Models,” arXiv:2307.05445v1 [cs.CV] Jul. 7, 2023z. [cited by applicant]
Obukhov et al.: High-fidelity performance metrics for generative models in PyTorch, 2020. URL https://github.com/toshas/torch-fidelity. Version: 0.3.0, DOI: 10.5281/zenodo.4957738. [cited by applicant]
Park et al.: PhotoShape: Photorealistic Materials for Large-Scale Shape Collections. In ACM Transactions on Graphics, 2018. [cited by applicant]
Park et al: “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019 (Jun. 15, 2019), pp. 165-174… [cited by applicant]
Paszke et al.: Automatic Differentiation in PyTorch, 2017. [cited by applicant]
Paszke et al.: PyTorch: an Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Neural Information Processing Systems Conference, 2019. [cited by applicant]
Poole et al.: Dreamfusion: Text-to-3d using 2d diffusion. In arXiv, 2022. [cited by applicant]
Raffel et al.: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. In The Journal of Machine Learning Research, 2020. [cited by applicant]
Ravi et al.: Accelerating 3D Deep Learning with PyTorch3D. In arXiv, 2020. [cited by applicant]
Rombach et al.: High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Schönberger et al.: Structure-from-Motion Revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016. [cited by applicant]
Schwarz et al.: GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis. In Proceedings of the Neural Information Processing Systems Conference, 2020. [cited by applicant]
Siarohin et al.: First Order Motion Model for Image Animation. In Proceedings of the Neural Information Processing Systems Conference, 2019. [cited by applicant]
Siarohin et al.: Unsupervised Volumetric Animation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Siarohin et al: “Unsupervised Volumetric Animation”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 26, 2023 (Jan. 26, 2023), XP091421870. [cited by applicant]
Simonyan et al.: Very deep convolutional networks for large-scale image recognition. In arXiv, 2014. [cited by applicant]
Skorokhodov et al.: 3D Generation on ImageNet. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Skorokhodov et al.: Stylegan-v: a continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Skorokhodov et al.: EpiGRAF: Rethinking Training of 3D GANs. In Proceedings of the Neural Information Processing Systems Conference, 2022. [cited by applicant]
Sohl-Dickstein et al.: Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. [cited by applicant]
Song et al.: Denoising Diffusion Implicit Models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Song et al.: Generative modeling by estimating gradients of the data distribution. In Proceedings of the Neural Information Processing Systems Conference, 2019. [cited by applicant]
Song et al.: Score-based generative modeling through stochastic differential equations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Tian et al.: A good image generator is what you need for high-resolution video synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. [cited by applicant]
Vahdat et al.: Score-based generative modeling in latent space. In Proceedings of the Neural Information Processing Systems Conference, 2021. [cited by applicant]
Vaswani et al.: Attention is all you need. In Proceedings of the Neural Information Processing Systems Conference, 2017. [cited by applicant]
Voleti et al.: MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation. In Proceedings of the Neural Information Processing Systems Conference, 2022. [cited by applicant]
Wang et al.: NeRF—: Neural Radiance Fields Without Known Camera Parameters. In arXiv, 2021. [cited by applicant]
Weng et al.: HumanNeRF: Free-Viewpoint Rendering of Moving People from Monocular Video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Whaley: The Interquartile Range: Theory and Estimation. PhD thesis, East Tennessee State University, 2005. [cited by applicant]
Xiao et al.: Tackling the generative learning trilemma with denoising diffusion GANs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Xu et al.: Pose for Everything: Towards Category-Agnostic Pose Estimation. In Proceedings of the European Conference on Computer Vision, 2022. [cited by applicant]
Xue et al.: GIRAFFE HD: a High-Resolution 3D-aware Generative Model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Yin et al.: NUWA-XL: Diffusion over Diffusion for extremely Long Video Generation. In arXiv, 2023. [cited by applicant]
Yu et al.: CelebV-Text: a Large-Scale Facial Text-Video Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Yu et al.: Generating videos with dynamics-aware implicit generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Yu et al.: MVImgNet: a Large-scale Dataset of Multi-view Images. In arXiv, 2023. [cited by applicant]
Zhang et al.: The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Zhu et al.: Discrete contrastive diffusion for cross-modal music and image generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. [cited by applicant]
Zhu et al.: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In arXiv, 2023. [cited by applicant]