IP Library › Granted Patent US 12,400,388
Granted Patent B2
US 12,400,388 · App. 18/089,984 · Granted Aug 26, 2025

Unsupervised volumetric animation

Inventors: Menglei Chai (Los Angeles, CA); Hsin-Ying Lee (San Jose, CA); Willi Menapace (Santa Monica, CA); Kyle Olszewski (Los Angeles, CA); Jian Ren (Hermosa Beach, CA); Aliaksandr Siarohin (Los Angeles, CA); Ivan Skorokhodov (Los Angeles, CA); Sergey Tulyakov (Santa Monica, CA)
Assignee: Snap Inc.
G06T13/40G06T7/70G06T19/20G06T2207/20081G06T2207/20084G06T2207/30201G06T2219/2004G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,388
App. No.
18/089,984
Granted
Aug 26, 2025
Kind
B2
Abstract

Unsupervised volumetric 3D animation (UVA) of non-rigid deformable objects without annotations learns the 3D structure and dynamics of objects solely from single-view red/green/blue (RGB) videos and decomposes the single-view RGB videos into semantically meaningful parts that can be tracked and animated. Using a 3D autodecoder framework, paired with a keypoint estimator via a differentiable perspective-n-point (PnP) algorithm, the UVA model learns the underlying object 3D geometry and parts decomposition in an entirely unsupervised manner from still or video images. This allows the UVA model to perform 3D segmentation, 3D keypoint estimation, novel view synthesis, and animation. The UVA model can obtain animatable 3D objects from a single or a few images. The UVA method also features a space in which all objects are represented in their canonical, animation-ready form. Applications include the creation of lenses from images or videos for social media applications.

Claims (110)

1. An unsupervised volumetric animation system for three-dimensional (3D) animation of a non-rigid deformable object, comprising:

a canonical voxel generator to produce a 3D volumetric representation of the non-rigid deformable object in a canonical pose parameterized as a voxel grid, wherein the non-rigid deformable object is represented as a set of moving rigid parts, and assigns each 3D point of the non-rigid deformable object to a corresponding moving rigid part of the non-rigid deformable object;

a two-dimensional (2D) keypoint predictor to estimate a pose, in a given image frame, of each moving rigid part of an input object to be animated;

a volumetric skinning algorithm to map a canonical object volume of the non-rigid deformable object into a deformed volume that represents, as a deformed object, the input object to be animated with the pose in a current frame; and

a volumetric renderer to render the deformed object as an image of the input object.

2. The system of claim 1 , wherein the input object to be animated is extracted from a video or a still image.

3. The system of claim 1 , wherein the 2D keypoint predictor uses a pose extracted from the input object to be animated to predict a set of 2D keypoints that correspond to 3D keypoints of the input object to be animated.

4. The system of claim 3 , wherein the volumetric renderer takes a deformed density and radiance of the deformed volume produced via volumetric skinning using a canonical density (V DENSITY ) of the non-rigid deformable object, a radiance of the non-rigid deformable object, a set of poses for different moving rigid parts of the input object to be animated, and moving rigid parts of the input object to be animated represented as linear blend skinning (LBS) weights.

5. The system of claim 4 , wherein the volumetric renderer volumetrically renders the deformed radiance to produce the image.

6. The system of claim 1 , wherein the 2D keypoint predictor estimates the pose of each moving rigid part by learning a set of 3D keypoints in a canonical space and comprises a 2D convolutional neural network that detects 2D projections of the moving rigid part to provide a set of corresponding 2D keypoints in a current frame.

7. The system of claim 6 , further comprising a perspective-n-point (PnP) algorithm that processes a differentiable PnP formulation to recover the pose of each moving rigid part from corresponding 2D keypoints and 3D keypoints.

8. The system of claim 7 , wherein the 2D keypoint predictor introduces N k learnable canonical 3D keypoints for each moving rigid part, shares 3D keypoints K p 3D of the moving rigid part among objects in a dataset, defines a 2D keypoints prediction network C that takes frame F i as input and outputs 2D keypoints K p 2D for each part p, where each 2D keypoint corresponds to its respective 3D keypoint, and recovers the pose of moving rigid part p as:

T

p

-

1

=

PnP

⁡

(

K

p

2

⁢

D

,

K

p

3

⁢

D

)

=

PnP

⁡

(

C

⁡

(

F

i

)

,

K

p

3

⁢

D

)

.

9. A method of providing three-dimensional (3D) animation of a non-rigid deformable object, comprising:

producing, using a canonical voxel generator, a 3D volumetric representation of the non-rigid deformable object in a canonical pose parameterized as a voxel grid, wherein the non-rigid deformable object is represented as a set of moving rigid parts;

assigning, by the canonical voxel generator, each 3D point of the non-rigid deformable object to a corresponding moving rigid part of the non-rigid deformable object;

estimating, by a two-dimensional (2D) keypoint predictor, a pose, in a given image frame, of each moving rigid part of an input object to be animated;

mapping, by a volumetric skinning algorithm, a canonical object volume of the non-rigid deformable object into a deformed volume that represents, as a deformed object, the input object to be animated with the pose in a current frame; and

rendering, by a volumetric renderer, the deformed object as an image of the input object.

10. The method of claim 9 , further comprising extracting the input object to be animated from a video or a still image.

11. The method of claim 9 , wherein the assigning comprises learning, for each moving rigid part, a set of canonical 3D keypoints during training.

12. The method of claim 9 , further comprising using, by the 2D keypoint predictor, a pose extracted from the input object to be animated to predict a set of 2D keypoints that correspond to 3D keypoints of the input object to be animated.

13. The method of claim 12 , wherein the mapping comprises the volumetric renderer taking a deformed density and radiance of the deformed volume produced via volumetric skinning using a canonical density (V DENSITY ) of the non-rigid deformable object, a radiance of the non-rigid deformable object, a set of poses for different moving rigid parts of the input object to be animated, and moving rigid parts of the input object to be animated represented as linear blend skinning (LBS) weights.

14. The method of claim 13 , wherein the rendering comprises volumetrically rendering the deformed radiance by the volumetric renderer to produce the image.

15. The method of claim 9 , wherein estimating the pose of each moving rigid part comprises learning a set of 3D keypoints in a canonical space and detecting 2D projections of the moving rigid part to provide a set of corresponding 2D keypoints in a current frame using a 2D convolutional neural network.

16. The method of claim 15 , wherein the estimating the pose of each moving rigid part further comprises using, by a perspective-n-point (PnP) algorithm, a differentiable PnP formulation to recover the pose of each moving rigid part from corresponding 2D keypoints and 3D keypoints.

17. The method of claim 16 , further comprising introducing N k learnable canonical 3D keypoints for each moving rigid part, sharing 3D keypoints K p 3D of the moving rigid part among objects in a dataset, defining a 2D keypoints prediction network C that takes frame F i as input and outputs 2D keypoints K p 2D for each part p, where each 2D keypoint corresponds to its respective 3D keypoint, and recovering the pose of moving rigid part p as:

T

p

-

1

=

PnP

⁡

(

K

p

2

⁢

D

,

K

p

3

⁢

D

)

=

PnP

⁡

(

C

⁡

(

F

i

)

,

K

p

3

⁢

D

)

.

18. The method of claim 17 , further comprising sharing the 3D keypoints for all the objects in the dataset, whereby all objects in the dataset share a same canonical space for poses.

19. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a processor cause the processor to animate a three-dimensional (3D) non-rigid deformable object extracted from a video or a still image by performing operations comprising:

producing, using a canonical voxel generator, a 3D volumetric representation of the non-rigid deformable object in a canonical pose parameterized as a voxel grid, wherein the non-rigid deformable object is represented as a set of moving rigid parts;

assigning, by the canonical voxel generator, each 3D point of the non-rigid deformable object to a corresponding moving rigid part of the non-rigid deformable object;

estimating, by a two-dimensional (2D) keypoint predictor, a pose, in a given image frame, of each moving rigid part of an input object to be animated;

mapping, by a volumetric skinning algorithm, a canonical object volume of the non-rigid deformable object into a deformed volume that represents, as a deformed object, the input object to be animated with the pose in a current frame; and

rendering, by a volumetric renderer, the deformed object as an image of the input object.

20. The medium of claim 19 , wherein the instructions for assigning each 3D point of the non-rigid deformable object to the corresponding moving rigid part of the non-rigid deformable object comprises executing instructions to learn, for each moving rigid part, a set of canonical 3D keypoints during training.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2025
From: CHAI, MENGLEI; LEE, HSIN-YING; MENAPACE, WILLI; OLSZEWSKI, KYLE; REN, JIAN; SIAROHIN, ALIAKSANDR; SKOROKHODOV, IVAN; TULYAKOV, SERGEY
To: SNAP INC.
Reel/Frame 071865/0196 →
Continuity (1)
Related Publication 20240221258A1 · Jul 4, 2024
References Cited (15)
US 8139067B2 · Anguelov · 2012 [cited by examiner]
US 20170071508A1 · Kaiser · 2017 [cited by examiner]
US 20210248772A1 · Iqbal · 2021 [cited by examiner]
Chen, X., Jiang, T., Song, J., Rietmann, M., Geiger, A., Black, M. J., & Hilliges, O. (2023). Fast-SNARF: A fast deformer for articulated neural fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, … [cited by examiner]
Peng, S., Xu, Z., Dong, J., Wang, Q., Zhang, S., Shuai, Q., . . . & Zhou, X. (2024). Animatable implicit neural representations for creating realistic avatars from videos. IEEE Transactions on Pattern Analysis and Mach… [cited by examiner]
Huang, Z., Xu, Y., Lassner, C., Li, H., & Tung, T. (2020). Arch: Animatable reconstruction of clothed humans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3093-3102). (Year… [cited by examiner]
Wu, S., Rupprecht, C., & Vedaldi, A. (2020). Unsupervised learning of probably symmetric deformable 3d objects from images in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogniti… [cited by examiner]
Habermann, M., Xu, W., Zollhoefer, M., Pons-Moll, G., & Theobalt, C. (2021). A deeper look into deepcap. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4), 4009-4022. (Year: 2021). [cited by examiner]
J. Y. Chang, G. Moon and K. M. Lee, “V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recogn… [cited by examiner]
Kaneko, Takuhiro. “Ar-nerf: Unsupervised learning of depth and defocus effects from natural images with aperture rendering neural radiance fields.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R… [cited by examiner]
Aliaksandr Siarohin et al: “Motion Representations for Articulated Animation”, Arxiv.Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 22, 2021 (Apr. 22, 2021). [cited by applicant]
Ben Mildenhall et al: “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 4, 2020 (Aug. 4, 2020). [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/085298, dated May 6, 2024 (Jun. 5, 2024)—10 pages. [cited by applicant]
Lewis J P et al: “Pose space deformation a unified approach to shape interpolation and skeleton-driven deformation”, Computer Graphics. Siggraph 2000 Conference Proceedings. New Orleans, LA, Jul. 23-28, 2000; [Computer … [cited by applicant]
Vincent Lepetit et al: “EPnP: An Accurate O(n) Solution to the PnP Problem”, Arxiv.org, vo1. 81, No. 2, Jul. 19, 2008 (Jul. 19, 2008), pp. 155-166, XP000055627426, 201 Olin Library Cornell University Ithaca, NY 14853. [cited by applicant]