IP Library › Granted Patent US 12,430,863
Granted Patent B2
US 12,430,863 · App. 17/892,097 · Granted Sep 30, 2025

Deformable neural radiance field for editing facial pose and facial expression in neural 3D scenes

Inventors: Zhixin Shu (San Jose, CA); Zexiang Xu (San Jose, CA); Shahrukh Athar (Stony Brook, NY); Kalyan Sunkavalli (San Jose, CA); Elya Shechtman (Seattle, WA)
Assignee: Adobe Inc.
G06T19/20G06T17/00G06T2200/08G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,863
App. No.
17/892,097
Granted
Sep 30, 2025
Kind
B2
Abstract

A scene modeling system receives a video including a plurality of frames corresponding to views of an object and a request to display an editable three-dimensional (3D) scene that corresponds to a particular frame of the plurality of frames. The scene modeling system applies a scene representation model to the particular frame, and includes a deformation model configured to generate, for each pixel of the particular frame based on a pose and an expression of the object, a deformation point using a 3D morphable model (3DMM) guided deformation field. The scene representation model includes a color model configured to determine, for the deformation point, color and volume density values. The scene modeling system receives a modification to one or more of the pose or the expression of the object including a modification to a location of the deformation point and renders an updated video based on the received modification.

Claims (43)

1. A method, comprising:

receiving, at a scene modeling system, a video including a plurality of frames corresponding to a plurality of views of an object and a request to display an editable three-dimensional (3D) scene that includes the object and that corresponds to a particular frame of the plurality of frames;

generating, by the scene modeling system, the editable 3D scene by applying a scene representation model to the particular frame, wherein the scene representation model comprises:

a deformation model configured to generate a 3D morphable model (“3DMM”) guided deformation field based on a 3DMM deformation field and a residual predicted by the deformation model, and further generate, for each pixel of the particular frame and based on a pose and an expression of the object, a deformation point using the 3DMM guided deformation field, and

a color model configured to determine, for the deformation point and using a volume rendering process, a color value and a volume density value; and

providing, by the scene modeling system, the editable 3D scene to a computing device executing a scene modeling application configured to generate a modified video using the editable 3D scene, wherein the scene modeling application generates the modified video by:

receiving a modification to one or more of the pose or the expression of the object including at least a modification to a location of the deformation point,

receiving a modification to a view of the editable 3D scene, wherein the modification to the view comprises one or more of a change to a camera position or to a camera orientation within the editable 3D scene,

rendering an updated editable 3D scene based on the received modification to the one or more of the pose or the expression of the object and by changing the view of the editable 3D scene to the modified view, and

generating the modified video including an updated frame to replace the particular frame, the updated frame generated based on the updated editable 3D scene.

2. The method of claim 1 , wherein the edit to the object comprises a change to one or more of the pose or the expression.

3. The method of claim 1 , wherein the pose of the object represents an orientation of the object with respect to a default pose and wherein the edit to the pose of the object corresponds a change in the orientation of the object.

4. The method of claim 1 , wherein the expression of the object corresponds to a first semantic category and wherein the edit to the expression of the object comprises a selection of a second semantic category.

5. The method of claim 1 , wherein the pose and the expression are extracted from the particular frame using a detailed expression capture and animation (“DECA”) method.

6. The method of claim 1 , wherein deforming the point using the 3DMM guided deformation field comprises transforming the point to a canonical space, wherein the color model is applied to the transformed point.

7. The method of claim 1 , wherein each of the deformation model and the color model comprise a multilayer perceptron model, wherein the deformation model is configured to generate the 3DMM guided deformation field, residual, and deformation points in a first stage, and the color model is configured to determine, in a second stage, a color value and volume density value based on the deformation point generated in the first stage, and further based on the pose and expression of the object.

8. A system, comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

generating, for a particular frame of a video, an editable 3D scene that includes an object that is depicted in the particular frame of the video by applying a scene representation model to the particular frame, wherein the video includes a plurality of frames corresponding to a plurality of views of an object including the particular frame, wherein the scene representation model comprises:

a deformation model configured to generate a 3D morphable model (“3DMM”) guided deformation field based on a 3DMM deformation field, and a residual predicted by the deformation model, and further generate, for each pixel of the particular frame and based on a pose and an expression of the object, a deformation point using the 3DMM guided deformation field, and

a color model configured to determine, for the deformation point and using a volume rendering process, a color value and a volume density value; and

providing the editable 3D scene to a computing device executing a scene modeling application configured to generate a modified video using the editable 3D scene, wherein the scene modeling application generates the modified video by:

receiving a modification to one or more of the pose or the expression of the object including at least a modification to a location of the deformation point,

receiving a modification to a view of the editable 3D scene, wherein the modification to the view comprises one or more of a change to a camera position or to a camera orientation within the editable 3D scene,

rendering an updated editable 3D scene based on the received modification to the one or more of the pose or the expression of the object and by changing the view of the editable 3D scene to the modified view, and

generating a modified video including an updated frame to replace the particular frame, the updated frame generated based on the updated editable 3D scene.

9. The system of claim 8 , wherein the edit to the object comprises a change to one or more of the pose or the expression.

10. The system of claim 8 , wherein the pose of the object represents an orientation of the object with respect to a default pose and wherein the edit to the pose of the object corresponds a change in the orientation of the object.

11. The system of claim 8 , wherein the expression of the object corresponds to a first semantic category and wherein the edit to the expression of the object comprises a selection of a second semantic category.

12. The system of claim 8 , wherein deforming the point using the 3DMM guided deformation field comprises transforming the point to a canonical space, wherein the color model is applied to the transformed point.

13. The system of claim 8 , wherein each of the deformation model and the color model comprise a multilayer perceptron model, wherein the deformation model is configured to generate the 3DMM guided deformation field, residual, and deformation points in a first stage, and the color model is configured to determine, in a second stage, a color value and volume density value based on the deformation point generated in the first stage, and further based on the pose and expression of the object.

14. A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

generating, for a particular frame of a video, an editable 3D scene that includes an object that is depicted in the particular frame of the video by applying a scene representation model to the particular frame, wherein the video includes a plurality of frames corresponding to a plurality of views of an object including the particular frame, wherein the scene representation model comprises:

a deformation model configured to generate a 3D morphable model (“3DMM”) guided deformation field based on a 3DMM deformation field, and a residual predicted by the deformation model, and further generate, for each pixel of the particular frame and based on a pose and an expression of the object, a deformation point using the 3DMM guided deformation field, and

a color model configured to determine, for the deformation point and using a volume rendering process, a color value and a volume density value;

receiving a modification to one or more of the pose or the expression of the object including at least a change in location of the deformation point;

receiving a modification to a view of the editable 3D scene, wherein the modification to the view comprises one or more of a change to a camera position or to a camera orientation within the editable 3D scene;

rendering an updated editable 3D scene based on the received modification to the one or more of the pose or the expression of the object and by changing the view of the editable 3D scene to the modified view; and

generating a modified video including an updated frame to replace the particular frame, the updated frame generated based on the updated editable 3D scene.

15. The non-transitory computer-readable medium of claim 14 , wherein the edit to the object comprises a change to one or more of the pose or the expression.

16. The non-transitory computer-readable medium of claim 14 , wherein deforming the point using the 3DMM guided deformation field comprises transforming the point to a canonical space, wherein the color model is applied to the transformed point.

17. The non-transitory computer-readable medium of claim 14 , wherein each of the deformation model and the color model comprise a multilayer perceptron model, wherein the deformation model is configured to generate the 3DMM guided deformation field, residual, and deformation points in a first stage, and the color model is configured to determine, in a second stage, a color value and volume density value based on the deformation point generated in the first stage, and further based on the pose and expression of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2022
From: SHU, ZHIXIN; XU, ZEXIANG; ATHAR, SHAHRUKH; SUNKAVALLI, KALYAN; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 060853/0217 →
Continuity (1)
Related Publication 20240062495A1 · Feb 22, 2024
References Cited (23)
US 20080025569A1 · Gordon · 2008 [cited by examiner]
US 20140043329A1 · Wang · 2014 [cited by examiner]
US 20140156398A1 · Li · 2014 [cited by examiner]
US 20210279956A1 · Chandran · 2021 [cited by examiner]
US 20220028150A1 · Soulvie · 2022 [cited by examiner]
US 20220245911A1 · Hu · 2022 [cited by examiner]
US 20230115531A1 · Zohar · 2023 [cited by examiner]
US 20230130281A1 · Brown · 2023 [cited by examiner]
US 20230196593A1 · Raghoebardajal · 2023 [cited by examiner]
US 20240005590A1 · Martin Brualla · 2024 [cited by examiner]
US 20240078773A1 · Kim · 2024 [cited by examiner]
US 20240095999A1 · Szabo · 2024 [cited by examiner]
US 20250037366A1 · Zoss · 2025 [cited by examiner]
Yao et al. (“Learning an Animatable Detailed 3D Face Model from In-The-Wild Images” Aug. 2021, ACM Trans. Graph., vol. 40, No. 4, Article 88). (Year: 2021). [cited by examiner]
Blanz et al., A Morphable Model for the Synthesis of 3D Faces, SIGGRAPH '99, Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, Jul. 1999, 8 pages. [cited by applicant]
Feng et al., Learning an Animatable Detailed 3D Face Model from In-the-wild Images, ACM Transactions on Graphics, vol. 40, Issue 4, Aug. 2021, pp. 1-13. [cited by applicant]
Gafni et al., Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction, Available online at https://openaccess.thecvf.com/content/CVPR2021/papers/Gafni_Dynamic_Neural_Radiance_Fields_for_Monocular_4D… [cited by applicant]
Guo et al., Towards Fast, Accurate and Stable 3D Dense Face Alignment, In Proceedings of the European Conference on Computer Vision (ECCV), Aug. 2020, 21 pages. [cited by applicant]
Li et al., Learning a Model of Facial Shape and Expression from 4D Scans, ACM Transactions on Graphics, vol. 36, Issue 6, Nov. 2017, pp. 1-17. [cited by applicant]
Mildenhall et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, European Conference on Computer Vision, Aug. 3, 2020, pp. 1-25. [cited by applicant]
Park et al., HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields, ACM Transactions on Graphics, vol. 40, Issue 6, Dec. 2021, pp. 1-12. [cited by applicant]
Park et al., Nerfies: Deformable Neural Radiance Fields, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 5865-5874. [cited by applicant]
Pumarola et al., D-NeRF: Neural Radiance Fields for Dynamic Scenes, Available online at https://openaccess.thecvf.com/content/CVPR2021/papers/Pumarola_D-NeRF_Neural_Radiance_Fields_for_Dynamic_Scenes_CVPR_2021_paper.pdf… [cited by applicant]