IP Library Granted Patent US 12,205,213
Granted Patent B2
US 12,205,213 · App. 17/526,647 · Granted Jan 21, 2025

Synthesizing sequences of images for movement-based performance

Inventors: Derek Edward Bradley (Zurich, CH); Prashanth Chandran (Zurich, CH); Paulo Fabiano Urnau Gotardo (Zurich, CH); Gaspard Zoss (Zurich, CH)
Assignees: Disney Enterprises, INC.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
G06T13/40G06N3/08G06T7/215G06T13/80G06T15/04G06T17/20G06T19/20G06T2207/20081G06T2207/20084G06T2207/30201G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,213
App. No.
17/526,647
Granted
Jan 21, 2025
Kind
B2
Abstract

A technique for rendering an input geometry includes generating a first segmentation mask for a first input geometry and a first set of texture maps associated with one or more portions of the first input geometry. The technique also includes generating, via one or more neural networks, a first set of neural textures for the one or more portions of the first input geometry. The technique further includes rendering a first image corresponding to the first input geometry based on the first segmentation mask, the first set of texture maps, and the first set of neural textures.

Claims (47)

1. A computer-implemented method for rendering an input geometry, the computer-implemented method comprising:

generating a first segmentation mask for a first input three-dimensional (3D) geometry and a first plurality of texture maps associated with a plurality of portions of the first input 3D geometry;

generating, via execution of a plurality of generator blocks included in one or more neural networks, a first plurality of neural textures for the plurality of portions of the first input 3D geometry, wherein each generator block included in the plurality of generator blocks generates a neural texture for a different portion included in the plurality of portions of the first input 3D geometry, wherein portions of the first input 3D geometry corresponding to different generator blocks included in the plurality of generator blocks are non-overlapping; and

rendering a first image corresponding to the first input 3D geometry based on the first segmentation mask, the first plurality of texture maps, and the first plurality of neural textures.

2. The computer-implemented method of claim 1 , further comprising training the one or more neural networks based on a training dataset that includes a second plurality of texture maps and a plurality of segmentation masks for a plurality of synthetic geometries.

3. The computer-implemented method of claim 1 , further comprising training the one or more neural networks based on one or more predictions generated by a discriminator neural network from one or more images produced by the one or more neural networks.

4. The computer-implemented method of claim 1 , wherein generating the first segmentation mask and the first plurality of texture maps comprises:

deforming a template mesh to match the first input 3D geometry; and

generating the first segmentation mask and the first plurality of texture maps based on a pose associated with the first input 3D geometry.

5. The computer-implemented method of claim 1 , further comprising rendering a second image corresponding to a second input geometry based on a second segmentation mask for the second input geometry, a second plurality of texture maps associated with a plurality of portions of the second input geometry, and the first plurality of neural textures.

6. The computer-implemented method of claim 1 , wherein generating the first plurality of neural textures comprises, for each portion included in the plurality of portions of the first input 3D geometry:

generating an input vector based on sampling a distribution of latent variables associated with the portion; and

inputting the input vector into the generator block corresponding to the portion.

7. The computer-implemented method of claim 1 , wherein rendering the first image comprises:

sampling the first plurality of neural textures based on the first plurality of texture maps to generate a plurality of screen-space neural features;

generating a composited set of screen-space neural features based on the first segmentation mask and the plurality of screen-space neural features; and

applying one or more convolutional layers to the composited set of screen-space neural features to render the first image.

8. The computer-implemented method of claim 1 , wherein the one or more neural networks comprise a generative neural network.

9. The computer-implemented method of claim 1 , wherein the first input 3D geometry comprises a face.

10. The computer-implemented method of claim 1 , wherein the plurality of portions of the first input 3D geometry comprises at least one of a skin, a hair, one or more eyes, a mouth, or a background.

11. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

generating one or more maps associated with a plurality of portions of a first input three-dimensional (3D) geometry;

generating, via execution of a plurality of generator blocks included in one or more neural networks, a first plurality of neural textures for the plurality of portions of the first input 3D geometry, wherein each generator block included in the plurality of generator blocks generates a neural texture for a different portion included in the plurality of portions of the first input 3D geometry, wherein portions of the first input 3D geometry corresponding to different generator blocks included in the plurality of generator blocks are non-overlapping; and

rendering a first image corresponding to the first input 3D geometry based on the one or more maps and the first plurality of neural textures.

12. The one or more non-transitory computer readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of training the one or more neural networks based on a training dataset that includes a plurality of maps associated with a plurality of synthetic geometries.

13. The one or more non-transitory computer readable media of claim 12 , wherein training the one or more neural networks comprises updating parameters of the one or more neural networks based on one or more predictions generated by a discriminator neural network from one or more images produced by the one or more neural networks.

14. The one or more non-transitory computer readable media of claim 11 , wherein generating the one or more maps comprises:

deforming a template mesh to match the first input 3D geometry; and

generating a segmentation mask and a set of texture maps based on a pose associated with the first input 3D geometry.

15. The one or more non-transitory computer readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of rendering a second image corresponding to the first input 3D geometry based on the one or more maps and a second set of neural textures for the plurality of portions of the first input 3D geometry.

16. The one or more non-transitory computer readable media of claim 11 , wherein generating the first plurality of neural textures comprises:

generating one or more input vector based on sampling one or more distributions of latent variables associated with the plurality of portions;

executing a mapping network included in the one or more neural networks to convert one or more sampled vectors into one or more input vectors; and

executing the plurality of generator blocks to convert the one or more input vectors into the first plurality of neural textures.

17. The one or more non-transitory computer readable media of claim 11 , wherein rendering the first image comprises:

sampling the first plurality of neural textures based on a first plurality of texture maps included in the one or more maps to generate a plurality of screen-space neural features;

generating a composited set of screen-space neural features based on a first segmentation mask included in the one or more maps and the plurality of screen-space neural features; and

applying one or more convolutional layers to the composited set of screen-space neural features to render the first image.

18. The one or more non-transitory computer readable media of claim 11 , wherein the first input 3D geometry comprises a face.

19. The one or more non-transitory computer readable media of claim 11 , wherein the one or more maps comprises at least one of a skin texture map, a hair texture map, an eye texture map, a mouth texture map, or a background texture map.

20. A system, comprising:

one or more memories that store instructions, and

one or more processors that are coupled to the one or more memories and,

when executing the instructions, are configured to:

generate a first segmentation mask for a first input three-dimensional (3D) geometry and a first plurality of texture maps associated with a plurality of portions of the first input 3D geometry;

generate, via execution of a plurality of generator blocks included in one or more neural networks, a first plurality of neural textures for the plurality of portions of the first input 3D geometry, wherein each generator block included in the plurality of generator blocks generates a neural texture for a different portion included in the plurality of portions of the first input 3D geometry, wherein portions of the first input 3D geometry corresponding to different generator blocks included in the plurality of generator blocks are non-overlapping; and

render a first image corresponding to the first input 3D geometry based on the first segmentation mask, the first plurality of texture maps, and the first plurality of neural textures.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 058306/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2021
From: BRADLEY, DEREK EDWARD; CHANDRAN, PRASHANTH; URNAU GOTARDO, PAULO FABIANO; ZOSS, GASPARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 058263/0359 →
Continuity (1)
Related Publication 20230154090A1 · May 18, 2023
References Cited (70)
US 10719742B2 · Shechtman et al. · 2020 [cited by applicant]
US 10762337B2 · Sharma et al. · 2020 [cited by applicant]
US 11373352B1 · Gafni · 2022 [cited by examiner]
US 20100321386A1 · Lin et al. · 2010 [cited by applicant]
US 20120229463A1 · Yeh et al. · 2012 [cited by applicant]
US 20120231886A1 · Gomez et al. · 2012 [cited by applicant]
US 20160065497A1 · Cole et al. · 2016 [cited by applicant]
US 20170213112A1 · Sachs · 2017 [cited by examiner]
US 20180030084A1 · Feng et al. · 2018 [cited by applicant]
US 20180300842A1 · Feng et al. · 2018 [cited by applicant]
US 20180374242A1 · Li · 2018 [cited by examiner]
US 20190035149A1 · Chen · 2019 [cited by examiner]
US 20190017190A1 · Salavon · 2019 [cited by applicant]
US 20190171908A1 · Salavon · 2019 [cited by examiner]
US 20200143171A1 · Lee · 2020 [cited by examiner]
US 20200265567A1 · Hu et al. · 2020 [cited by applicant]
US 20200302029A1 · Holm et al. · 2020 [cited by applicant]
US 20210049468A1 · Karras et al. · 2021 [cited by applicant]
US 20210232803A1 · Fu et al. · 2021 [cited by applicant]
US 20220014723A1 · Pandey · 2022 [cited by examiner]
US 20220036626A1 · Vo · 2022 [cited by examiner]
US 20220051485A1 · Martin Brualla · 2022 [cited by examiner]
US 20220121876A1 · Kalarot · 2022 [cited by examiner]
US 20220130111A1 · Martin Brualla · 2022 [cited by examiner]
US 20220180602A1 · Hao · 2022 [cited by examiner]
US 20220188696A1 · Yang et al. · 2022 [cited by applicant]
US 20230062924A1 · Henley · 2023 [cited by examiner]
US 20230077187A1 · Zafeiriou et al. · 2023 [cited by applicant]
WO 2020117657A1 · 2020 [cited by applicant]
WO 2020256471A1 · 2020 [cited by applicant]
WO 2021096192A1 · 2021 [cited by applicant]
WO 2021145862A1 · 2021 [cited by applicant]
Karras et al., “Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion”, ACM Transactions on Graphics, DOI: http://dx.doi.org/10.1145/3072959.3073658, vol. 36, No. 4, Article 94, Jul. 2017, pp. 9… [cited by applicant]
Taylor et al., “A Deep Learning Approach for Generalized Speech Animation”, ACM Transactions on Graphics, DOI: 10.1145/3072959.3073699, vol. 36, No. 4, Article 93, Jul. 2017, pp. 93:1-93:11. [cited by applicant]
Holden et al., “Learned Motion Matching”, ACM Transactions on Graphics, https://doi.org/10.1145/3386569.3392440, vol. 39, No. 4, Article 53, Jul. 2020, pp. 53:1-53:13. [cited by applicant]
Harvey et al., “Robust Motion In-betweening”, ACM Transactions on Graphics, https://doi.org/10.1145/3386569.3392480, vol. 39, No. 4, Article 60, Jul. 2020, pp. 60:1-60:12. [cited by applicant]
Zoss et al., “Data-driven Extraction and Composition of Secondary Dynamics in Facial Performance Capture”, ACM Transactions on Graphics, https://doi.org/10.1145/3386569.3392463, vol. 39, No. 4, Article 1, Jul. 2020, pp.… [cited by applicant]
Pons-Moll et al., “Dyna: A Model of Dynamic Human Shape in Motion”, ACM Transactions on Graphics, DOI: http://dx.doi.org/10.1145/2766993, vol. 34, No. 4, Article 120, Aug. 2015, pp. 120:1-120:14. [cited by applicant]
Santesteban et al., “SoftSMPL: Data-driven Modeling of Nonlinear Soft-tissue Dynamics for Parametric Humans”, EUROGRAPHICS 2020, vol. 39, No. 2, Apr. 1, 2020, 11 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need”, 31st Conference on Neural Information Processing Systems, arXiv:1706.03762, Dec. 6, 2017, p. 1-15. [cited by applicant]
Shaw et al., “Self-Attention with Relative Position Representations”, arXiv:1803.02155, Apr. 12, 2018, 5 pages. [cited by applicant]
Jiang et al., “TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up”, arXiv:2102.07074, Jun. 14, 2021, p. 1-19. [cited by applicant]
Esser et al., “Taming Transformers for High-Resolution Image Synthesis”, arXiv:2012.09841, Jun. 23, 2021, pp. 1-52. [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of StyleGAN”, arXiv:1912.04958, Mar. 23, 2020, pp. 1-21. [cited by applicant]
Gecer et al., “Synthesizing Coupled 3D Face Modalities by Trunk-Branch Generative Adversarial Networks”, arXiv:1909.02215, Dec. 2, 2020, pp. 1-19. [cited by applicant]
Richardson et al., “Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation”, arXiv:2008.00951, Apr. 21, 2021, 21 pages. [cited by applicant]
Park et al., “Semantic Image Synthesis with Spatially-Adaptive Normalization”, arXiv:1903.07291, Nov. 5, 2019, pp. 1-19. [cited by applicant]
Zhu et al., “SEAN: Image Synthesis with Semantic Region-Adaptive Normalization”, arXiv:1911.12861, May 24, 2020, 19 pages. [cited by applicant]
Petrov et al., “DeepFaceLab: Integrated, flexible and extensible face-swapping framework”, arxiv:2005.05535, Jun. 29, 2021, pp. 1-10. [cited by applicant]
Naruniec et al., “High-Resolution Neural Face Swapping for Visual Effects”, Eurographics Symposium on Rendering 2020, vol. 39, No. 4, 2020, 16 pages. [cited by applicant]
Gafni et al., “Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction”, arXiv:2012.03065, Dec. 5, 2020, pp. 1-11. [cited by applicant]
Huang et al., “Music Transformer: Generating Music with Long-Term Structure”, arXiv:1809.04281, Dec. 12, 2018, pp. 1-14. [cited by applicant]
Wang et al., “Video-to-Video Synthesis”, arXiv:1808.06601, Dec. 3, 2018, pp. 1-14. [cited by applicant]
Karras et al., “Alias-Free Generative Adversarial Networks”, 35th Conference on Neural Information Processing Systems, arXiv:2106.12423, Oct. 18, 2021, pp. 1-31. [cited by applicant]
Turkoglu et al., “A Layer-Based Sequential Framework for Scene Generation with GANs”, AAAI 2019, arXiv:1902.00671, 2019, 9 pages. [cited by applicant]
Wang et al., “Generative Image Modeling Using Style and Structure Adversarial Networks”, European Conference on Computer Vision, https://link.springer.com/chapter/10.1007 /978-3-319-46493-0 20, 2016, pp. 318-335. [cited by applicant]
GB Combined Search and Examination Report for Application No. 2217056.7 dated May 16, 2023. [cited by applicant]
Regateiro et al., “Dynamic Surface Animation using Generative Networks”, International Conference on 3D Vision, DOI 10.1109/3DV.2019.00049, 2019, pp. 376-385. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 17/526,608 dated Jul. 28, 2023, 31 pages. [cited by applicant]
Phi, Michael, “Illustrated Guide to Transformers—Step by Step Explanation”, Towards Data Science, Apr. 30, 2020, 29 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/526,608, dated Nov. 16, 2023, 19 pages. [cited by applicant]
Non-final Office Action received for U.S. Appl. No. 17/526,608, dated Feb. 29, 2024, 24 pages. [cited by applicant]
Canadian Office Action, Application Serial No. 3,180,427, dated Apr. 26, 2024, 9 pages. [cited by applicant]
Phi, Michael, “Illustrated Guide to Transformers—Step by Step Explanation”, Towards Data Science, Apr. 30, 2020, 31 pages. https://towardsdatascience.com/illustrated-guide-to-transformers-step-by-step-explanation-f74876… [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/669,053 dated Apr. 25, 2024, 26 pages. [cited by applicant]
New Zealand Examination Report for NZ Application Serial No. 794049, dated May 8, 2024, 3 pages. [cited by applicant]
Great Britain Examination Report, Application No. GB2217056.7, dated Apr. 25, 2024, 4 pages. [cited by applicant]
Regateiro, Joao et al., “Dynamic Surface Animation using Generative Networks”, 2019 International Conference on 3D Vision (3DV), 2019, p. 376-385. [cited by applicant]
Australian Examination Report No. 2, Application No. 2022263508, dated Mar. 26, 2024, 3 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/526,608, dated Sep. 20, 2024, 25 pages. [cited by applicant]