IP Library Granted Patent US 12,190,428
Granted Patent B2
US 12,190,428 · App. 18/333,647 · Granted Jan 7, 2025

Deep relightable appearance models for animatable face avatars

Inventors: Jason Saragih (Pittsburgh, PA); Stephen Anthony Lombardi (Pittsburgh, PA); Shunsuke Saito (Pittsburgh, PA); Tomas Simon Kreuz (Pittsburgh, PA); Shih-En Wei (Pittsburgh, PA); Kevyn Alex Anthony McPhail (Pittsburgh, PA); Yaser Sheikh (Pittsburgh, PA); Sai Bi (La Jolla, CA)
Assignee: Meta Platforms Technologies, LLC
G06T13/40G06T11/001G06T15/04G06T15/506G06T15/60G06T2215/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,428
App. No.
18/333,647
Granted
Jan 7, 2025
Kind
B2
Abstract

A method for providing a relightable avatar of a subject to a virtual reality application is provided. The method includes retrieving multiple images including multiple views of a subject and generating an expression-dependent texture map and a view-dependent texture map for the subject, based on the images. The method also includes generating, based on the expression-dependent texture map and the view-dependent texture map, a view of the subject illuminated by a light source selected from an environment in an immersive reality application, and providing the view of the subject to an immersive reality application running in a client device. A non-transitory, computer-readable medium storing instructions and a system that executes the instructions to perform the above method are also provided.

Claims (36)

1. A computer-implemented method, comprising:

generating texture maps for a subject, based on multiple images including multiple views of the subject under a plurality of lighting conditions, the texture maps including spatially varied expressions and viewpoints across the plurality of lighting conditions;

generating a final texture map based on the texture maps and intensity-defined weights corresponding to the plurality of lighting conditions;

generating, based on the final texture map, a view of the subject illuminated by a light source selected from an environment of the subject; and

providing the view of the subject to an immersive reality application running in a client device.

2. The computer-implemented method of claim 1 , wherein generating the texture maps further comprises interpolating a lighting configuration based on a first lighting condition and a second lighting condition available in the multiple images, the texture maps including an expression-dependent texture map and a view-dependent texture map with spatially varied expressions and viewpoints across the plurality of lighting conditions.

3. The computer-implemented method of claim 1 , further comprising retrieving the multiple images including the multiple views of the subject from one or more frames from one or more headset mounted cameras facing a user of the client device, wherein the client device is a virtual reality headset.

4. The computer-implemented method of claim 1 , wherein generating the texture maps for the subject comprises selecting a lighting configuration for the immersive reality application.

5. The computer-implemented method of claim 1 , wherein generating the texture maps for the subject comprises determining a lighting configuration based on an environment map including multiple lighting configurations in the environment of the subject in the immersive reality application.

6. The computer-implemented method of claim 1 , wherein generating the texture maps for the subject comprises determining a location of the environment of the subject in the immersive reality application, a subject orientation in the environment, and a view direction.

7. The computer-implemented method of claim 2 , wherein generating the texture maps comprises retrieving a shadow map to encode a geometric association between the light source selected from the environment in the immersive reality application and the view-dependent texture map.

8. The computer-implemented method of claim 2 , wherein generating the expression-dependent texture map comprises linearly combining multiple expression dependent texture maps based on a lighting condition of the expression-dependent texture map.

9. The computer-implemented method of claim 1 , wherein generating the view of the subject comprises identifying a clear shadow boundary from a self-occlusion from a portion of a face of the subject.

10. The computer-implemented method of claim 1 , further comprising providing a video of the subject based on animated views of the subject in the immersive reality application.

11. A system, comprising:

a memory storing multiple instructions; and

one or more processors configured to execute the instructions to cause the system to:

generate texture maps for a subject, based on multiple images including multiple views of the subject under a plurality of lighting conditions, the texture maps including spatially varied expressions and viewpoints across the plurality of lighting conditions;

generate a final texture map based on the texture maps and intensity-defined weights corresponding to the plurality of lighting conditions;

generate, based on the final texture map, a view of the subject illuminated by a light source selected from an environment of the subject; and

provide the view of the subject to an immersive reality application running in a client device.

12. The system of claim 11 , wherein the one or more processors further execute instructions to generate the texture maps including interpolating a lighting configuration based on a first lighting condition and a second lighting condition available in the multiple images, the texture maps including an expression-dependent texture map and a view-dependent texture map.

13. The system of claim 11 , wherein the one or more processors further execute instructions to retrieve the multiple images including the multiple views of the subject from one or more frames from one or more headset mounted cameras facing a user of the client device, wherein the client device is a virtual reality headset.

14. The system of claim 11 , wherein to generate the texture maps for the subject the one or more processors execute instructions to select a lighting configuration for the immersive reality application.

15. The system of claim 11 , wherein to generate the texture maps for the subject the one or more processors execute instructions to determine a lighting configuration based on an environment map including multiple lighting configurations in the environment of the subject in the immersive reality application.

16. The system of claim 11 , wherein to generate the texture maps for the subject the one or more processors execute instructions to determine a location of the environment of the subject in the immersive reality application, a subject orientation in the environment, and a view direction.

17. The system of claim 12 , wherein to generate the expression-dependent texture map and the view-dependent texture map for the subject, the one or more processors execute instructions to retrieve a shadow map to encode a geometric association between the light source in the immersive reality application and the view-dependent texture map.

18. The system of claim 12 , wherein to generate the expression-dependent texture map the one or more processors execute instructions to linearly combine multiple expression dependent texture maps based on a lighting condition of the expression-dependent texture map.

19. The system of claim 11 , wherein to generate the view of the subject the one or more processors execute instructions to identify a clear shadow boundary from a self-occlusion from a portion of a face of the subject.

20. A computer-implemented method for training a model to generate a relightable, three-dimensional representation of a subject, comprising:

generating, with a relightable appearance model, an expression-dependent texture map and a view-dependent texture map for a subject, based on images including multiple views of the subject under multiple space-multiplexed and time-multiplexed illumination patterns;

generating a final texture based on the expression-dependent texture map, the view-dependent texture map, and intensity-defined weights corresponding to the multiple space-multiplexed and the time-multiplexed illumination patterns;

generating, based on the final texture, a synthetic view of the subject illuminated by each of the space-multiplexed and the time-multiplexed illumination patterns;

determining a loss value indicative of a difference between the synthetic view of the subject and at least one of the images;

updating the relightable appearance model based on the loss value; and

generating a three-dimensional relightable representation of the subject based on an updated relightable appearance model.

Assignments (2)
CHANGE OF NAME Recorded Aug 15, 2023
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 064599/0030 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: SARAGIH, JASON; LOMBARDI, STEPHEN ANTHONY; SAITO, SHUNSUKE; KREUZ, TOMAS SIMON; WEI, SHIH-EN; MCPHAIL, KEVYN ALEX ANTHONY; SHEIKH, YASER; BI, SAI
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 064216/0283 →
Continuity (3)
Continuation 17580486 · Jan 20, 2022
Provisional Application 63141871 · Jan 26, 2021
Related Publication 20230326112A1 · Oct 12, 2023
References Cited (38)
US 20130222408A1 · Lee · 2013 [cited by examiner]
US 20190213772A1 · Lombardi et al. · 2019 [cited by applicant]
US 20210067840A1 · Mate · 2021 [cited by examiner]
US 20210287416A1 · O'Hagan et al. · 2021 [cited by applicant]
US 20210366184A1 · Leroux et al. · 2021 [cited by applicant]
WO 2017029488A2 · 2017 [cited by applicant]
Huang, Xiang, et al. “Near light correction for image relighting and 3D shape recovery.” 2015 Digital Heritage. vol. 1. IEEE, 2015. (Year: 2015). [cited by examiner]
Loscos, Céline, et al. “Interactive virtual relighting and remodeling of real scenes.” Rendering Techniques' 99: Proceedings of the Eurographics Workshop in Granada, Spain, Jun. 21-23, 1999 10. Springer Vienna, 1999. (Y… [cited by examiner]
Havran, Vlastimil, et al. “Interactive System for Dynamic Scene Lighting using Captured Video Environment Maps.” Rendering Techniques. 2005. (Year: 2005). [cited by examiner]
Busbridge I.W., “The Mathematics of Radiative Transfer,” Cambridge University Press, 1960, No. 50, 81 pages. [cited by applicant]
Cao C., et al., “Real-Time High-Fidelity Facial Performance Capture,” ACM Transactions on Graphics (TOG), 2015, vol. 34, No. 4, pp. 1-9. [cited by applicant]
Debevec P., et al., “Acquiring the Reflectance Field of a Human Face,” Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, 2000, pp. 145-156. [cited by applicant]
Garrido P., et al., “Reconstructing Detailed Dynamic Face Geometry from Monocular Video,” ACM Transactions on Graphics, 2013, vol. 32, pp. 1-10. [cited by applicant]
Ghosh A., et al., “Practical Modeling and Acquisition of Layered Facial Reflectance,” In ACM SIGGRAPH Asia 2008 papers, 2008, pp. 1-10. [cited by applicant]
Gotardo P., et al., “Practical Dynamic Facial Appearance Modeling and Acquisition,” ACM Transactions on Graphics (ToG), Dec. 2018, vol. 37, No. 6, Article 232, pp. 1-13, Retrieved from the Internet: URL: https://doi.org… [cited by applicant]
Guo K., et al., “The Relightables: Volumetric Performance Capture of Humans with Realistic Relighting,” ACM Transactions on Graphics, Article 217, vol. 38(6), Nov. 2019, pp. 1-19. [cited by applicant]
Ha D., et al., “HyperNetworks,” ArXiv Preprint Arxiv: 1609.09106V4, Dec. 1, 2016, 29 pages. [cited by applicant]
EPO—International Search Report and Written Opinion for International Application No. PCT/US2022/013820, mailed Jun. 7, 2022, 11 pages. [cited by applicant]
Jensen H.W., “A Practical Model for Subsurface Light Transport,” In Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques, 2001, pp. 511-518. [cited by applicant]
Kingma; et al., “ADAM: A Method for Stochastic Optimization,” ArXir:1412.6980v1, Dec. 22, 2014, 9 pages. [cited by applicant]
Lombardi S., et al., “Deep Appearance Models for Face Rendering,” ACM Transactions on Graphics, Aug. 2018, vol. 37 (4), Article 68, pp. 1-13. [cited by applicant]
Ma W-C., et al., “Rapid Acquisition of Specular and Diffuse Normal Maps from Polarized Spherical Gradient Illumination,” Rendering Techniques, 2007, vol. 9, 12 pages. [cited by applicant]
Meka A., et al., “Deep Reflectance Fields: High-Quality Facial Reflectance Field Inference from Color Gradient Illumination,” ACM Transactions on Graphics (TOG), 2019, vol. 38, No. 4, pp. 1-12. [cited by applicant]
Meka A., et al., “Deep Relightable Textures—Volumetric Performance Capture with Neural Rendering,” ACM Transactions on Graphics Proceedings SIGGRAPH Asia, 2020, vol. 39, No. 6, Article 259, pp. 1-21, Retrieved from the … [cited by applicant]
Nagano K., et al., “paGAN: Real-time Avatars Using Dynamic Textures,” ACM Transactions on Graphics (TOG), vol. 37, No. 6, Nov. 2018, 12 pages. [cited by applicant]
Pighin F., et al., “Synthesizing Realistic Facial Expressions from Photographs,” International Conference on Computer Graphics and Interactive Techniques, ACM SIGGRAPH, Jul. 30, 2006, 10 pages. [cited by applicant]
Schwartz G., et al., “The Eyes Have It: An Integrated Eye and Face Model for Photorealistic Facial Animation,” ACM Transactions on Graphics (TOG), Jul. 2020, vol. 39, No. 4, 15 Pages. [cited by applicant]
Sevastopolsky A., et al., “Relightable 3D Head Portraits from a Smartphone Video,” Arxiv, Dec. 17, 2020, 15 pages. [cited by applicant]
Seymour M., “Meet Mike: Epic Avatars,” In ACM SIGGRAPH VR Village, 2017, 2 pages. [cited by applicant]
Shu Z., et al., “Neural Face Editing with Intrinsic Image Disentangling,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5541-5550. [cited by applicant]
Sun T., et al., “Single Image Portrait Relighting,” ACM Transactions on Graphics (TOG), Jul. 2019, vol. 38, No. 4, Article 79, pp. 1-12, Retrieved from the Internet: URL: https://doi.org/10.1145/3306346.3323008. [cited by applicant]
Tewari A., et al., State of the Art on Neural Rendering, State of The Art Report (STAR), May 2020, vol. 39, No. 2, 27 pages. [cited by applicant]
Wenger A., et al., “Performance Relighting and Reflectance Transformation with Time-Multiplexed Illumination,” ACM Transactions on Graphics (TOG), 2005, vol. 24, No. 3, pp. 756-764. [cited by applicant]
Weyrich T., et al., “Analysis of Human Faces using a Measurement-Based Skin Reflectance Model,” ACM Transactions on Graphics (TOG), 2006, vol. 25, No. 3, pp. 1013-1024. [cited by applicant]
Williams L., “Casting Curved Shadows on Curved Surfaces,” In Proceedings of the 5th Annual Conference on Computer Graphics and Interactive Techniques, 1978, pp. 270-274. [cited by applicant]
Xu Z., et al., “Deep Image-Based Relighting from Optimal Sparse Samples,” ACM Transactions on Graphics (TOG), 2018, vol. 37, No. 4, Article 126, pp. 1-13. [cited by applicant]
Yamaguchi S., et al., “High-Fidelity Facial Reflectance and Geometry Inference from an Unconstrained Image,” ACM Transactions on Graphics (TOG), 2018, vol. 37, No. 4, Article 162, pp. 1-14. [cited by applicant]
Zhang X., et al., “Neural Light Transport for Relighting and View Synthesis,” ACM Transactions on Graphics (TOG), 2020, vol. 40, No. 1, Article 9, pp. 1-17. [cited by applicant]