IP Library Granted Patent US 12,620,189
Granted Patent B2
US 12,620,189 · App. 18/521,009 · Granted May 5, 2026

Method, electronic device and computer program

Inventors: Lev Markhasin (Stuttgart, DE); Iheb Belgacem (Stuttgart, DE); Shivangi Aneja (Munich, DE); Matthias Nießner (Munich, DE); Angela Dai (Stuttgart, DE)
Assignee: Sony Semiconductor Solutions Corporation
G06T19/20G06T15/04G06T17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,189
App. No.
18/521,009
Granted
May 5, 2026
Kind
B2
Abstract

A method for user command-guided editing of an initial textured 3D morphable model of an object comprising: obtaining the initial textured 3D morphable model of the object comprising an initial texture map and an initial 3D mesh model of the object; and determining an edited texture map of the object corresponding to the user command by editing the initial texture map of the object based on a first artificial neural network; and/or determining an edited 3D mesh model of the object corresponding to the user command by editing the initial 3D mesh model of the object based on a second artificial neural network; and generating an edited textured 3D morphable model of the object corresponding to the user command based on the edited texture map of the object and/or the edited 3D mesh model of the object.

Claims (46)

1 . A method for user command-guided editing of an initial textured 3D morphable model of an object comprising:

obtaining the initial textured 3D morphable model of the object comprising an initial texture map of the object and an initial 3D mesh model of the object; and

determining an edited texture map of the object corresponding to the user command by editing the initial texture map of the object based on a first artificial neural network; and/or

determining an edited 3D mesh model of the object corresponding to the user command by editing the initial 3D mesh model of the object based on a second artificial neural network; and

generating an edited textured 3D morphable model of the object corresponding to the user command based on the edited texture map of the object and/or the edited 3D mesh model of the object,

wherein the edited texture map is determined by a third artificial neural network based on a sum of the initial texture latent code and an offset texture latent code,

wherein the third artificial neural network is trained by adversarial self-supervised training on a plurality of two-dimensional RGB images using differentiable rendering.

2 . The method of claim 1 , further comprising:

generating the initial texture map of the object and a corresponding initial texture latent code based on a third artificial neural network; and/or

generating the initial 3D mesh model of the object based on an initial general appearance parameter of the 3D mesh model.

3 . The method of claim 2 , further comprising:

determining the offset texture latent code based on the initial texture latent code by the first artificial neural network corresponding to the user command; and/or

determining an offset general appearance parameter of the 3D mesh model based on the initial texture latent code by the second artificial neural network corresponding to the user command, and determining the edited 3D mesh model of the object based on the offset general appearance parameter.

4 . The method of claim 2 ,

wherein the object is a human face, and the initial 3D mesh model is a FLAME model and the initial general appearance parameter of the 3D mesh model is linear expression coefficients; and/or

wherein the object is a human person or parts thereof, and the initial 3D mesh model is a SMPL-X model, and the initial general appearance parameter of the 3D mesh model is a jaw joint parameter, finger joints parameter, remaining body joints parameter, combined body, face, hands shape parameters and/or facial expression parameters.

5 . The method of claim 1 , wherein the first artificial neural network and/or the second artificial neural network are trained based on one or more texture maps and corresponding texture latent codes, wherein the texture maps and corresponding texture latent codes are generated based on a third artificial neural network.

6 . The method of claim 5 , wherein the third artificial neural network is trained by an adversarial self-supervised training.

7 . The method of claim 6 , wherein the third artificial neural network is trained based on a plurality of RGB images.

8 . The method of claim 1 , wherein the first artificial neural network and/or the second artificial neural network are trained with regards to a loss function which is based on a pre-trained vision-language model supervision and the user command.

9 . The method of claim 8 , wherein a difference measure is determined between the user command and a descriptive text, which is generated by the pre-trained vision-language model, of a visual context in a rendered image based on the edited textured 3D morphable model of the object.

10 . The method of claim 8 , wherein the first artificial neural network and/or the second artificial neural network are trained based on a plurality of different user commands.

11 . The method of claim 1 , further comprising:

rendering the edited texture map of the object or the initial texture map of the object together with the edited 3D mesh model of the object or the initial 3D mesh model of the object to obtain an image corresponding to the user command.

12 . The method of claim 1 , further comprising

transcribing the user command, which is a speech input in human language, into a text.

13 . The method of claim 1 , further comprising:

obtaining a plurality of different initial 3D mesh models of the object; and

determining a plurality of edited texture maps of the object corresponding to the plurality of the initial 3D mesh models of the object and the user command by jointly editing the initial texture map of the object a plurality of times based on the first artificial neural network and based on and the plurality of initial 3D mesh models; and

generating a plurality of edited textured 3D morphable models of the object corresponding to the user command based on the edited texture maps of the object and on the plurality of initial 3D mesh models of the object; and

rendering the plurality of edited texture maps of the object together with the plurality of initial 3D mesh models of the object to obtain a plurality of images corresponding to the user command.

14 . A non-transitory computer-readable storage medium storing computer-readable instructions thereon which, when executed by a computer, cause the computer to carry out the steps of claim 1 .

15 . An electronic device for performing a user command-guided editing of an initial textured 3D morphable model, comprising:

processing circuitry configured to

obtain the initial textured 3D morphable model of the object comprising an initial texture map of the object and an initial 3D mesh model of the object; and

determine an edited texture map of the object corresponding to the user command by editing the initial texture map of the object based on a first artificial neural network; and/or

determine an edited 3D mesh model of the object corresponding to the user command by editing the initial 3D mesh model of the object based on a second artificial neural network (E); and

generate an edited textured 3D morphable model of the object corresponding to the user command based on the edited texture map of the object and/or the edited 3D mesh model of the object,

wherein the edited texture map is determined by a third artificial neural network based on a sum of the initial texture latent code and an offset texture latent code,

wherein the third artificial neural network is trained by adversarial self-supervised training on a plurality of two-dimensional RGB images using differentiable rendering.

16 . A method for user command-guided editing of an initial textured 3D morphable model of an object comprising:

obtaining the initial textured 3D morphable model of the object comprising an initial texture map of the object and an initial 3D mesh model of the object; and

determining an edited texture map of the object corresponding to the user command by editing the initial texture map of the object based on a first artificial neural network; and/or

determining an edited 3D mesh model of the object corresponding to the user command by editing the initial 3D mesh model of the object based on a second artificial neural network; and

generating an edited textured 3D morphable model of the object corresponding to the user command based on the edited texture map of the object and/or the edited 3D mesh model of the object,

wherein the third artificial neural network is trained using both a full-image discriminator and a patch discriminator that evaluates texture patches to generate high-frequency texture details.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: MARKHASIN, LEV; BELGACEM, IHEB
To: SONY EUROPE BV
Reel/Frame 070325/0310 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: SONY EUROPE BV
To: SONY SEMICONDUCTOR SOLUTIONS CORPORATION
Reel/Frame 070325/0486 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: ANEJA, SHIVANGI; NIEßNER, MATTHIAS; DAI, ANGELA
To: TECHNISCHE UNIVERSITÄT MÜNCHEN
Reel/Frame 070325/0567 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: TECHNISCHE UNIVERSITÄT MÜNCHEN
To: SONY EUROPE BV
Reel/Frame 070325/0605 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2025
From: SONY EUROPE BV
To: SONY SEMICONDUCTOR SOLUTIONS CORPORATION
Reel/Frame 070325/0662 →
Priority Claims (1)
EP 22211197 · Dec 2, 2022 · regional
Continuity (1)
Related Publication 20240193891A1 · Jun 13, 2024
References Cited (59)
US 20130106867A1 · Joo et al. · 2013 [cited by applicant]
US 20170256257A1 · Froelich · 2017 [cited by applicant]
US 20190035149A1 · Chen · 2019 [cited by examiner]
US 20220134225A1 · Omote · 2022 [cited by applicant]
Ramesh et al., “Zero-Shot Text-to-Image Generation”, arXiv:2102.12092v2 [cs.CV], Feb. 26, 2021, 20 pages. [cited by applicant]
Ramesh et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv:2204.06125v1 [cs.CV], Apr. 13, 2022, pp. 1-27. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Proceedings of the 38th International Conference on Machine Learning, PMLR, vol. 139, 2021, 16 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, pp. 1-48. [cited by applicant]
Lattas et al., “AvatarMe: Realistically Renderable 3D Facial Reconstruction “in-the-wild””, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 757-766. [cited by applicant]
Lattas et al., “AvatarMe++: Facial Shape and BRDF Inference With Photorealistic Rendering-Aware GANs”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 12, Dec. 2022, pp. 9269-9284. [cited by applicant]
Tewari et al., “StyleRig: Rigging StyleGAN for 3D Control over Portrait Images”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 6141-6150. [cited by applicant]
Tewari et al., “PIE: Portrait Image Embedding for Semantic Control”, ACM Trans. Graph., vol. 39, No. 6, Article 223. Publication date: Dec. 2020, pp. 223:1-223:14. [cited by applicant]
Gecer et al., “Synthesizing Coupled 3D Face Modalities by Trunk-Branch Generative Adversarial Networks”, arXiv:1909.02215v3 [cs.CV], Dec. 2, 2020, pp. 1-19. [cited by applicant]
Gecer et al., “Fast-GANFIT: Generative Adversarial Network for High Fidelity 3D Face Reconstruction”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 9, Sep. 2022, pp. 4879-4893. [cited by applicant]
Gecer et al., “GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 1155-1164. [cited by applicant]
Canfes et al., “Text and Image Guided 3D Avatar Generation and Manipulation”, arXiv:2202.06079v1 [cs.CV], Feb. 12, 2022, pp. 1-12. [cited by applicant]
Grassal et al., “Neural Head Avatars from Monocular RGB Videos”, arXiv:2112.01554v2 [cs.CV], Mar. 28, 2022, pp. 1-18. [cited by applicant]
Bau et al., “Paint by Word”, arXiv:2103.10951v2 [cs.CV], Mar. 24, 2021, 10 pages. [cited by applicant]
Laine et al., “Modular Primitives for High-Performance Differentiable Rendering”, arXiv:2011.03277v1 [cs.GR], Nov. 6, 2020, p. 194:1-194:14. [cited by applicant]
Hong et al., “AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars”, arXiv:2205.08535v1 [cs.CV], ACM Trans. Graph., vol. 41, No. 4, Article 161. Publication date: Jul. 2022, pp. 161:1-161:19. [cited by applicant]
Pavlakos et al., “Expressive Body Capture: 3D Hands, Face, and Body from a Single Image”, Computer Vision and Pattern Recognition (CVPR) 2019, pp. 1-4. [cited by applicant]
Romero et al., “Embodied Hands: Modeling and Capturing Hands and Bodies Together”, SIGGRAPH Asia 2017, Available Online at: https://mano.is.tue.mpg.de, pp. 1-5. [cited by applicant]
Osman et al., “SUPR:ASparseUnifiedPart-BasedHuman Representation”, Available Online At: https://github.com/ahmedosman/SUPR, ECCV 2022, pp. 1-18. [cited by applicant]
Aleksa Gordić, “OpenAI Whisper: Robust Speech Recognition via Large-Scale Weak Supervision | Paper and Code”, Available Online at: https://www.youtube.com/watch?v=AwJf8aQfChE, 2022, pp. 1-3. [cited by applicant]
“Introducing Whisper”, Available Online at: https://openai.com/blog/whisper/, Sep. 21, 2022, pp. 1-6. [cited by applicant]
Luo et al., “Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 11657-11667. [cited by applicant]
Crowson et al., “VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance”, 2021, pp. 1-17. [cited by applicant]
Youwang et al., “CLIP-Actor: Text-Driven Recommendation and Stylization for Animating Human Meshes”, In ECCV, 2022, pp. 1-18. [cited by applicant]
Kowalski et al., “CONFIG: Controllable Neural Face Image Generation” In European Conference on Computer Vision (ECCV), 2020, pp. 1-17. [cited by applicant]
Petrovich et al., “TEMOS: Generating diverse human motions from textual descriptions”, In European Conference on Computer Vision (ECCV), 2022, pp. 1-18. [cited by applicant]
Loper et al., “SMPL: A Skinned Multi-Person Linear Model”. ACM Trans. Graphics (Proc. SIGGRAPH Asia), vol. 34, No. 6, Oct. 2015, pp. 248:1-248:16. [cited by applicant]
Michel et al., “Text2Mesh: Text-Driven Neural Stylization for Meshes” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 18, 2022, pp. 13482-13492. [cited by applicant]
Lee et al., “StyleUV: Diverse and High-quality UV Map Generative Model”, arXiv:2011.12893v1 [cs.GR], Nov. 25, 2020, pp. 1-12. [cited by applicant]
Khalid et al., “CLIP-Mesh: Generating textured meshes from text using pretrained image-text models”, 22 Conference Papers, Dec. 6-9, 2022, 8 pages. [cited by applicant]
Ruiz et al., “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, arXiv:2208.12242v1 [cs.CV], Aug. 25, 2022, pp. 1-21. [cited by applicant]
Nikolay Jetchev, “ClipMatrix: Text-controlled Creation of 3D Textured Meshes”, arXiv:2109.12922v1 [cs.LG], Sep. 27, 2021, pp. 1-4. [cited by applicant]
Chakladar et al., “3D Avatar Approach for Continuous Sign Movement Using Speech/Text”, Applied Science, vol. 11, No. 3439, Available Online at: https://doi.org/10.339011083439, Apr. 12, 2021, pp. 1-13. [cited by applicant]
Avrahami et al., “Blended Diffusion for Text-driven Editing of Natural Images”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18187-18197. [cited by applicant]
Patashnik et al., “StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery”. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 2065-2074. [cited by applicant]
Michel et al., “Text2Mesh: Text-Driven Neural Stylization for Meshes” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 13482-13492. [cited by applicant]
Ghosh et al., “GIF: Generative Interpretable Faces”, In International Conference on 3D Vision (3DV), 2020., pp. 868-878. [cited by applicant]
Abdal et al., “CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions”, SIGGRAPH '22 Conference Proceedings, Aug. 7-11, 2022, 9 pages. [cited by applicant]
Abdal et al., “StyleFlow: Attribute-conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing Flows”, ACM Transactions on Graphics, vol. 40, No. 3, Article 21. Publication date: Apr. … [cited by applicant]
Marriott et al., “A 3D Gan for Improved Large-pose Facial Recognition”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 13440-13450. [cited by applicant]
Gal et al., “StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators”, arXiv:2108.00946v2 [cs.CV], Dec. 16, 2021, pp. 1-25. [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv:2112.10752v2 [cs.CV], Apr. 13, 2022, pp. 1-45. [cited by applicant]
Slossberg et al., “Unsupervised High-Fidelity Facial Texture Generation and Reconstruction”, arXiv:2110.04760v1 [cs.CV], Oct. 10, 2021, pp. 1-13. [cited by applicant]
Laine et al., “Modular Primitives for High-Performance Differentiable Rendering”, ACM Transactions on Graphics, vol. 39, No. 6, Dec. 2020, pp. 194:1-194:14. [cited by applicant]
Aneja et al., “ClipFace: Text-guided Editing of Textured 3D Morphable Models”, arXiv:2212.01406v1, Dec. 2, 2022, pp. 1-16. [cited by applicant]
Karras et al., “Training Generative Adversarial Networks with Limited Data”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), pp. 1-11. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4396-4405. [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of StyleGAN”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8107-8116. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, arXiv: 1710.10196v3 [cs.NE], Feb. 26, 2018, pp. 1-26. [cited by applicant]
Li et al., “Learning a model of facial shape and expression from 4D scans”, ACM Transactions on Graphics, vol. 36, No. 6, Article 194. Publication date: Nov. 2017, pp. 194:1-194:17. [cited by applicant]
Kocasari et al., “StyleMC: Multi-Channel Based Fast Text-Guided Image Generation and Manipulation”. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Jan. 2022, pp. 895-904. [cited by applicant]
Blanz et al., “A Morphable Model for the Synthesis of 3D Faces”, In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH '99, pp. 187-194. [cited by applicant]
Feng et al., “Learning an Animatable Detailed 3D Face Model from In-the-Wild Images”, ACM Trans. Graph., vol. 40, No. 4, Article 88. Publication date: Aug. 2021, pp. 88:1-88:13. [cited by applicant]
Deng et al., “Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 5153-5162. [cited by applicant]
Liu et al., “3D-FM GAN: Towards 3D-Controllable Face Manipulation”, arXiv:2208.11257v1 [cs.CV], Aug. 24, 2022, pp. 1-26. [cited by applicant]