IP Library Granted Patent US 11,138,781
Granted Patent B1
US 11,138,781 · App. 16/837,573 · Granted Oct 5, 2021

Creation of photorealistic 3D avatars using deep neural networks

Inventors: Jeb R. Linton (Manassas, VA); Satya Sreenivas (Los Alamos, NM); Naeem Atlaf (Round Rock, TX); Sanjay Nadhavajhala (Cupertino, CA)
Assignee: International Business Machines Corporation
G06T13/40G06N3/08G06T3/4053G06T17/20G10L13/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,138,781
App. No.
16/837,573
Granted
Oct 5, 2021
Kind
B1
Abstract

A computer-implemented method for the automatic generation of photorealistic 3D avatars. The method includes generating, using machine learning network, a plurality of 2D photorealistic facial images; generating, using a model agnostic meta-learner, a plurality of 2D photorealistic facial expression images based on the plurality of 2D photorealistic facial images; and generating a plurality of 3D photorealistic avatars based on the plurality of 2D photorealistic facial expression images.

Claims (29)

1. A method for automatic generation of photorealistic 3D avatars, the method comprising:

generating, using a machine learning network, a plurality of 2D photorealistic artificial human-like facial images;

generating, using a model agnostic meta-learner, a plurality of 2D photorealistic artificial human-like facial expression images based on the plurality of 2D photorealistic artificial human-like facial images, wherein the model agnostic meta-learner employs a Generator/Discriminator/Embedder architecture to create a comprehensive set of expressions with various mouth and eye positions; and

generating a plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images.

2. The method of claim 1 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images comprises transforming the plurality of 2D photorealistic artificial human-like facial expression images by applying digital coordinate transform to a stretched skin-sized and shaped 2D facial expression image for wrapping around a wireframe of a head and shoulders of an avatar.

3. The method of claim 2 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises filling in gaps in a back of the head using a U-Network trained on a set of photographic skins and their equivalent with the back of the head removed.

4. The method of claim 3 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises filling in hair at the back of the head based on a color, style, and length of the hair in an original image.

5. The method of claim 4 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises applying a super-resolution technique to sharpen the 3D photorealistic artificial human-like avatars.

6. The method of claim 1 , further comprising generating a voice generation network for converting a real voice to that of an avatar.

7. The method of claim 1 , wherein the model agnostic meta-learner employs the Generator/Discriminator/Embedder architecture is used to animate a face.

8. The method of claim 1 , wherein the machine learning network used to generate the plurality of 2D photorealistic artificial human-like facial images is a Generative Adversarial Network (GAN).

9. The method of claim 1 , wherein the machine learning network used to generate the plurality of 2D photorealistic artificial human-like facial images is a Style Transfer Network.

10. The method of claim 1 , wherein the machine learning network used to generate the plurality of 2D photorealistic artificial human-like facial images is a combination of a Generative Adversarial Network (GAN) and a Style Transfer Network.

11. A system configured to automatically generate photorealistic 3D avatars, the system comprising memory for storing instructions, and a processor configured to execute the instructions to:

generate a plurality of 2D photorealistic artificial human-like facial images by a machine learning network;

generate a plurality of 2D photorealistic artificial human-like facial expression images based on the plurality of 2D photorealistic artificial human-like facial images using a model agnostic meta-learner, wherein the model agnostic meta-learner employs a Generator/Discriminator/Embedder architecture to create a comprehensive set of expressions with various mouth and eye positions; and

generate a plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images.

12. The system of claim 11 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images comprises transforming the plurality of 2D photorealistic artificial human-like facial expression images by applying digital coordinate transform to a stretched skin-sized and shaped 2D facial expression image for wrapping around a wireframe of a head and shoulders of an avatar.

13. The system of claim 12 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises filling in gaps in a back of the head using a U-Network trained on a set of photographic skins and their equivalent with the back of the head removed.

14. The system of claim 13 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises applying a deep learning technique to fill in hair at the back of the head based on a color, style, and length of the hair in an original image.

15. The system of claim 14 , wherein generating the plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images further comprises applying a super-resolution technique to sharpen the 3D photorealistic artificial human-like avatars.

16. The system of claim 11 , the processor is further configured to execute the instructions to generate a voice generation network for converting a real voice to that of an avatar.

17. The system of claim 11 , wherein the model agnostic meta-learner employs the Generator/Discriminator/Embedder architecture to animate a face.

18. The system of claim 11 , wherein the machine learning network used to generate the plurality of 2D photorealistic artificial human-like facial images is a generative Adversarial Network (GAN).

19. The system of claim 11 , wherein the machine learning network used to generate the plurality of 2D photorealistic artificial human-like facial images is a Style Transfer Network.

20. A computer program product for automatically generating photorealistic 3D avatars, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of a system to cause the system to:

generate a plurality of 2D photorealistic artificial human-like facial images by a machine learning network;

generate a plurality of 2D photorealistic artificial human-like facial expression images based on the plurality of 2D photorealistic artificial human-like facial images using a model agnostic meta-learner, wherein the model agnostic meta-learner employs a Generator/Discriminator/Embedder architecture to create a comprehensive set of expressions with various mouth and eye positions; and

generate a plurality of 3D photorealistic artificial human-like avatars based on the plurality of 2D photorealistic artificial human-like facial expression images.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2025
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: BLUE HERON DEVELOPMENT LLC
Reel/Frame 070130/0844 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRECT THE CONVEYING PARTY DATA TO CORRECT ASSIGNOR NAEEM ATLAF TO NAEEM ALTAF PREVIOUSLY RECORDED AT REEL: 52287 FRAME: 750. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 12, 2024
From: LINTON, JEB R.; SREENIVAS, SATYA; ALTAF, NAEEM; NADHAVAJHALA, SANJAY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069336/0323 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2020
From: LINTON, JEB R.; SREENIVAS, SATYA; ATLAF, NAEEM; NADHAVAJHALA, SANJAY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052287/0750 →
Cited By (1)
US 12,536,739