IP Library › Granted Patent US 11,727,617
Granted Patent B2
US 11,727,617 · App. 17/695,902 · Granted Aug 15, 2023

Single image-based real-time body animation

Inventors: Egor Nemchinov (Saint-Petersburg, RU); Sergei Gorbatyuk (Saint-Petersburg, RU); Aleksandr Mashrabov (Sochi, RU); Egor Spirin (Saint-Petersburg, RU); Iaroslav Sokolov (Saint-Petersburg, RU); Andrei Smirdin (Saint-Petersburg, RU); Igor Tukh (Saint-Petersburg, RU)
Assignee: Snap Inc.
G06T13/40G06N3/08G06T3/0037G06T7/194G06T15/04G06T17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,617
App. No.
17/695,902
Filed
Mar 16, 2022
Granted
Aug 15, 2023
Kind
B2
Art Unit
2612
USPC
345/473
Abstract

Disclosed are systems and methods for single image-based body animation. An example method includes receiving an input image, the input image including a body image of a person, extracting the body image of the person from the input image, fitting a generic model to the body image, where the generic model is configured to receive a set of pose parameters corresponding to a pose of the person and generate a generic body shape adopting the pose, generating a three-dimensional (3D) model, where the 3D model is configured to receive a set of further pose parameters corresponding to the pose of the person and generate an output image of the person adopting the pose, the output image including a feature of the body image being omitted from the generic body shape, and providing a further set of further pose parameters to generate a frame of an output video.

Claims (86)

1. A method for single image-based body animation, the method comprising:

receiving, by a computing device, an input image, the input image including a body image of a person;

extracting, by the computing device, the body image of the person from the input image;

fitting, by the computing device, a generic model to the body image, wherein the generic model is configured to:

receive a set of pose parameters corresponding to a pose of the person; and

generate, based on the set of pose parameters, a generic body shape adopting the pose;

generating, by the computing device and based on the body image and the generic model, a three-dimensional (3D) model, wherein the 3D model is configured to:

receive a set of further pose parameters corresponding to the pose of the person; and

generate, based on the set of further pose parameters, an output image of the person adopting the pose, the output image including at least one feature of the body image of the person being omitted from the generic body shape; and

providing, by the computing device to the 3D model, at least one further set of further pose parameters corresponding to a further pose of the person to generate a frame of an output video, the frame including a further output image of the person adopting the further pose.

2. The method of claim 1 , wherein the at least one feature of the body image of the person omitted from the generic body shape includes one or more of the following: hair, a position of a finger on a hand, and at least one piece of clothing.

3. The method of claim 1 , wherein:

a parameter of the set of pose parameters determines axis-angle rotations of at least one joint in the generic body shape; and

the generic model depends on a set of shape parameters, the set of shape parameters including a vector of 3D points corresponding to the generic body shape.

4. The method of claim 3 , wherein:

a portion of the set of pose parameters is trained on a first dataset; and

the set of shape parameters is trained on a second dataset.

5. The method of claim 4 , wherein:

the first dataset includes 3D scans of a first set of people adopting different poses; and

the second dataset includes 3D scans of a second set of people having different body shapes.

6. The method of claim 3 , wherein the set of pose parameters includes head pose parameters associated with the head of the person, the head pose parameters being trained on a third dataset, the third dataset including images of facial shapes and facial expressions of different people.

7. The method of claim 1 , wherein:

the generic model includes a first mesh, the first mesh including first 3D points corresponding to the generic body shape;

the 3D model includes a second mesh, the second mesh including second 3D points corresponding to the body image of the person; and

the second mesh is obtained by warping the first mesh to fit boundaries of the body image.

8. The method of claim 7 , wherein:

the generic model includes a first texture map, the first texture map being used to texture faces of the first mesh;

the 3D model includes a second texture map, the second texture map being used to texture faces of the second mesh; and

the second texture map is generated based on the first texture map.

9. The method of claim 8 , wherein the generation of the second texture map includes:

unwarping the first mesh to generate a two-dimensional (2D) representation of the first texture map, the two-dimensional (2D) representation including a set of parts; and

performing the following for a face of the second mesh:

determining coordinates of three edge points corresponding to points A, B, and C in the 2D representation;

determining that the points A, B, and C belong to the same part of the 2D representation; and

in response to the determining that the points A, B, and C belong to the same part of the 2D representation, using a triangle of the first texture map, the triangle being formed by the points A, B, and C, to generate a portion of the second texture map for texturing the face of the second mesh.

10. The method of claim 9 , wherein the generation of the second texture map include, after the determination of the coordinates of the points A, B, and C:

determining that the points A, B, and C belong to different parts of the 2D representation; and

in response to the determination that the points A, B, and C belong to different parts of the 2D representation:

splitting a triangle formed by points A, B, and C into a first triangle and a second triangle, wherein the first triangle belongs to a first part of the 2D presentation and the second triangle belong to a second part of the 2D presentation;

using the first triangle of the first texture map to generate a first portion of the second texture map for texturing the face of the second mesh; and

using the second triangle of the first texture map to generate a second portion of the second texture map for texturing the face of the second mesh.

11. A system for single image-based body animation, the system comprising at least one processor, a memory storing processor-executable codes, wherein the at least one processor is configured to implement the following operations upon executing the processor-executable codes:

receiving, by a computing device, an input image, the input image including a body image of a person;

extracting, by the computing device, the body image of the person from the input image;

fitting, by the computing device, a generic model to the body image, wherein the generic model is configured to:

receive a set of pose parameters corresponding to a pose of the person; and

generate, based on the set of pose parameters, a generic body shape adopting the pose;

generating, by the computing device and based on the body image and the generic model, a three-dimensional (3D) model, wherein the 3D model is configured to:

receive a set of further pose parameters corresponding to the pose of the person; and

generate, based on the set of further pose parameters, an output image of the person adopting the pose, the output image including at least one feature of the body image of the person being omitted from the generic body shape; and

providing, by the computing device to the 3 D model, at least one further set of further pose parameters corresponding to a further pose of the person to generate a frame of an output video, the frame including a further output image of the person adopting the further pose.

12. The system of claim 11 , wherein the at least one feature of the body image of the person omitted from the generic body shape includes one or more of the following: hair, a position of a finger on a hand, and at least one piece of clothing.

13. The system of claim 11 , wherein:

a parameter of the set of pose parameters determines axis-angle rotations of at least one joint in the generic body shape; and

the generic model depends on a set of shape parameters, the set of shape parameters including a vector of 3D points corresponding to the generic body shape.

14. The system of claim 13 , wherein:

a portion of the set of pose parameters is trained on a first dataset; and

the set of shape parameters is trained on a second dataset.

15. The system of claim 14 , wherein:

the first dataset includes 3D scans of a first set of people adopting different poses; and

the second dataset includes 3D scans of a second set of people having different body shapes.

16. The system of claim 13 , wherein the set of pose parameters includes head pose parameters associated with the head of the person, the head pose parameters being trained on a third dataset, the third dataset including images of facial shapes and facial expressions of different people.

17. The system of claim 11 , wherein:

the generic model includes a first mesh, the first mesh including first 3D points corresponding to the generic body shape;

the 3D model includes a second mesh, the second mesh including second 3D points corresponding to the body image of the person; and

the second mesh is obtained by warping the first mesh to fit boundaries of the body image.

18. The system of claim 17 , wherein:

the generic model includes a first texture map, the first texture map being used to texture faces of the first mesh;

the 3D model includes a second texture map, the second texture map being used to texture faces of the second mesh; and

the second texture map is generated based on the first texture map.

19. The system of claim 18 , wherein the generation of the second texture map includes:

unwarping the first mesh to generate a two-dimensional (2D) representation of the first texture map, the two-dimensional (2D) representation including a set of parts; and

performing the following for a face of the second mesh:

determining coordinates of three edge points corresponding to points A, B, and C in the 2D representation;

determining that the points A, B, and C belong to the same part of the 2D representation; and

in response to the determining that the points A, B, and C belong to the same part of the 2D representation, using a triangle of the first texture map, the triangle being formed by the points A, B, and C, to generate a portion of the second texture map for texturing the face of the second mesh.

20. A non-transitory processor-readable medium having instructions stored thereon, which when executed by one or more processors, cause the one or more processors to implement a method for single image-based body animation, the method comprising:

receiving, by a computing device, an input image, the input image including a body image of a person;

extracting, by the computing device, the body image of the person from the input image;

fitting, by the computing device, a generic model to the body image, wherein the generic model is configured to:

receive a set of pose parameters corresponding to a pose of the person; and

generate, based on the set of pose parameters, a generic body shape adopting the pose;

generating, by the computing device and based on the body image and the generic model, a three-dimensional (3D) model, wherein the 3D model is configured to:

receive a set of further pose parameters corresponding to the pose of the person; and

generate, based on the set of further pose parameters, an output image of the person adopting the pose, the output image including at least one feature of the body image of the person being omitted from the generic body shape; and

providing, by the computing device to the 3D model, at least one further set of further pose parameters corresponding to a further pose of the person to generate a frame of an output video, the frame including a further output image of the person adopting the further pose.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2023
From: NEMCHINOV, EGOR; GORBATYUK, SERGEI; MASHRABOV, ALEKSANDR; SPIRIN, EGOR; SOKOLOV, IAROSLAV; SMIRDIN, ANDREI; TUKH, IGOR
To: SNAP INC.
Reel/Frame 063984/0036 →
Continuity (3)
Continuation 17062309 · Oct 2, 2020
Continuation 16434185 · Jun 7, 2019
Related Publication 20220207810A1 · Jun 30, 2022