IP Library Granted Patent US 11,430,247
Granted Patent B2
US 11,430,247 · App. 16/949,773 · Granted Aug 30, 2022

Image generation using surface-based neural synthesis

Inventors: Iason Kokkinos (London, GB); Georgios Papandreou (London, GB); Riza Alp Guler (London, GB)
Assignee: Snap Inc.
G06V40/103G06F17/18G06K9/629G06K9/6215G06T5/50G06V20/647G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,430,247
App. No.
16/949,773
Granted
Aug 30, 2022
Kind
B2
Abstract

Aspects of the present disclosure involve a system and a method for performing operations comprising: receiving a two-dimensional continuous surface representation of a three-dimensional object, the continuous surface comprising a plurality of landmark locations; determining a first set of soft membership functions based on a relative location of points in the two-dimensional continuous surface representation and the landmark locations; receiving a two-dimensional input image, the input image comprising an image of the object; extracting a plurality of features from the input image using a feature recognition model; generating an encoded feature representation of the extracted features using the first set of soft membership functions; generating a dense feature representation of the extracted features from the encoded representation using a second set of soft membership functions; and processing the second set of soft membership functions and dense feature representation using a neural image decoder model to generate an output image.

Claims (62)

1. A computer implemented method of neural image synthesis, the method comprising:

receiving a two-dimensional continuous surface representation of a three-dimensional object, the continuous surface comprising a plurality of landmark locations;

determining a first set of soft membership functions based on a relative location of points in the two-dimensional continuous surface representation and the landmark locations;

receiving a two-dimensional input image, the input image comprising an image of the object;

extracting a plurality of features from the input image using a feature recognition model;

generating an encoded feature representation of the extracted features using the first set of soft membership functions;

generating a dense feature representation of the extracted features from the encoded representation using a second set of soft membership functions;

processing the second set of soft membership functions and dense feature representation using a neural image decoder model to generate an output image; and

causing presentation of the output image on a client device.

2. The method of claim 1 , wherein determining the first set of soft membership functions comprises:

determining distances between a plurality of points in the two-dimensional continuous surface representation and the landmark locations; and

assigning each point in the plurality of points to a landmark based on the determined distances.

3. The method of claim 1 , further comprising determining the landmark locations using a landmark recognition model.

4. The method of claim 1 , wherein the neural image decoder model comprises a convolutional neural network conditioned on the two-dimensional continuous surface representation.

5. The method of claim 1 , wherein generating an encoded feature representation of the extracted features using the first set of soft membership functions comprises performing a membership-weighted estimate of a mean and variance for each channel of the extracted features.

6. The method of claim 5 , wherein generating a dense feature representation of the extracted features from the encoded representation using a second set of soft membership functions comprises applying a dual operation to the membership-weighted estimate of a mean and variance for each channel of the extracted features.

7. The method of claim 1 , wherein the object is a human body and wherein the landmarks comprise joints of the human body.

8. The method of claim 1 , wherein the first set of soft membership functions and the second set of soft membership functions are the same.

9. The method of claim 8 , further comprising:

generating a three-dimensional model of the three-dimensional object from the input image; and

generating the two-dimensional continuous surface representation from the three-dimensional model.

10. The method of any of claim 9 , further comprising modifying values in the encoded representation prior to generating the dense feature representation.

11. The method of claim 1 , wherein the two-dimensional continuous surface representation of a three-dimensional object is generated from the input image, and wherein the method further comprises:

receiving a further two-dimensional input image, the further input image comprising a further object of the same type as the three-dimensional object;

generating a further two-dimensional continuous surface representation of said three-dimensional object from the further two-dimensional input image, the further continuous surface comprising the plurality of landmark locations; and

determining the second set of soft membership functions based on relative locations of points in the further two-dimensional continuous surface representation and the landmark locations.

12. The method of claim 11 , wherein the input image comprises an image of the object in a first pose and the further input image comprises an image of the further object in a second pose, wherein the second image comprises portions corresponding to unseen portions of the first image; and

wherein generating the two-dimensional continuous surface representation from the input image comprises generating portions of the two-dimensional continuous surface representation corresponding to the unseen portions of the first image from the encoded representation using a learned attention mechanism.

13. The method of claim 12 , wherein the learned attention mechanism is based on the first set of soft membership functions.

14. The method of claim 1 , further comprising:

determining a set of content features from a source image using an encoder neural network;

determining a set of style features from a style image using the encoder neural network;

determining position dependent content features using joint statistics of position and content features in regions of the source image;

determining position dependent style features using joint statistics of position and style features in regions of the style image;

generating a set of transformed content features from the set of position dependent content features based on the joint statistics of position and content features;

generating a set of transformed style features from the set of transformed content features based on the joint statistics of position and style features; and

generating another output image from the transformed set of style features and the transformed set of content features using a decoder neural network.

15. The method of claim 14 , further comprising combining the output image with the another output image to generate a combined image.

16. The method of claim 14 , wherein the joint statistics of position and content features comprise a content feature mean, a content position mean and covariances between content features and content positions, and wherein determining position dependent content features comprises determining a conditional model of the content features conditioned on position.

17. The method of claim 16 , wherein the conditional model of the content features comprises a position dependent content mean and a conditional content covariance.

18. The method of claim 17 , wherein generating the set of transformed content features from the set of position dependent content features comprises:

centring the position dependent content features based on the position dependent content mean; and

applying a whitening transformation based on the conditional content covariance.

19. A system for neural image analysis, comprising:

a processor configured to perform operations comprising:

receiving a two-dimensional continuous surface representation of a three-dimensional object, the continuous surface comprising a plurality of landmark locations;

determining a first set of soft membership functions based on a relative location of points in the two-dimensional continuous surface representation and the landmark locations;

receiving a two-dimensional input image, the input image comprising an image of the object;

extracting a plurality of features from the input image using a feature recognition model;

generating an encoded feature representation of the extracted features using the first set of soft membership functions;

generating a dense feature representation of the extracted features from the encoded representation using a second set of soft membership functions;

processing the second set of soft membership functions and dense feature representation using a neural image decoder model to generate an output image; and

causing presentation of the output image on a client device.

20. A non-transitory machine-readable storage medium that includes instructions that, when executed by one or more processors of a machine, cause the machine to perform operations for neural image analysis comprising:

receiving a two-dimensional continuous surface representation of a three-dimensional object, the continuous surface comprising a plurality of landmark locations;

determining a first set of soft membership functions based on a relative location of points in the two-dimensional continuous surface representation and the landmark locations;

receiving a two-dimensional input image, the input image comprising an image of the object;

extracting a plurality of features from the input image using a feature recognition model;

generating an encoded feature representation of the extracted features using the first set of soft membership functions;

generating a dense feature representation of the extracted features from the encoded representation using a second set of soft membership functions;

processing the second set of soft membership functions and dense feature representation using a neural image decoder model to generate an output image; and

causing presentation of the output image on a client device.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: ARIEL AI, INC.
To: ARIEL AI, LLC
Reel/Frame 060497/0239 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: ARIEL AI, LLC
To: SNAP INTERMEDIATE INC.
Reel/Frame 060497/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: SNAP INTERMEDIATE INC.
To: SNAP INC.
Reel/Frame 060498/0524 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2021
From: KOKKINOS, IASON; PAPANDREOU, GEORGIOS; GULER, RIZA ALP
To: ARIEL AI LTD
Reel/Frame 056804/0387 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2021
From: ARIEL AI LTD
To: ARIEL AI, INC.
Reel/Frame 056804/0478 →
Continuity (2)
Provisional Application 62936328 · Nov 15, 2019
Related Publication 20210150197A1 · May 20, 2021
Cited By (2)
US 12,354,211 US 12,380,611