IP Library › Granted Patent US 12,075,190
Granted Patent B2
US 12,075,190 · App. 18/221,702 · Granted Aug 27, 2024

Generating an image mask using machine learning

Inventors: Lidiia Bogdanovych (Los Angeles, CA); William Brendel (Los Angeles, CA); Samuel Edward Hare (Los Angeles, CA); Fedir Poliakov (Marina Del Rey, CA); Guohui Wang (Los Angeles, CA); Xuehan Xiong (Los Angeles, CA); Jianchao Yang (Los Angeles, CA); Linjie Yang (Los Angeles, CA)
Assignee: Snap Inc.
H04N7/147G06F18/214G06F18/24765G06N3/04G06N3/08G06T7/11G06T7/194G06V10/82G06V30/19173G06V30/242G06T2207/10016G06T2207/10024G06T2207/20024G06T2207/20081G06T2207/20084G06T2207/20221G06T2207/30201H04N5/44504H04N5/76H04N7/141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,075,190
App. No.
18/221,702
Granted
Aug 27, 2024
Kind
B2
Abstract

A machine learning system can generate an image mask (e.g., a pixel mask) comprising pixel assignments for pixels. The pixels can be assigned to classes, including, for example, face, clothes, body skin, or hair. The machine learning system can be implemented using a convolutional neural network that is configured to execute efficiently on computing devices having limited resources, such as mobile phones. The pixel mask can be used to more accurately display video effects interacting with a user or subject depicted in the image.

Claims (55)

1. A method comprising:

generating, by a processor of a device, an image of a user of the device;

generating an initial mask by applying a segmentation neural network to the image of the user;

refining borders of labeled areas of the initial mask by applying a post-processing engine to the initial mask to generate an output mask;

generating a modified image based on the output mask; and

storing the modified image on the device,

wherein the post-processing engine comprises a guided filter using a downsampled image of the image and a grayscale image of the downsampled image to generate the output mask,

wherein the guided filter further uses an eroded mask of the initial mask and a resized mask of the eroded mask to generate the output mask.

2. The method of claim 1 , wherein the post-processing engine applies the guided filter to a cropped area of the image.

3. The method of claim 1 , wherein the post-processing engine is configured to:

erode the initial mask to generate a new erode mask;

downsample the image to generate the downsampled image;

convert the downsampled image into the grayscale image; and

resize the new erode mask to generate a new resized mask image that matches a size of the grayscale image.

4. The method of claim 3 , wherein the post-processing engine is further configured to:

apply the guided filter based on the grayscale image and the new resized mask image to generate a new mask;

combining pixel information from both the new erode mask and the new mask to generate a combined mask;

apply thresholding operation on the combined mask; and

rescale pixel values in the combined mask to generate the output mask.

5. The method of claim 1 , wherein the image includes a portrait image of the user, the portrait image comprising a portrait background and a portrait foreground that depicts the user.

6. The method of claim 5 , wherein applying the segmentation neural network to the portrait image, the modified image displaying the portrait foreground depicting the user without the portrait background, the segmentation neural network trained on training data comprising a plurality of multi-labeled portrait images of different users, each multi-labeled portrait image depicting the portrait foreground that comprises a plurality of labeled user regions within the portrait foreground that corresponds to one of the different users depicted in the multi-labeled portrait image, the segmentation neural network trained to identify the portrait foreground in the portrait image by identifying each of the plurality of labeled user regions within the portrait image.

7. The method of claim 6 , wherein the training data to train the segmentation neural network further comprises a reduced size version of each of the plurality of multi-labeled portrait images, each reduced size version comprising a reduced size version of the plurality of labeled user regions within a reduced size portrait foreground, the reduced size versions being refined by using the reduced size versions of the plurality of multi-labeled portrait images as the guided filter,

wherein applying the segmentation neural network to the portrait image comprises rescaling the refined reduced size versions to a larger size to generate image masks.

8. The method of claim 7 , wherein the training data to train the segmentation neural network further comprises an increased size version of each of the plurality of multi-labeled portrait images, each increased size version comprising an increased size version of the plurality of labeled user regions within an increased size portrait foreground.

9. The method of claim 8 , wherein the device natively generates images at a same size as one of:

the plurality of multi-labeled portrait images,

the reduced size version of each of the plurality of multi-labeled portrait images, or

the increased size version of each of the plurality of multi-labeled portrait images.

10. The method of claim 7 , wherein each of the multi-labeled portrait images is labeled manually by users that label each of the plurality of labeled user regions in each image.

11. The method of claim 10 , wherein the labels correspond to regions of pixels that are assigned to one region of the plurality of labeled user regions.

12. The method of claim 7 , wherein the plurality of labeled user regions comprises one or more of: a hair area of a depicted user, a face area of the depicted user, and a clothes area of the depicted user.

13. The method of claim 5 , wherein the modified image is generated by applying an image effect using the portrait foreground.

14. The method of claim 13 , wherein the image effect modified an appearance of the portrait foreground that depicts the user.

15. The method of claim 1 , wherein the segmentation neural network comprises a convolutional neural network.

16. The method of claim 15 , wherein the convolutional neural network comprises a single deconvolutional layer.

17. The method of claim 1 , further comprising:

publishing the modified image as an ephemeral message on a network site.

18. A system comprising:

one or more processors of a machine; and

a memory storing instructions that, when executed by the one or more processors, cause the machine to perform operations comprising:

generating, by a processor of a device, an image of a user of the device;

generating an initial mask by applying a segmentation neural network to the image of the user;

refining borders of labeled areas of the initial mask by applying a post-processing engine to the initial mask to generate an output mask;

generating a modified image based on the output mask; and

storing the modified image on the device,

wherein the post-processing engine comprises a guided filter using a downsampled image of the image and a grayscale image of the downsampled image to generate the output mask,

wherein the guided filter further uses an eroded mask of the initial mask and a resized mask of the eroded mask to generate the output mask.

19. A non-transitory machine-readable storage device embodying instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

generating, by a processor of a device, an image of a user of the device;

generating an initial mask by applying a segmentation neural network to the image of the user;

refining borders of labeled areas of the initial mask by applying a post-processing engine to the initial mask to generate an output mask;

generating a modified image based on the output mask; and

storing the modified image on the device,

wherein the post-processing engine comprises a guided filter using a downsampled image of the image and a grayscale image of the downsampled image to generate the output mask,

wherein the guided filter further uses an eroded mask of the initial mask and a resized mask of the eroded mask to generate the output mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: BOGDANOVYCH, LIDIIA; BRENDEL, WILLIAM; EDWARD HARE, SAMUEL; POLIAKOV, FEDIR; WANG, GUOHUI; XIONG, XUEHAN; YANG, JIANCHAO; YANG, LINJIE
To: SNAP INC.
Reel/Frame 064438/0800 →
Continuity (5)
Continuation 16992968 · Aug 13, 2020
Continuation 16521956 · Jul 25, 2019
Continuation 15706057 · Sep 15, 2017
Provisional Application 62481415 · Apr 4, 2017
Related Publication 20230362331A1 · Nov 9, 2023
Cited By (1)
US 12,238,404