IP Library › Granted Patent US 11,908,041
Granted Patent B2
US 11,908,041 · App. 17/648,363 · Granted Feb 20, 2024

Object replacement system

Inventors: Viacheslav Ivanov (London, GB); Aleksei Zhuravlev (London, GB)
Assignee: Snap Inc.
G06T11/00G06V10/74G06V10/774G06V10/776G06V10/82G06V20/20G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,041
App. No.
17/648,363
Filed
Jan 19, 2022
Granted
Feb 20, 2024
Kind
B2
Art Unit
2613
USPC
345/633
Abstract

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and a method for performing operations comprising: receiving an image that includes a depiction of a real-world environment; processing the image to obtain data indicating presence of a real-world object in the real-world environment; receiving input that selects an AR experience comprising an AR object; determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object; applying a machine learning technique to the image to generate a new image that depicts the real-world environment without the real-world object; and applying the AR object to the new image to generate a modified new image that depicts the real-world environment including the AR object in place of the real-world object.

Claims (73)

1. A method comprising:

receiving, by one or more processors, an image that includes a depiction of a real-world environment;

processing the image to obtain data indicating presence of a real-world object in the real-world environment;

receiving input that selects an augmented reality (AR) experience comprising an AR object;

determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience;

in response to determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience, applying a machine learning technique to the image to generate a new image that depicts the real-world environment without the real-world object, the machine learning technique trained by:

receiving training data comprising a plurality of training images that depict real-world environments with a given real-world object and ground truth images that depict the real-world environments without the given real-world object;

applying the machine learning technique to a first training image of the plurality of training images to estimate an image that depicts the real-world environment without the given real-world object;

computing a deviation between the estimated image and a ground truth image associated with the first training image; and

updating parameters of the machine learning technique based on the computed deviation; and

applying the AR object to the new image to generate a modified new image that depicts the real-world environment including the AR object in place of the real-world object.

2. The method of claim 1 , further comprising:

accessing a plurality of machine learning techniques, a first of the plurality of machine learning techniques being configured to remove a first type of object from a first image that depicts a first environment including the first type of object, a second of the plurality of machine learning techniques being configured to remove a second type of object from a second image that depicts a second environment including the second type of object.

3. The method of claim 2 , further comprising:

determining a type associated with the real-world object detected on the real-world environment depicted in the image indicated in the obtained data; and

selecting the first machine learning technique from the plurality of machine learning techniques as the machine learning technique in response to determining that the type matches the first type of object.

4. The method of claim 1 , wherein the real-world object comprises at least one of hardware on a cabinet, a door on a house, a faucet on a sink, a wheel on a vehicle, eyeglasses, headwear, one or more earrings, one or more piercings, or a necklace, and wherein the AR object comprises at least one of AR hardware on a cabinet, an AR door on a house, an AR faucet on a sink, an AR wheel on a vehicle, AR eyeglasses, AR headwear, one or more AR earrings, one or more AR piercings, or an AR necklace.

5. The method of claim 1 , wherein applying the AR object to the new image comprises adding a depiction of the AR object to the modified new image.

6. The method of claim 1 , wherein the machine learning technique comprises a first neural network.

7. The method of claim 6 , further comprising:

training a second neural network from the first neural network using a knowledge distillation process, the second neural network being a compact version of the first neural network and being configured to be implemented on a mobile device.

8. The method of claim 6 , wherein the deviation is computed based on perceptual loss.

9. The method of claim 6 , further comprising generating the training data by performing operations comprising:

generating the first training image that depicts a first real-world environment;

determining that the first training image includes a depiction of the given real-world object on the first real-world environment;

applying a generative model comprising a latent code modification process to remove the given real-world object from the first real-world environment, an output of the generative model comprising an intermediate image that depicts the first real-world environment without the given real-world object; and

generating the ground truth image corresponding to the first training image based on the intermediate image.

10. The method of claim 9 , wherein the ground truth image corresponding to the first training image is generated by performing operations comprising:

applying an environment segmentation model on the first training image to generate a binary mask corresponding to the given real-world object; and

combining feature representations of a region of interest in the intermediate image, where the region of interest is given by the binary mask, with the feature representations of the first training image to generate a blended representation as a ground truth image corresponding to the first training image.

11. The method of claim 10 , further comprising verifying integrity of the ground truth image corresponding to the first training image by performing operations comprising:

applying an attribute extractor on the first training image to generate a first set of attributes;

applying the attribute extractor on the ground truth image corresponding to the first training image to generate a second set of attributes; and

determining that one or more verification criteria are satisfied by comparing the first and second sets of attributes.

12. The method of claim 11 , wherein the one or more verification criteria comprises:

an indication that the first set of attributes indicates presence of the given real-world object and that the second set of attributes indicates absence of the given real-world object;

a demographic of the first set of attributes matching the demographic in the second set of attributes;

a scene illumination of the first set of attributes matching the scene illumination in the second set of attributes; and

a pose of the first set of attributes matching the pose in the second set of attributes.

13. The method of claim 12 , wherein the one or more verification criteria further comprises a facial expression of the first set of attributes matching the facial expression in the second set of attributes.

14. The method of claim 9 , further comprising:

computing a first quantity of images in the training data that is associated with a first demographic;

computing a second quantity of images in the training data that is associated with a second demographic; and

selecting a new image associated with the first demographic for adding to the training data in response to determining that the second quantity exceeds the first quantity by more than a threshold value.

15. The method of claim 9 , further comprising searching attributes associated with a collection of images to identify the first training image that is associated with an attribute indicating presence of the given real-world object.

16. The method of claim 1 , wherein the image is a frame of a real-time video feed received from a client device of a person.

17. The method of claim 1 , further comprising:

receiving a video comprising the image depicting the real-world object being moved from a first position to a second position, the second position being closer in proximity to the real-world environment than the first position; and

as the real-world object is being moved, applying the machine learning technique to gradually fade and remove the real-world object from the video.

18. A system comprising:

at least one processor configured to perform operations comprising:

receiving an image that includes a depiction of a real-world environment;

processing the image to obtain data indicating presence of a real-world object in the real-world environment;

receiving input that selects an augmented reality (AR) experience comprising an AR object;

determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience;

in response to determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience, applying a machine learning technique to the image to generate a new image that depicts the real-world environment without the real-world object, the machine learning technique trained by:

receiving training data comprising a plurality of training images that depict real-world environments with a given real-world object and ground truth images that depict the real-world environments without the given real-world object;

applying the machine learning technique to a first training image of the plurality of training images to estimate an image that depicts the real-world environment without the given real-world object;

computing a deviation between the estimated image and a ground truth image associated with the first training image; and

updating parameters of the machine learning technique based on the computed deviation; and

applying the AR object to the new image to generate a modified new image that depicts the real-world environment including the AR object in place of the real-world object.

19. The system of claim 18 , wherein the operations further comprise accessing a plurality of machine learning techniques, a first of the plurality of machine learning techniques being configured to remove a first type of object from a first image that depicts a first environment including the first type of object, a second of the plurality of machine learning techniques being configured to remove a second type of object from a second image that depicts a second environment including the second type of object.

20. A non-transitory machine-readable storage medium that includes instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving an image that includes a depiction of a real-world environment;

processing the image to obtain data indicating presence of a real-world object in the real-world environment;

receiving input that selects an augmented reality (AR) experience comprising an AR object;

determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience;

in response to determining that the real-world object detected in the real-world environment depicted in the image indicated in the obtained data corresponds to the AR object of the AR experience, applying a machine learning technique to the image to generate a new image that depicts the real-world environment without the real-world object, the machine learning technique trained by:

receiving training data comprising a plurality of training images that depict real-world environments with a given real-world object and ground truth images that depict the real-world environments without the given real-world object;

applying the machine learning technique to a first training image of the plurality of training images to estimate an image that depicts the real-world environment without the given real-world object;

computing a deviation between the estimated image and a ground truth image associated with the first training image; and

updating parameters of the machine learning technique based on the computed deviation; and

applying the AR object to the new image to generate a modified new image that depicts the real-world environment including the AR object in place of the real-world object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: IVANOV, VIACHESLAV; ZHURAVLEV, ALEKSEI
To: SNAP INC.
Reel/Frame 058697/0545 →
Continuity (1)
Related Publication 20230230292A1 · Jul 20, 2023