IP Library Granted Patent US 11,164,384
Granted Patent B2
US 11,164,384 · App. 16/521,359 · Granted Nov 2, 2021

Mobile device image item replacements

Inventors: Xiaoyi Huang (Palo Alto, CA); Jingwen Wang (Palo Alto, CA); Yi Wu (Palo Alto, CA); Xin Ai (Palo Alto, CA)
Assignee: Houzz, Inc.
G06T19/006G06F3/0482G06F3/04883G06T15/506G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,164,384
App. No.
16/521,359
Granted
Nov 2, 2021
Kind
B2
Abstract

A system for replacing physical items in images is discussed. A depicted item can be selected and removed from an image via image mask data and pixel merging techniques. Virtual light source positions can be generated based on real-world light source data from the image. A rendered simulation of a virtual item can then be integrated into the image to create a modified image for display.

Claims (62)

1. A method comprising:

generating, using one or more processors of a mobile device, an image of a physical environment;

receiving, on a touchscreen of the mobile device, a selection of an object to be replaced in the image;

classifying the object into an object category using an object classification neural network;

selecting a pose detection neural network from a plurality of pose detection neural networks based on the object being classified in the object category, each of the plurality of pose detection neural networks being trained for different types of objects;

determining a three-dimensional orientation of the object as depicted within the image using the pose detection neural network comprising a convolutional neural network trained to detect three-dimensional orientation of objects in a plurality of object training images, the objects of the plurality of object training images being of a same type as the object detected in the image;

removing, from the image, the object using regions that are proximate to the object in the image;

generating a render of a virtual model in the three-dimensional orientation and as illuminated by one or more virtual light sources based on a lighting scheme in the image; and

generating a modified image that depicts the render replacing the object in the physical environment.

2. The method of claim 1 , further comprising:

determining the lighting scheme of the image.

3. The method of claim 2 , wherein determining the lighting scheme comprises determining one or more bright regions of the image.

4. The method of claim 3 , further comprising:

positioning, in a virtual environment, the one or more virtual light sources based on locations of the one or more bright regions of the image.

5. The method of claim 3 , wherein the determining of the one or more bright regions of the image comprises determining an area of pixels in the image having higher brightness values.

6. The method of claim 1 , wherein, in the image, the object is depicted in an object image region, and the regions that are proximate to the object in the image are proximate regions that are external to the object image region.

7. The method of claim 6 , wherein the object is removed by merging the proximate regions and the object image region.

8. The method of claim 7 , wherein the proximate regions and the object image region are merged using a neural network that implements partial convolutional layers.

9. The method of claim 6 , wherein the object is removed by interpolating the proximate regions and the object image region.

10. The method of claim 1 , further comprising:

displaying the image on a display device of the mobile device; and

receiving selection of the object through the display device of the mobile device.

11. The method of claim 10 , wherein receiving selection of the object comprises receiving selection of a selected region of the image that depicts the object.

12. The method of claim 11 , further comprising:

generating an image mask using the selected region.

13. The method of claim 11 , further comprising:

segmenting the image into segment regions using an image segmentation convolutional neural network (CNN), wherein the selected region is identified from a user input on the image as displayed on the touchscreen of the mobile device.

14. The method of claim 13 , wherein the user input is one of: a tap gesture or a click.

15. The method of claim 11 , wherein receiving selection of the object through the display device comprises:

receiving, on the touchscreen of the mobile device, a swipe gesture over at least a portion of the object as depicted in the image.

16. The method of claim 1 , further comprising:

receiving, on the touchscreen of the mobile device, an additional selection that selects an additional object to be replaced in the image, the additional object and the object being different types of objects;

classifying the additional object into another object category using the object classification neural network;

selecting another pose detection neural network from the plurality of pose detection neural networks based on the additional object being classified in the another object category, wherein each of the plurality of pose detection neural networks are trained for different types of objects;

determining an additional three-dimensional orientation of the additional object as depicted within the image using the another pose detection neural network;

removing, from the image, the additional object from the image;

generating an additional render of an additional virtual model in the additional three-dimensional orientation and as illuminated by the one or more virtual light sources based on the lighting scheme in the image; and

generating a modified image that depicts the additional render replacing the additional object.

17. A system comprising:

one or more processors;

a touchscreen

a memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

generating an image of a physical environment;

receiving, on the touchscreen, a selection of an object to be replaced in the image;

classifying the object into an object category using an object classification neural network;

selecting a pose detection neural network from a plurality of pose detection neural networks based on the object being classified in the object category, each of the plurality of pose detection neural networks being trained for different types of objects;

determining a three-dimensional orientation of the object as depicted within the image using the pose detection neural network comprising a convolutional neural network trained to detect three-dimensional orientation of objects in a plurality of object training images, the objects of the plurality of object training images being of a same type as the object detected in the image;

removing, from the image, the object using regions that are proximate to the object in the image;

generating a render of a virtual model in the three-dimensional orientation and as illuminated by one or more virtual light sources based on a lighting scheme in the image; and

generating a modified image that depicts the render replacing the object in the physical environment.

18. The system of claim 17 , the operations further comprising:

determining the lighting scheme of the image.

19. The system of claim 18 , wherein determining the lighting scheme comprises determining one or more bright regions of the image.

20. A machine-readable storage device embodying instructions that, when executed by a device, cause the device to perform operations comprising:

generating an image of a physical environment;

receiving, on a touchscreen, a selection of an object to be replaced in the image;

classifying the object into an object category using an object classification neural network;

selecting a pose detection neural network from a plurality of pose detection neural networks based on the object being classified in the object category, each of the plurality of pose detection neural networks being trained for different types of objects;

determining a three-dimensional orientation of the object as depicted within the image using the pose detection neural network comprising a convolutional neural network trained to detect three-dimensional orientation of objects in a plurality of object training images, the objects of the plurality of object training images being of a same type as the object detected in the image;

removing, from the image, the object using regions that are proximate to the object in the image;

generating a render of a virtual model in the three-dimensional orientation and as illuminated by one or more virtual light sources based on a lighting scheme in the image; and

generating a modified image that depicts the render replacing the object in the physical environment.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Feb 18, 2022
From: HERCULES CAPITAL, INC., AS COLLATERAL AND ADMINISTRATIVE AGENT
To: HOUZZ INC.
Reel/Frame 059191/0501 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2020
From: HUANG, XIAOYI; WANG, JINGWEN; WU, YI; AI, XIN
To: HOUZZ, INC.
Reel/Frame 053295/0095 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Nov 5, 2019
From: HOUZZ INC.
To: HERCULES CAPITAL, INC., AS COLLATERAL AND ADMINISTRATIVE AGENT
Reel/Frame 050928/0333 →
Continuity (1)
Related Publication 20210027539A1 · Jan 28, 2021
Cited By (1)
US 12,488,535