IP Library › Granted Patent US 12,586,314
Granted Patent B1
US 12,586,314 · App. 18/362,354 · Granted Mar 24, 2026

Systems for generation of 3D models based on images of items

Inventors: Meher Gitika Karumuri (Santa Clara, CA); Sunil Sharadchandra Hadap (Dublin, CA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06T17/10G06F30/10G06T7/11G06T7/74G06T15/04G06T19/00G06T2207/20021G06T2207/20081G06T2207/30196G06T2210/12G06T2210/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,314
App. No.
18/362,354
Granted
Mar 24, 2026
Kind
B1
Abstract

A three-dimensional model for an item is generated using two-dimensional images of the item. Based on the category of the item, bounding boxes in input images that include portions of the item are determined, and a segmentation algorithm is used to generate masks that include specific pixels that represent portions of the item. The generated masks are used in combination with a constructive solid geometry algorithm to create a three-dimensional model of the item. The model is then used to generate a two-dimensional image with pixels that represent portions of the item, each pixel being mapped to a corresponding pixel of an input image. The color and texture values of the input image are associated with the corresponding pixels in the two-dimensional image, which are used to provide the model with colors and textures of the item. This provides a simulated image of the item as worn by a user.

Claims (102)

1 . A system comprising:

one or more non-transitory memories storing computer-executable instructions; and

one or more hardware processors to execute the computer-executable instructions to:

access a plurality of images that depict an item available in a catalog, the plurality of images including at least a first image and a second image, wherein the plurality of images are associated with previously stored category data indicative of a category of the item;

access previously stored bounding box data associated with the category, wherein the bounding box data includes a set of bounding boxes that represent one or more locations of a portion of the item;

based on correspondence between the first image and the bounding box data, determine at least a first bounding box that represents a first location of a first portion of the item within the first image, and a second bounding box that represents a second location of a second portion of the item within the first image;

based on correspondence between the second image and the bounding box data, determine at least a third bounding box that represents a third location of the first portion of the item within the second image, and a fourth bounding box that represents a fourth location of the second portion of the item within the second image;

determine a first mask based on characteristics of a first set of pixels within the first bounding box and a second set of pixels within the second bounding box;

determine a second mask based on characteristics of a third set of pixels within the third bounding box and a fourth set of pixels within the fourth bounding box;

use a constructive solid geometry algorithm to determine a three-dimensional model based on the first mask and the second mask, wherein the three-dimensional model represents the item;

use a mapping algorithm to determine a two-dimensional image based on the three-dimensional model;

determine a first mapping between one or more first pixels of the first image and one or more second pixels of the two-dimensional image;

determine a second mapping between one or more third pixels of the second image and one or more fourth pixels of the two-dimensional image;

associate a first set of values of the one or more first pixels with the one or more second pixels based on the first mapping, wherein the first set of values represents one or more of color or texture of the one or more first pixels;

associate a second set of values of the one or more third pixels with the one or more fourth pixels based on the second mapping, wherein the second set of values represents one or more of color or texture of the one or more third pixels; and

associate one or more values that represent one or more of color or texture with the three-dimensional model, based on the two-dimensional image, the first set of values, and the second set of values.

2 . The system of claim 1 , further comprising computer-executable instructions to:

receive an input image that depicts at least a portion of a user;

determine a fifth set of pixels of the input image that represent a portion of a body of the user;

determine a third mapping between a portion of the three-dimensional model and the fifth set of pixels; and

based on the third mapping, generate an output image that depicts the item in association with the body of the user.

3 . The system of claim 1 , further comprising computer-executable instructions to:

determine that the category of the item is associated with an irregular shape; and

in response to determination of the category associated with the item, use a surface modeling algorithm to determine a surface topology associated with at least a portion of the first mask.

4 . A system comprising:

one or more non-transitory memories storing computer-executable instructions; and

one or more hardware processors to execute the computer-executable instructions to:

access a first image of an item available in a catalog;

determine a first set of pixels of the first image that represent a first portion of the item;

determine a second set of pixels of the first image that represent a second portion of the item;

determine a first mask based on the first set of pixels and a second mask based on the second set of pixels;

determine a first solidified mesh based on the first mask, using a constructive solid geometry algorithm;

determine a second solidified mesh based on the second mask, using the constructive solid geometry algorithm;

determine a three-dimensional model based on a Boolean operation using the first solidified mesh and the second solidified mesh;

based on the three-dimensional model, determine a second image that includes a plurality of pixels, and a first mapping that associates each pixel of the plurality of pixels with a corresponding pixel of the three-dimensional model;

determine a second mapping between one or more first pixels of the first image and one or more second pixels of the second image;

associate a first set of values of the one or more first pixels with the one or more second pixels based on the second mapping, wherein the first set of values represents one or more of color or texture of the one or more first pixels; and

associate the first set of values with one or more third pixels of the three-dimensional model based on the first mapping.

5 . The system of claim 4 , wherein the first image depicts the item having a first orientation, the system further comprising computer-executable instructions to:

access a third image that depicts the item, wherein the third image depicts the item from a second orientation;

determine a third set of pixels of the third image that represent the first portion of the item; and

determine a fourth set of pixels of the third image that represent the second portion of the item;

wherein the first mask is further determined based on the third set of pixels, and the second mask is further determined based on the fourth set of pixels.

6 . The system of claim 4 , further comprising computer-executable instructions to:

access a previously stored category associated with one or more of the first image or the item; and

access previously stored bounding box data associated with the category, wherein the bounding box data includes a set of bounding boxes that represent one or more locations within the first image of the first portion of the item relative to one or more locations within the first image of the second portion of the item;

wherein the first set of pixels and the second set of pixels are determined based on the bounding box data, the first image, and a machine learning algorithm that is trained to determine portions of an image that correspond to locations of bounding boxes based on characteristics of the pixels.

7 . The system of claim 4 , wherein the first image comprises a second plurality of pixels, the system further comprising computer-executable instructions to:

determine one or more pixel characteristics for each pixel of the second plurality of pixels;

wherein the first set of pixels is determined based on a segmentation algorithm and a first subset of the one or more pixel characteristics associated with the first set of pixels, and wherein the second set of pixels is determined based on the segmentation algorithm and a second subset of the one or more pixel characteristics associated with the second set of pixels.

8 . The system of claim 4 , further comprising computer-executable instructions to:

determine a first orientation of the first portion of the item, relative to a normal orientation, based on the first set of pixels using a surface normal algorithm, wherein the first mask has the normal orientation and is generated based on a first difference between the first orientation and the normal orientation; and

determine a second orientation of the second portion of the item, relative to the normal orientation, based on the second set of pixels using the surface normal algorithm, wherein the second mask is generated based on a second difference between the second orientation and the normal orientation.

9 . The system of claim 4 , wherein the Boolean operation comprises a Boolean intersection between the first solidified mesh and the second solidified mesh.

10 . The system of claim 4 , wherein the three-dimensional model is further determined based on one or more of: a signed distance field, an unsigned distance field, a marching cubes algorithm, or an occupancy field.

11 . The system of claim 4 , further comprising computer-executable instructions to:

determine a first orientation of the item associated with the first image;

determine a second orientation associated with one or more of the three-dimensional model or the item associated with the second image; and

determine a difference between the first orientation and the second orientation;

wherein the second mapping between the one or more first pixels of the first image and the one or more second pixels of the second image is determined based on the difference between the first orientation and the second orientation.

12 . The system of claim 4 , further comprising computer-executable instructions to:

determine a plurality of vertices associated with the three-dimensional model; and

determine the second image by associating each vertex of the plurality of vertices with a corresponding UV coordinate in the second image.

13 . The system of claim 4 , further comprising computer-executable instructions to:

receive an input image that depicts a user;

determine a third set of pixels that represent a portion of a body of the user in the input image;

determine a third mapping between a portion of the three-dimensional model and the third set of pixels; and

based on the third mapping, generate an output image that depicts the item in association with the body of the user.

14 . A system comprising:

one or more non-transitory memories storing computer-executable instructions; and

one or more hardware processors to execute the computer-executable instructions to:

access a first image of an item available in a catalog;

determine a first set of pixels of the first image that represent at least a first portion of the item;

determine a first mask based on the first set of pixels;

determine a first solidified mesh based on the first mask, using a constructive solid geometry algorithm;

determine a three-dimensional model based on the first solidified mesh;

determine a second image based on the three-dimensional model, wherein the second image represents at least a first portion and a second portion of the three-dimensional model;

determine a first mapping between one or more first pixels of the first image and one or more second pixels of the second image and a second mapping between one or more third pixels of the second image and one or more fourth pixels of the three-dimensional model; and

associate a first set of values of the one or more first pixels with the one or more second pixels based on the first mapping, wherein the first set of values represents one or more of color or texture of the one or more first pixels.

15 . The system of claim 14 , further comprising computer-executable instructions to:

receive an input image that depicts a user;

determine one or more fifth pixels of the input image that represent a portion of a body of the user;

determine a second mapping between one or more sixth pixels of the three-dimensional model and the one or more fifth pixels of the input image; and

based on the second mapping, generate an output image that depicts the item in association with the body of the user.

16 . The system of claim 14 , further comprising computer-executable instructions to:

determine a second set of pixels of the first image that represent a second portion of the item;

determine a second mask based on the second set of pixels; and

determine a second solidified mesh based on the second mask, using the constructive solid geometry algorithm;

wherein the three-dimensional model is further determined based on a Boolean intersection of the first solidified mesh and the second solidified mesh.

17 . The system of claim 14 , further comprising computer-executable instructions to:

determine a category associated with one or more of the first image or the item; and

determine bounding box data associated with the category, wherein the bounding box data includes at least one bounding box that represents one or more locations within the first image of the first portion of the item;

wherein the first mask is determined based on characteristics of a plurality of pixels within the bounding box, wherein the plurality of pixels includes the first set of pixels.

18 . The system of claim 17 , wherein the first image comprises a plurality of pixels, the system further comprising computer-executable instructions to:

determine one or more pixel characteristics for each pixel of the plurality of pixels of the bounding box;

wherein the first set of pixels is determined based on a segmentation algorithm and a first subset of the one or more pixel characteristics associated with the first set of pixels.

19 . The system of claim 18 , further comprising computer-executable instructions to:

determine a first orientation of the first set of pixels relative to a normal orientation using a surface normal algorithm, wherein the first mask has the normal orientation and is generated based on a first difference between the first orientation and the normal orientation.

20 . The system of claim 14 , further comprising computer-executable instructions to:

determine a category associated with one or more of the first image or the item;

determine that the category is associated with items having irregular geometry; and

in response to determination of the category, use a surface modeling algorithm to determine a surface topology associated with at least a portion of the first mask.

References Cited (43)
US 11830127B1 · Arbit · 2023 [cited by examiner]
US 12094133B2 · Ramanathan · 2024 [cited by examiner]
US 20090177454A1 · Bronstein · 2009 [cited by examiner]
US 20200312008A1 · Cowburn · 2020 [cited by examiner]
US 20210018608A1 · Charpentier · 2021 [cited by examiner]
US 20210326722A1 · Morin · 2021 [cited by examiner]
US 20220136860A1 · Dong · 2022 [cited by examiner]
US 20220269895A1 · Barkan · 2022 [cited by examiner]
CN 115705653A · 2023 [cited by examiner]
JP 2019547805 · 2020 [cited by examiner]
KR 20190028349A · 2019 [cited by examiner]
WO WO2022271838A1 · 2022 [cited by examiner]
“Blender 3D Software”, 18 pgs. Retrieved from the Internet: URL: https://www.blender.org. [cited by applicant]
“ZBrush Pixologic 3D Software”, 6 pgs. Retrieved from the Internet: URL: https://pixologic.com/. [cited by applicant]
“Delaunay triangulation”, 9 pgs. https://en.wikipedia.org/wiki/Delaunay_triangulation. [cited by applicant]
Aberman, et al., “Neural Best-Buddies: Sparse Cross-Domain Correspondence”, ACM Transactions on Graphics, Aug. 21, 2018, 14 pgs. Retrieved from the Internet: URL: https://arxiv.org/pdf/1805.04140.pdf. [cited by applicant]
Alldieck, et al., “Tex2Shape: Detailed Full Human Body Geometry From a Single Image”, Sep. 15, 2019, 13 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/1904.08645. [cited by applicant]
Chan, et al., “Efficient Geometry-aware 3D Generative Adversarial Networks”, Stanford University, Apr. 27, 2022, 27 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2112.07945. [cited by applicant]
Efros, Alexei, “Image Pyramids and Blending”, Computational Photography, CMU, Fall 2005, 53 pgs. Retrieved from the Internet: URL: http://graphics.cs.cmu.edu/courses/15-463/2005_fall/www/Lectures/Pyramids.pdf. [cited by applicant]
Gao, et al., “Get3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images”, Sep. 22, 2022, 39 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2209.11163. [cited by applicant]
Goodfellow, et al., “Generative Adversarial Nets”, University of Montreal, Department of Information, Jun. 10, 2014, 8 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/1406.2661. [cited by applicant]
Jang, et al., “CodeNeRF: Disentangled Neural Radiance Fields for Object Categories”, University College London, Department of Computer Science, Sep. 3, 2021, 10 pgs. Retrieved from the Internet: URL: https://arxiv.org/a… [cited by applicant]
Johari, et al., “GeoNeRF: Generalizing NeRF with Geometry Priors”, Mar. 21, 2022, 19 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2111.13539. [cited by applicant]
Kato, et al., “Neural 3D Mesh Renderer”, The University of Texas at Tokyo, Nov. 20, 2017, 17 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/1711.07566. [cited by applicant]
Li, et al., “Learning a Model of Facial Shape and Expression from 4D Scans”, Nov. 2017, 17 pgs. Retrieved from the Internet: URL: https://ps.is.mpg.de/uploads_file/attachment/attachment/400/paper.pdf. [cited by applicant]
Lin, et al., “SDF-SRN: Learning Signed Distance 3D Object Reconstruction from Static Images”, Carnegie Mellon University, 2020, 12 pgs. Retrieved from the Internet: URL: https://scholar.google.com/scholar?q=sdf-srn:+lea… [cited by applicant]
Long, et al., “NeuralUDF: Learning Unsigned Distance Fields for Multi-view Reconstruction of Surfaces with Arbitrary Topologies”, Nov. 25, 2022, 16 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2211.14173. [cited by applicant]
Loper, et al., “SMPL: A Skinned Multi-Person Linear Model”, Oct. 16, 2015, 16 pgs. Retrieved from the Internet: URL: https://files.is.tue.mpg.de/black/papers/SMPL2015.pdf. [cited by applicant]
Mildenhall, et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, Aug. 3, 2020, 25 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2003.08934. [cited by applicant]
Moon, et al., “3D Clothed Human Reconstruction in the Wild”, Jul. 20, 2022, 25 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2207.10053. [cited by applicant]
Saito, et al., “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitalization”, Dec. 3, 2019, 15 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/1905.05172. [cited by applicant]
Saito, et al., “PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitalization”, Apr. 1, 2020, 10 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2004.00452. [cited by applicant]
Shu, et al., “Portrait Lighting Transfer Using a Mass Transport Approach”, Oct. 2017, 6 pgs. Retrieved from the Internet: URL: https://www3.cs.stonybrook.edu/˜cvl/content/papers/2017/shu_tog2017.pdf. [cited by applicant]
Skorokhodov, et al., “EpiGRAF: Rethinking Training of 3D GANs”, Dec. 15, 2022, 23 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2206.10535. [cited by applicant]
Smigel, Mason, “Deformation Cage” 8 pgs. Retrieved from the Internet: URL: https : / / www.masonsmigel.com/post/deformation-cage. [cited by applicant]
Sorkine, et al., “As-Rigid-As-Possible Surface Modeling”, TU Berlin, Germany, 2007, 8 pgs. Retrieved from the Internet: URL: https://igl.ethz.ch/projects/ARAP/arap_web.pdf. [cited by applicant]
Stutz, David, “A Formal Definition of Watertight Meshes”, Jan. 2018. 11 pgs. Retrieved from the Internet: URL: https://davidstutz.de/a-formal-definition-of-watertight-meshes/. [cited by applicant]
Tewari, et al., “Advances in Neural Rendering”, Mar. 30, 2022, 33 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2111.05849. [cited by applicant]
Wang, et al., “Eyeglasses 3D Shape Reconstruction from a Single Face Image”, 2020, 16 pgs. Retrieved from the Internet: URL: http://xufeng.site/publications/2020/wyt_Eyeglasses%203D%20shape%20reconstruction%20from%20a%2… [cited by applicant]
Wimbauer, et al., “De-rendering 3D Objects in the Wild”, Sep. 27, 2022, 15 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2201.02279. [cited by applicant]
Xu, et al., “3D Human Texture Estimation from a Single Image with Transformers”, Nanyang Technological University, S-Lab, 2021, 10 pgs. Retrieved from the Internet: URL: https://scholar.google.com/scholar?q=3d+human+tex… [cited by applicant]
Yu, et al., “pixelNeRF: Neural Radiance Fields from One or Few Images”, UC Berkeley, May 30, 2021, 20 pgs. Retrieved from the Internet: URL: https://arxiv.org/abs/2012.02190. [cited by applicant]
Zheng, et al., “DeepHuman: 3D Reconstruction From a Single Image”, 2019, 11 pgs. Retrieved from the Internet: URL: https://scholar.google.com/scholar?q=deephuman:+3d+human+reconstruction+from+a+single+image+zheng&hl=en&… [cited by applicant]