IP Library Granted Patent US 12,400,400
Granted Patent B2
US 12,400,400 · App. 17/882,694 · Granted Aug 26, 2025

Techniques for producing three-dimensional models from one or more two-dimensional images

Inventor: Rachelle Villalon (Cambridge, MA)
Assignee: HL Acquisition, Inc.
G06T17/20G06F17/16G06N3/047G06T3/4046G06T7/11G06T7/187G06T7/50G06T7/70G06T15/205G06V10/225G06V10/761G06V10/7715G06V20/70G06T2200/08G06T2207/10028G06T2207/20084G06T2210/12G06T2215/16G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,400
App. No.
17/882,694
Granted
Aug 26, 2025
Kind
B2
Abstract

Described are techniques for producing a three-dimensional model of a scene from one or more two dimensional images. The techniques include receiving by a computing device one or more two dimensional digital images of a scene, the image including plural pixels, applying the received image data to scene generator/scene understanding engine that produces from the one or more digital images a metadata output that includes depth prediction data for at least some of the plural pixels in the two dimensional image and that produces metadata for a controlling a three-dimensional computer model engine, and outputting the metadata to a three-dimensional computer model engine to produce a three-dimensional digital computer model of the scene depicted in the two dimensional image.

Claims (89)

1. A system comprising:

a computing device that includes a processor, memory and storage that stores a computer program that comprises instructions to configure the computing device to:

receive one or more two-dimensional digital images of a scene, each of the one or more two-dimensional digital images including plural pixels;

transform the received plural pixels of each of the one or more, two-dimensional digital images into superpixels that are combinations of the plural pixels that function as single inputs;

produce metadata that includes depth prediction data for at least some of the plural pixels in the one or more two-dimensional digital images;

produce metadata instructions for controlling a three-dimensional computer modeling engine;

apply the metadata that includes a spatial shell and placed objects corresponding to objects in the one or more two-dimensional digital images, with the placed objects being placeholders that are swapped for three-dimensional representations;

receive reference measurements of two-dimensional objects depicted in the one or more two-dimensional digital images;

produce a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffold mesh used to calculate one or more bounding boxes;

calculate original coordinate positions of the one or more bounding boxes;

translate placed object positions and orientations based on the placed objects, the reference measurements, and a scaling;

place the placed objects into the spatial shell; and

output the metadata to the three-dimensional computer modeling engine to produce a three-dimensional digital computer model of the scene including the two-dimensional objects depicted in the one or more two-dimensional digital images.

2. The system of claim 1 wherein the system further comprises instructions to configure the computing device to:

store the reference measurements and reference objects in a database server.

3. The system of claim 1 wherein the metadata contains image processed specific and referenced information.

4. The system of claim 1 wherein the system further comprises instructions to configure the computing device to:

apply depth prediction processing to depth values to extrude a depth image of the depth values into the volumetric mesh that serves as the scaffold mesh for the final mesh.

5. The system of claim 1 wherein the placed objects include properties of the placed objects.

6. The system of claim 1 wherein the system further comprises instructions to configure the computing device to:

swap the placed objects for the three-dimensional representations placed into the spatial shell.

7. The system of claim 1 wherein the system further comprises instructions to configure the computing device to:

semantically segment pixels in the one or more two-dimensional images as belonging to an object category, segmenting a scene into plural segments according to recognized features;

classify pixels as belonging to one of the plural segments, with the segmented image having regions that isolate architectural properties of the scene as captured in an image.

8. The system of claim 1 wherein the metadata instructions include one or more of floor plan, section, elevation drawings, object information, and measurement information.

9. The system of claim 1 wherein the system further comprises instructions to configure the computing device to:

pull a catalog of purchase items from vendors; and

add selected purchase items into the three-dimensional model.

10. A system comprising:

a computing device that includes a processor, memory and storage that stores a computer program that comprises instructions to configure the computing device to:

receive by a transformation engine, one or more two-dimensional digital images of a scene, the one or more two-dimensional digital images including plural pixels and image data;

transform the received plural pixels into superpixels that are combined to function as a single input;

produce from the one or more two-dimensional digital images, metadata that includes depth prediction data for at least some of the plural pixels in the one or more two-dimensional digital images and that produces metadata for controlling a three-dimensional computer modeling engine;

apply the metadata that includes a three-dimensional computer modeling shell and placed objects that are placed in the one or more two-dimensional image, with the placed objects being placeholders that are swapped for three-dimensional representations;

produce a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffold mesh used to calculate one or more bounding boxes;

calculate original coordinate positions of the one or more bounding boxes;

place the placed objects into the three-dimensional computer modeling shell;

receive reference measurements of two-dimensional objects depicted in the one or more two-dimensional digital image;

translate placed object positions, orientations and sizing based on the original coordinate positions, the two-dimensional objects, the reference measurements, and a scaling; and

output the metadata to the three-dimensional computer modeling engine to produce a three-dimensional digital computer model of the scene including the two-dimensional objects depicted in each of the one or more two-dimensional digital images.

11. The system of claim 10 wherein the system further comprises instructions to configure the computing device to:

store the reference measurements and placed objects in a database server.

12. The system of claim 10 wherein the metadata contains image processed specific and referenced information.

13. The system of claim 10 wherein the system further comprises instructions to configure the computing device to:

apply depth prediction processing to depth values to extrude a depth image of the depth values into the volumetric mesh that serves as the scaffold mesh for the final mesh.

14. The system of claim 10 wherein the placed objects include properties of the placed objects.

15. The system of claim 10 wherein the system further comprises instructions to configure the computing device to:

swap the placed objects for the three-dimensional representations placed into the three-dimensional computer modeling shell.

16. The system of claim 10 wherein the system further comprises instructions to configure the computing device to:

semantically segment pixels in the one or more two-dimensional images as belonging to an object category, segmenting a scene into plural segments according to recognized features;

classify pixels as belonging to one of the plural segments, with the segmented image having regions that isolate architectural properties of the scene as captured in an image.

17. The system of claim 10 wherein the metadata includes one or more of a floor plan, a section, elevation drawings, object information, and measurement information.

18. The system of claim 10 wherein the system further comprises instructions to configure the computing device to:

pull a catalog of purchase items from retail vendors; and

add selected purchase items into the three-dimensional model.

19. A method of inserting 3D objects into a spatial 3-dimensional model comprises

producing from metadata a spatial shell to place objects from one or more two-dimensional digital images, with the placed objects being placeholders that are swapped for three-dimensional representations;

producing a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffold mesh used to calculate one or more bounding boxes;

calculating original coordinate positions of the one or more bounding boxes;

placing the placed objects into the spatial shell;

receiving reference measurements of two-dimensional objects depicted in the one or more two-dimensional digital images;

translating the placed object positions, orientations and sizing based on the original coordinate positions, the two-dimensional objects, the reference measurements, and a scaling; and

outputting the metadata to a three-dimensional computer modeling engine to produce a three-dimensional digital computer model of a scene including the two-dimensional objects depicted in the one or more two-dimensional digital images.

20. The method of claim 19 further comprising:

storing the reference measurements and placed objects in a database server.

21. The method of claim 19 wherein the metadata contains image processed specific and referenced information.

22. The method of claim 19 further comprising:

applying depth prediction processing to depth values to extrude a depth image of the depth values into the volumetric mesh that serves as the scaffold mesh for the final mesh.

23. The method of claim 19 wherein the placed objects include properties of the placed objects.

24. The method of claim 19 further comprising:

swapping the placed objects for the three-dimensional representations placed into the spatial shell.

25. The method of claim 19 further comprising:

semantically segmenting pixels in the one or more two-dimensional images as belonging to an object category, segmenting a scene into plural segments according to recognized features;

classifying pixels as belonging to one of the plural segments, with the segmented image having regions that isolate architectural properties of the scene as captured in an image.

26. The method of claim 19 wherein the metadata includes at one or more of a floor plan, a section, elevation drawings, object information, and measurement information.

27. The method of claim 19 further comprising:

pulling a catalog of purchase items from vendors; and

adding selected purchase items into the three-dimensional model.

28. A method of scaling, comprises:

receiving by a transformation engine, one or more two-dimensional digital images of a scene, the one or more two-dimensional digital images including plural pixels and image data;

produce from the one or more two-dimensional digital images, metadata that includes depth prediction data for at least some of the plural pixels in the one or more two-dimensional digital images;

apply the metadata that includes a three-dimensional computer modeling shell and placed objects being placeholders that are swapped for three-dimensional representations;

produce a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffold mesh used to calculate one or more bounding boxes;

calculate original coordinate positions of the one or more bounding boxes;

place the placed objects into the three-dimensional computer modeling shell;

receive reference measurements of two-dimensional objects depicted in the one or more two-dimensional digital images to scale the two-dimensional objects depicted in the one or more two-dimensional digital images; and

scaling placed object positions, orientations and sizing based on the original coordinate positions, the two-dimensional objects, and the reference measurements.

29. The method of claim 28 further comprises:

transforming the received plural pixels into superpixels that are combined to function as single inputs for at least some of the plural pixels.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2025
From: VILLALON, RACHELLE
To: HOSTA LABS INC.
Reel/Frame 071688/0857 →
CHANGE OF NAME Recorded Jul 14, 2025
From: HOSTA LABS INC.
To: HL ACQUISITION, INC. DBA HOSTA AI
Reel/Frame 071931/0778 →
Continuity (3)
Continuation 16423257 · May 28, 2019
Provisional Application 62677219 · May 29, 2018
Related Publication 20220392165A1 · Dec 8, 2022
References Cited (7)
US 20070088531A1 · Yuan · 2007 [cited by examiner]
US 20120299920A1 · Coombe · 2012 [cited by examiner]
US 20180082435A1 · Whelan · 2018 [cited by examiner]
US 20210209855A1 · Stokking · 2021 [cited by examiner]
Eigen, Depth Map Prediction from a Single Image using a Multi-Scale Deep Network, Dept. of Computer Science, Courant Institute, New York University, pp. 9 (Year: 2014). [cited by examiner]
Liu, Deep Convolutional Neural Fields for Depth Estimation from a Single Image, URL: https://www.arxiv-vanity.com/papers/1411.6387/ , pp. 13 (Year: 2014). [cited by examiner]
Sousa, Convolutional Neural Networks for Depth Estimation on 2D Images, 2015, pp. 6 (Year: 2015). [cited by examiner]