IP Library Granted Patent US 11,423,615
Granted Patent B1
US 11,423,615 · App. 16/423,257 · Granted Aug 23, 2022

Techniques for producing three-dimensional models from one or more two-dimensional images

Inventor: Rachelle Villalon (Cambridge, MA)
Assignee: HL Acquisition, Inc.
G06T17/20G06F17/16G06N3/0472G06T15/205G06T2207/10028G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,615
App. No.
16/423,257
Granted
Aug 23, 2022
Kind
B1
Abstract

Described are techniques for producing a three-dimensional model of a scene from one or more two dimensional images. The techniques include receiving by a computing device one or more two dimensional digital images of a scene, the image including plural pixels, applying the received image data to scene generator/scene understanding engine that produces from the one or more digital images a metadata output that includes depth prediction data for at least some of the plural pixels in the two dimensional image and that produces metadata for a controlling a three-dimensional computer model engine, and outputting the metadata to a three-dimensional computer model engine to produce a three-dimensional digital computer model of the scene depicted in the two dimensional image.

Claims (72)

1. A method comprises:

receiving by a computing device, a single, two-dimensional digital image of a scene, the single, two-dimensional digital image including plural pixels and image data;

applying the received image data to a transformation engine that transforms the single, two-dimensional digital image by superpixel segmentation that combines small homogenous regions of pixels into superpixels that are combined to function as a single input and produces from the transformed single two-dimensional digital image, a metadata output that includes depth prediction data for at least some of the plural pixels in the single, two-dimensional digital image and that produces metadata for controlling a three-dimensional computer model engine that includes a spatial shell to provide placed objects in the single two-dimensional image, with the placed objects being placeholders that are swapped for three-dimensional representations of the placed objects within the spatial shell;

receiving reference measurements of two dimensional objects depicted in the single two-dimensional digital image;

producing a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffolding mesh used to calculate one or more bounding boxes;

calculating original coordinate positions of the one or more bounding boxes;

transforming object positions and orientations based on the original coordinate positions; and

placing the placed objects into the spatial shell; and

outputting the metadata to the three-dimensional computer modeling engine to produce a three-dimensional digital computer model of the scene including the two dimensional objects depicted in the single, two dimensional digital image.

2. The method of claim 1 wherein applying the received single two-dimensional digital image to the transformation engine, further comprises:

identifying the two dimensional objects within the image scene in the single, two-dimensional digital image; and

applying labels to the identified two dimensional objects.

3. The method of claim 2 wherein identifying the two dimensional objects further comprises:

extracting each labeled object's region; and

determining and outputting pixel corner coordinates, height and width, and confidence values into the metadata output to provide specific instructions for the three-dimensional modeling engine to produce the three-dimensional model.

4. The method of claim 2 further comprises:

generating using the metadata, statistical information on the identified two dimensional objects within the image scene.

5. The method of claim 1 wherein the depth prediction data produced by the transformation engine is the result of the transformation engine inferring the depths of the at least some of the pixels in the single, two-dimensional digital image.

6. The method of claim 5 wherein inferring depths of pixels in the image further comprises:

determining a penalty function for superpixels by:

determining unary values over each of the superpixels;

determining pairwise values over each of the superpixels; and

determining a combination of the unary and the pairwise values.

7. The method of claim 6 wherein the unary processing returns a depth value for a single superpixel and the pairwise communicates with neighboring superpixels having similar appearance to produce similar depths for those neighboring superpixels.

8. The method of claim 6 wherein the unary processing for a single superpixel is determined by:

inputting the single superpixel into a fully convolutional neural net that produces a convolutional map that has been up-sampled to the original image size;

applying the up-sampled convolutional map and the superpixel segmentation over the original input image to a superpixel average pooling layer to produce feature vectors; and

input the feature vectors to a fully connected output layer to produce a unary output for the superpixel.

9. The method of claim 6 wherein the pairwise processing for a single superpixel is determined by:

collecting similar feature vectors are collected from all neighboring superpixel patches adjacent to the single superpixel;

cataloguing unique feature vectors of the superpixel and neighboring superpixel patches into collections of similar and unique features; and

input the collections into a fully connected layer that outputs a vector of similarities between the neighboring superpixel patches and the single superpixel.

10. The method of claim 9 wherein the unary output is fed into a conditional random fields graph model to produce an output depth map that contains the depth prediction data relating to the distance of surfaces of the two dimensional objects from a reference point.

11. The method of claim 2 wherein a depth prediction service processes input digital pixel data through a pre-trained convolutional neural network to produce an output depth map.

12. A system comprising:

a computing device that includes a processor, memory and storage that stores a computer program that comprises instructions to configure the computing device to:

receive by a transformation engine a single, two-dimensional digital image of a scene, the single, two-dimensional digital image including plural pixels and image data;

apply the received image data to transform the single, two-dimensional digital image by superpixel segmentation that combines small homogenous regions of pixels into superpixels that are combined to function as a single input;

produce from the transformed single two-dimensional digital image, a metadata output that includes depth prediction data for at least some of the plural pixels in the single, two-dimensional digital image and that produces metadata for controlling a three-dimensional computer model engine;

produce the metadata that includes a spatial shell to place objects in the single two-dimensional image, with the placed objects being placeholders that are swapped for three-dimensional representations;

produce a volumetric mesh that serves as a scaffold mesh for a final mesh, with the scaffolding mesh used to calculate one or more bounding boxes;

calculate the original coordinate position of the one or more bounding boxes;

transform object positions and orientations based on the original coordinates; and

place the placed objects into the three dimensional computer model shell;

receive reference measurements of two dimensional objects depicted in the single two-dimensional digital image; and

output the metadata to the three-dimensional computer modeling engine to produce a three-dimensional digital computer model of the scene including the two dimensional objects depicted in the single, two dimensional digital image.

13. The system of claim 12 wherein the computing system is further configured by the instructions to:

apply the received single two-dimensional digital image to the transformation engine, further comprises:

identify the two dimensional objects within the image scene in the single, two-dimensional digital image; and

apply labels to the identified objects.

14. The system of claim 12 wherein the system further comprises instructions to:

generate, using the metadata, statistical information on the identified objects within the image scene.

15. The system of claim 13 wherein identify the two dimensional objects further comprise instructions to:

extract each labeled object's region; and

determine and outputting pixel corner coordinates, height and width, and confidence values into the metadata output to provide specific instructions for the three-dimensional modeling engine to produce the three-dimensional model.

16. The system of claim 15 wherein the depth prediction data produced by the transformation engine is the result of the transformation engine inferring the depths of the at least some of the pixels in the single, two-dimensional digital image.

17. The system of claim 15 wherein the transformation engine infers depths of pixels in the image by instructions to:

determine a penalty function for superpixels by instructions to:

determine unary values over each of the superpixels;

determine pairwise values over each of the superpixels; and

determine a combination of the unary and the pairwise values.

18. The system of claim 17 wherein the instructions to determine unary values return a depth value for a single superpixel and the pairwise instructions communicate with neighboring superpixels having similar appearance to produce similar depths for those neighboring superpixels.

19. The system of claim 17 wherein the instructions to determine unary values for a single superpixel comprise instructions to:

input the single superpixel into a fully convolutional neural net that produces a convolutional map that has been up-sampled to the original image size;

apply the up-sampled convolutional map and the superpixel segmentation over the original input image to a superpixel average pooling layer to produce feature vectors; and

input the feature vectors to a fully connected output layer to produce a unary output for the superpixel.

20. The system of claim 17 wherein the instructions to determine pairwise values for a single superpixel comprise instructions to:

collect similar feature vectors are collected from all neighboring superpixel patches adjacent to the single superpixel;

catalogue unique feature vectors of the superpixel and neighboring superpixel patches into collections of similar and unique features; and

input the collections into a fully connected layer that outputs a vector of similarities between the neighboring superpixel patches and the single superpixel.

21. The system of claim 18 wherein the determined unary output is fed into a conditional random fields graph model to produce an output depth map that contains the depth prediction data relating to the distance of surfaces of the two dimensional objects from a reference point.

22. The system of claim 13 wherein a depth prediction service processes input digital pixel data through a pre-trained convolutional neural network to produce an output depth map.

Assignments (2)
CHANGE OF NAME Recorded Apr 19, 2022
From: HOSTA LABS INC.
To: HL ACQUISITION, INC. DBA HOSTA AI
Reel/Frame 059726/0030 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2019
From: VILLALON, RACHELLE
To: HOSTA LABS INC.
Reel/Frame 049289/0578 →
Continuity (1)
Provisional Application 62677219 · May 29, 2018
Cited By (3)
US 12,260,575 US 12,387,370 US 12,555,395