IP Library › Granted Patent US 11,170,569
Granted Patent B2
US 11,170,569 · App. 16/823,123 · Granted Nov 9, 2021

System and method for virtual modeling of indoor scenes from imagery

Inventors: Brian Totty (Mountain View, CA); Kevin Wong (Mountain View, CA); Jianfeng Yin (Mountain View, CA); Luis Puig Morales (Mountain View, CA); Paul Gauthier (Mountain View, CA); Salma Jiddi (Mountain View, CA); Qiqin Dai (Mountain View, CA); Brian Pugh (Mountain View, CA); Konstantinos Nektarios Lianos (Mountain View, CA); Angus Dorbie (Mountain View, CA); Yacine Alami (Mountain View, CA); Marc Eder (Mountain View, CA); Christopher Sweeney (Mountain View, CA); Javier Civera (Mountain View, CA)
Assignee: GEOMAGICAL LABS, INC.
G06T17/05G06T5/50G06T7/174G06T7/50G06T2207/10028G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,569
App. No.
16/823,123
Granted
Nov 9, 2021
Kind
B2
Abstract

A method for determining a visual scene virtual representation and a highly accurate visual scene-aligned geometric representation for virtual interaction.

Claims (54)

1. A method for generating a virtual model representative of a physical scene, comprising:

receiving scene data, captured in-situ within the physical scene;

generating a virtual scene visual representation (VSVR) based on the scene data;

determining scene information based on the VSVR and the scene data, wherein the scene information comprises segmentation masks, wall planes, and a floor plane;

generating a plurality of dense depth maps for the physical scene by biasing a set of neural networks, each configured to generate a dense depth map of the plurality, with the scene information as prior knowledge during inference;

generating a final segmentation masks, final wall planes, and a final floor plane based on the plurality of dense depth maps;

generating a virtual model, comprising fusing different scene components from the plurality of dense depth maps into the virtual model; and

transmitting the VSVR, the virtual model, the final segmentation masks, the final wall planes, and the final floor plane to a user.

2. The method of claim 1 , wherein the virtual model is aligned with the VSVR and comprises a depth for each pixel of the VSVR.

3. The method of claim 1 , wherein the scene information further comprises redundant depth maps, different from the dense depth maps, wherein the redundant depth maps comprise depth maps captured as scene data and depth maps generated using different photogrammatic techniques from the scene data.

4. The method of claim 1 , wherein the neural networks comprise alternating-direction neural networks that learn parameters using backpropagation.

5. A method for generating a virtual model representative of a physical scene, comprising:

receiving scene data, captured in-situ within the physical scene;

determining scene information based on the scene data, wherein the scene information comprises segmentation masks, wall planes, and a floor plane;

generating a plurality dense geometric representations of the physical scene by biasing a set of neural networks, each configured to generate a dense geometric representation of the plurality, with the scene information as prior knowledge during inference;

generating a virtual scene visual representation (VSVR) based on the scene data;

determining the virtual model, wherein the virtual model comprises a physical position for each pixel of the VSVR, wherein the physical position is determined from the plurality of dense geometric representations; and

transmitting the VSVR and the virtual model to a user.

6. The method of claim 5 , wherein the virtual model is a VSVR-aligned depth map and scaled to standard units.

7. The method of claim 5 :

wherein the scene data comprises a plurality of source images;

wherein each pixel in the VSVR is associated with a source pixel in a source image from the plurality of source images;

wherein each voxel in each dense geometric representation is associated with a source pixel in a source image from the plurality of source images;

wherein each point in the virtual model is associated with a pixel in the VSVR; and

wherein each point in the virtual model is associated with a position from a voxel of the dense geometric representations, wherein the voxel shares a common source pixel with the respective VSVR pixel.

8. The method of claim 5 , wherein the VSVR is generated before generating the dense geometric representations, wherein the dense geometric representations are generated based on the VSVR.

9. The method of claim 5 , wherein determining the scene information comprises determining the scene information using a set of photogrammetric techniques, the set of photogrammetric techniques comprises at least one of: structure from motion, multi-view stereo, simultaneous localization and mapping, and optical flow.

10. The method of claim 5 , wherein determining the scene information comprises determining redundant variants of the scene information using different techniques.

11. The method of claim 5 , wherein different neural networks of the set are biased with a different one of: the segmentation masks, the wall planes, and the floor plane.

12. The method of claim 5 , further comprising transmitting final wall planes, final floor planes, and final segmentation masks to the user.

13. The method of claim 12 , further comprising generating the final segmentation masks based on the segmentation masks and the dense geometric representations.

14. The method of claim 5 , wherein the set of neural networks comprise alternating-direction neural networks.

15. The method of claim 5 , wherein determining the virtual model comprises fusing different scene components from each of the dense geometric representations into a final dense geometric representation.

16. The method of claim 15 , wherein each of the dense geometric representations is aligned with the VSVR, wherein the final dense geometric representation is the virtual model.

17. The method of claim 5 , further comprising scaling the dense geometric representations based on a common physical point represented in both scaled sensor data and the dense geometric representation.

18. The method of claim 17 , wherein the scaled sensor data comprises a scaled 3D point generated by an augmented reality engine executing on a capture device, wherein the capture device captures the scene data.

19. The method of claim 5 , wherein the dense geometric representations comprise at least one of: a floor enhanced geometric representation, a wall enhanced geometric representation, and an occlusion enhanced geometric representation.

20. The method of claim 5 , wherein the VSVR comprises a photorealistic panoramic image.

21. An apparatus for generating a virtual model representative of a physical scene, the apparatus comprising:

one or more processors; and

one or more memories operatively coupled to at least one of the one or more processors and having instructions stored thereon that, when executed by at least one of the one or more processors, cause at least one of the one or more processors to:

receive scene data, captured in-situ within the physical scene;

determine scene information based on the scene data, wherein the scene information comprises segmentation masks, wall planes, and a floor plane;

generate a plurality dense geometric representations of the physical scene by biasing a set of neural networks, each configured to generate a dense geometric representation of the plurality, with the scene information as prior knowledge during inference;

generate a virtual scene visual representation (VSVR) based on the scene data;

determine the virtual model, wherein the virtual model comprises a physical position for each pixel of the VSVR, wherein the physical position is determined from the plurality of dense geometric representations; and

transmit the VSVR and the virtual model to a user.

22. At least one non-transitory computer-readable medium storing computer-readable instructions for generating a virtual model representative of a physical scene that, when executed by one or more computing devices, cause at least one of the one or more computing devices to:

receive scene data, captured in-situ within the physical scene;

determine scene information based on the scene data, wherein the scene information comprises segmentation masks, wall planes, and a floor plane;

generate a plurality dense geometric representations of the physical scene by biasing a set of neural networks, each configured to generate a dense geometric representation of the plurality, with the scene information as prior knowledge during inference;

generate a virtual scene visual representation (VSVR) based on the scene data;

determine the virtual model, wherein the virtual model comprises a physical position for each pixel of the VSVR, wherein the physical position is determined from the plurality of dense geometric representations; and

transmit the VSVR and the virtual model to a user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2020
From: TOTTY, BRIAN; WONG, KEVIN; YIN, JIANFENG; MORALES, LUIS PUIG; GAUTHIER, PAUL; JIDDI, SALMA; DAI, QIQIN; PUGH, BRIAN; LIANOS, KONSTANTINOS NEKTARIOS; DORBIE, ANGUS; ALAMI, YACINE; EDER, MARC; SWEENEY, CHRISTOPHER; CIVERA, JAVIER
To: GEOMAGICAL LABS, INC.
Reel/Frame 052422/0118 →
Continuity (2)
Provisional Application 62819817 · Mar 18, 2019
Related Publication 20200302686A1 · Sep 24, 2020