IP Library › Granted Patent US 11,797,724
Granted Patent B2
US 11,797,724 · App. 17/454,020 · Granted Oct 24, 2023

Scene layout estimation

Inventors: Shreyas Hampali (Styria, AT); Sinisa Stekovic (Graz, AT); Friedrich Fraundorfer (Graz, AT); Vincent Lepetit (Talence, FR)
Assignee: QUALCOMM Incorporated
G06F30/13G06T7/55G06T17/20G06T2207/10028G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,797,724
App. No.
17/454,020
Granted
Oct 24, 2023
Kind
B2
Abstract

Systems and techniques are provided for determining environmental layouts. For example, based on one or more images of an environment and depth information associated with the one or more images, a set of candidate layouts and a set of candidate objects corresponding to the environment can be detected. The set of candidate layouts and set of candidate objects can be organized as a structured tree. For instance, a structured tree can be generated including nodes corresponding to the set of candidate layouts and the set of candidate objects. A combination of objects and layouts can be selected in the structured tree (e.g., based on a search of the structured tree, such as using a Monte-Carlo Tree Search (MCTS) algorithm or adapted MCTS algorithm). A three-dimensional (3D) layout of the environment can be determined based on the combination of objects and layouts in the structured tree.

Claims (58)

1. An apparatus for determining one or more environmental layouts, comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

detect, based on one or more images of an environment and depth information associated with the one or more images, a set of candidate layouts and a set of candidate objects corresponding to the environment;

generate a structured tree comprising nodes corresponding to the set of candidate layouts and the set of candidate objects;

perform a search of the structured tree using a tree search algorithm to determine a respective score for each node in the structured tree, each score comprising a weight assigned for one or more views associated with each node of the structured tree, an exploration term derived for the tree search algorithm, and a fitness value;

select, based on the respective score determined for each node in the structured tree, a combination of objects and layouts in the structured tree; and

determine a three-dimensional (3D) layout of the environment based on the combination of objects and layouts in the structured tree.

2. The apparatus of claim 1 , wherein the tree search algorithm includes a Monte-Carlo Tree Search (MCTS) algorithm.

3. The apparatus of claim 2 , wherein the MCTS algorithm comprises an adapted MCTS algorithm, wherein the adapted MCTS algorithm assigns a fitness value to each node searched in the structured tree, the fitness value representing a probability that at least one of an object or layout associated with the node is present in the environment.

4. The apparatus of claim 1 , wherein the weight is at least partly based on a respective view score for each view of the one or more views, and wherein a view score for a view defines a consistency measurement between a candidate associated with a node associated with the view and data from the one or more images and the depth information associated with the view.

5. The apparatus of claim 1 , wherein the set of candidate layouts comprises a set of 3D layout models and associated poses, and wherein the set of candidate objects comprises a set of 3D object models and associated poses.

6. The apparatus of claim 1 , wherein the one or more images comprise one or more red-green-blue images (RGB) and the depth information comprises one or more depth maps of the environment.

7. The apparatus of claim 1 , wherein the one or more images and the depth information comprise one or more RGB-Depth (RGB-D) images.

8. The apparatus of claim 1 , wherein, to detect the set of candidate layouts, the at least one processor is configured to:

identify, based on a semantic segmentation of a point cloud associated with the one or more images, 3D points corresponding to at least one of a wall of the environment and a floor of the environment;

generate 3D planes based on the 3D points;

generate polygons based on intersections between at least some of the 3D planes; and

determine layout candidates based on the polygons.

9. The apparatus of claim 1 , wherein, to detect the set of candidate objects, the at least one processor is configured to:

detect 3D bounding box proposals for objects in a point cloud generated for the environment; and

for each bounding box proposal, retrieve a set of candidate object models from a dataset.

10. The apparatus of claim 1 , wherein the structured tree comprises multiple levels, wherein each level of multiple levels comprises a different set of incompatible candidates associated with the environment, and wherein the different set of incompatible candidates comprise at least one of incompatible objects and incompatible layouts.

11. The apparatus of claim 10 , wherein two or more candidates from the different set of incompatible candidates are incompatible when the two or more candidates intersect or are not spatial neighbors.

12. The apparatus of claim 1 , wherein the environment comprises a 3D scene.

13. The apparatus of claim 1 , wherein each candidate of the set of candidate layouts and the set of candidate objects comprises a polygon corresponding to intersecting planes.

14. The apparatus of claim 13 , wherein the intersecting planes include one or more two-dimensional planes.

15. The apparatus of claim 13 , wherein the polygon includes a three-dimensional polygon.

16. The apparatus of claim 1 , wherein the at least one processor is configured to:

generate virtual content based on the determined 3D layout of the environment.

17. The apparatus of claim 1 , wherein the at least one processor is configured to:

share the determined 3D layout of the environment with a computing device.

18. The apparatus of claim 1 , wherein the apparatus is a mobile device including a camera for capturing the one or more images.

19. A method for determining one or more environmental layouts, comprising:

detecting, based on one or more images of an environment and depth information associated with the one or more images, a set of candidate layouts and a set of candidate objects corresponding to the environment;

generating a structured tree comprising nodes corresponding to the set of candidate layouts and the set of candidate objects;

perform a search of the structured tree using a tree search algorithm to determine a respective score for each node in the structured tree, each score comprising a weight assigned for one or more views associated with each node of the structured tree, an exploration term derived for the tree search algorithm, and a fitness value;

selecting, based on the respective score determined for each node in the structured tree, a combination of objects and layouts in the structured tree; and

determining a three-dimensional (3D) layout of the environment based on the combination of objects and layouts in the structured tree.

20. The method of claim 19 , wherein the tree search algorithm includes a Monte-Carlo Tree Search (MCTS) algorithm.

21. The method of claim 20 , wherein the MCTS algorithm comprises an adapted MCTS algorithm, wherein the adapted MCTS algorithm assigns a fitness value to each node searched in the structured tree, the fitness value representing a probability that at least one of an object or layout associated with the node is present in the environment.

22. The method of claim 19 , wherein the set of candidate layouts comprises a set of 3D layout models and associated poses, and wherein the set of candidate objects comprises a set of 3D object models and associated poses.

23. The method of claim 19 , wherein the one or more images comprise one or more red-green-blue images (RGB) and the depth information comprises one or more depth maps of the environment.

24. The method of claim 19 , wherein the one or more images and the depth information comprise one or more RGB-Depth (RGB-D) images.

25. The method of claim 19 , wherein detecting the set of candidate layouts comprises:

identifying, based on a semantic segmentation of a point cloud associated with the one or more images, 3D points corresponding to at least one of a wall of the environment and a floor of the environment;

generating 3D planes based on the 3D points;

generating polygons based on intersections between at least some of the 3D planes; and

determining layout candidates based on the polygons.

26. The method of claim 19 , wherein detecting the set of candidate objects comprises:

detecting 3D bounding box proposals for objects in a point cloud generated for the environment; and

for each bounding box proposal, retrieve a set of candidate object models from a dataset.

27. The method of claim 19 , wherein the structured tree comprises multiple levels, wherein each level of the multiple levels comprises a different set of incompatible candidates associated with the environment, and wherein the different set of incompatible candidates comprise at least one of incompatible objects and incompatible layouts.

28. The method of claim 19 , further comprising:

generating virtual content based on the determined 3D layout of the environment.

29. The method of claim 19 , further comprising:

sharing the determined 3D layout of the environment with a computing device.

30. The method of 19 , wherein the weight is at least partly based on a respective view score for each view of the one or more views, and wherein a view score for a view defines a consistency measurement between a candidate associated with a node associated with the view and data from the one or more images and the depth information associated with the view.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2022
From: HAMPALI, SHREYAS; STEKOVIC, SINISA; FRAUNDORFER, FRIEDRICH; LEPETIT, VINCENT
To: TECHNISCHE UNIVERSITAT GRAZ
Reel/Frame 059589/0790 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2022
From: TECHNISCHE UNIVERSITÄT GRAZ
To: QUALCOMM TECHNOLOGIES, INC.
Reel/Frame 059589/0912 →
Continuity (2)
Provisional Application 63113722 · Nov 13, 2020
Related Publication 20220156426A1 · May 19, 2022