IP Library › Granted Patent US 12,738,078
Granted Patent B2
US 12,738,078 · App. 18/414,163 · Granted Sep 15, 2026

Systems and methods for efficient floorplan generation from 3D scans of indoor scenes

Inventor: Ameya Pramod Phalak (Sunnyvale, CA)
Assignee: Magic Leap, Inc.
G06V20/64G06F18/23G06F18/2431G06T7/55G06T17/00G06V10/7625G06V20/20G06T2207/10024G06T2207/10028G06T2207/20081G06T2210/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,078
App. No.
18/414,163
Granted
Sep 15, 2026
Kind
B2
Abstract

Methods, systems, and wearable extended reality devices for generating a floorplan of an indoor scene are provided. A room classification of a room and a wall classification of a wall for the room may be determined from an input image of the indoor scene. A floorplan may be determined based at least in part upon the room classification and the wall classification without constraining a total number of rooms in the indoor scene or a size of the room.

Claims (66)

1 . A method, comprising:

determining a plurality of entities in a higher-dimensional space from a plurality of images captured from the higher-dimensional space, wherein the higher-dimensional space has a dimensionality greater than two; and

determining an indoor scene having one or more rooms or areas and one or more walls in the higher-dimensional space at least by:

determining one or more first labels for a first subset of entities of the plurality of entities and one or more wall labels for the one or more walls using at least the first subset;

determining, by a neural network, one or more room labels for the one or more rooms or areas based at least in part upon the one or more first labels;

generating one or more respective shapes for the one or more rooms or areas at least by jointly utilizing the one or more first labels, the one or more wall labels, or the one or more room labels;

determining the indoor scene from the plurality of entities based at least in part upon the one or more respective shapes; and

generating a dataset for training the neural network at least by:

identifying a set of shapes, wherein a shape comprises a combination of bits on a binary kernel;

selecting a shape from the set based at least in part upon a constraint;

placing the shape in an occupancy grid in a grid space; and

selecting a separate shape and placing the separate shape in an adjacent occupancy grid in the grid space.

2 . The method of claim 1 , wherein at least some images of the plurality of images are captured via a higher-dimensional scan of the higher-dimensional space that comprises the indoor scene.

3 . The method of claim 1 , wherein the indoor scene is determined at least by determining the one or more areas or rooms from the first subset of entities, and the one or more first labels and the one or more wall labels are concurrently determined in parallel using the neural network.

4 . The method of claim 1 , wherein a first label corresponds to an entity of the first subset of entities, and the entity corresponds to a vertex in the first subset of entities that is perceptually uniform.

5 . The method of claim 1 , wherein the neural network is used to concurrently output the one or more room labels and the one or more wall labels as an output using the first subset of entities as an input.

6 . The method of claim 1 , determining the one or more room labels comprising:

determining a second subset of entities from the plurality of entities; and

obtaining, from a second subset entity in the second subset, two or more first data structures and one or more second data structures, wherein each second subset entity is used to cast at least two votes for the one or more rooms and at least one vote for the one or more walls.

7 . The method of claim 6 , determining the one or more room labels comprising:

determining whether the second subset entity belongs to a single room or area or more than one room or area based at least in part upon the two or more first data structures.

8 . The method of claim 7 , wherein

the two or more first data structures are identical when the second subset entity belongs to the single room of the one or more rooms and are distinct when the second subset entity belongs to two or more different rooms; and

a first data structure of the two or more first data structures is determined from the second subset entity to a predicted center of a room or area of the one or more rooms, and a second vector of the one or more second data structures is determined from the second subset entity to a center of a wall of the one or more walls.

9 . The method of claim 7 , determining the one or more room labels comprising:

obtaining a plurality of first data structures and a plurality of second data structures generated by using at least the second subset of entities; and

clustering the plurality of first data structures or the plurality of second data structures respectively based at least in part upon respective spatial density distributions pertaining to the plurality of first data structures or the plurality of second data structures, wherein a vote cast by the second subset entity represents that the second subset entity is determined to be located within a room or area or along a wall.

10 . The method of claim 1 , further comprising:

training the neural network at least by providing the dataset to the neural network as an input and determining whether a predicted output generated by the neural network satisfies a criterion.

11 . The method of claim 1 , further comprising:

generating a super-room at least by assigning a single room identification to multiple shapes selected from the set and placed in the grid space.

12 . A system, comprising:

a processor; and

memory operatively coupled to the processor and storing a sequence of instructions which, when executed by the processor, causes the processor to perform a set of acts, the set of acts comprising:

determining a plurality of entities in a higher-dimensional space from a plurality of images captured from the higher-dimensional space, wherein the higher-dimensional space has a dimensionality greater than two; and

determining an indoor scene having one or more rooms or areas and one or more walls in the higher-dimensional space at least by:

determining one or more first labels for a first subset of entities of the plurality of entities and one or more wall labels for the one or more walls using at least the first subset;

determining, by a neural network, one or more room labels for the one or more rooms or areas based at least in part upon the one or more first labels;

generating one or more respective shapes for the one or more rooms or areas at least by jointly utilizing the one or more first labels, the one or more wall labels, or the one or more room labels;

determining the indoor scene from the plurality of entities based at least in part upon the one or more respective shapes;

determining a second subset of entities from the plurality of entities;

obtaining, from a second subset entity in the second subset, two or more first data structures and one or more second data structures, wherein each second subset entity is used to cast at least two votes for the one or more rooms and at least one vote for the one or more walls; and

determining whether the second subset entity belongs to a single room or area or more than one room or area based at least in part upon the two or more first data structures,

wherein

the two or more first data structures are identical when the second subset entity belongs to the single room of the one or more rooms and are distinct when the second subset entity belongs to two or more different rooms, and

a first data structure of the two or more first data structures is determined from the second subset entity to a predicted center of a room or area of the one or more rooms, and a second data structure of the one or more second data structures is determined from the second subset entity to a center of a wall of the one or more walls.

13 . The system of claim 12 , wherein at least some images of the plurality of images are captured via a higher-dimensional scan of the higher-dimensional space that comprises the indoor scene.

14 . The system of claim 12 , wherein the indoor scene is determined at least by determining one or more areas or rooms from the first subset of entities, a first label corresponds to an entity of the first subset of entities, the entity corresponds to a vertex in the first subset of entities that is perceptually uniform, and the neural network is used to concurrently output the one or more room labels and the one or more wall labels as an output using the first subset of entities as an input.

15 . A wearable extended reality device for generating a floorplan of an indoor scene, comprising:

an optical system having an array of micro-displays or micro-projectors to present digital contents to an eye of a user;

a processor coupled to the optical system; and

memory operatively coupled to the processor and storing a sequence of instructions which, when executed by the processor, causes the processor to perform a set of acts, the set of acts comprising:

determining a plurality of entities in a higher-dimensional space from a plurality of images captured from the higher-dimensional space, wherein the higher-dimensional space has a dimensionality greater than two; and

determining an indoor scene having one or more rooms or areas and one or more walls in the higher-dimensional space at least by:

determining one or more first labels for a first subset of entities of the plurality of entities and one or more wall labels for the one or more walls using at least the first subset;

determining, by a neural network, one or more room labels for the one or more rooms or areas based at least in part upon the one or more first labels;

generating one or more respective shapes for the one or more rooms or areas at least by jointly utilizing the one or more first labels, the one or more wall labels, or the one or more room labels;

determining the indoor scene from the plurality of entities based at least in part upon the one or more respective shapes;

determining a second subset of entities from the plurality of entities;

obtaining, from a second subset entity in the second subset, two or more first data structures and one or more second data structures, wherein each second subset entity is used to cast at least two votes for the one or more rooms and at least one vote for the one or more walls; and

determining whether the second subset entity belongs to a single room or area or more than one room or area based at least in part upon the two or more first data structures,

wherein

the two or more first data structures are identical when the second subset entity belongs to the single room of the one or more rooms and are distinct when the second subset entity belongs to two or more different rooms, and

a first data structure of the two or more first data structures is determined from the second subset entity to a predicted center of a room or area of the one or more rooms, and a second data structure of the one or more second data structures is determined from the second subset entity to a center of a wall of the one or more walls.

16 . The wearable extended reality device of claim 15 , wherein at least some images of the plurality of images are captured via a higher-dimensional scan of the higher-dimensional space that comprises the indoor scene.

17 . The wearable extended reality device of claim 15 , wherein the indoor scene is determined at least by determining one or more areas or rooms from the first subset of entities, a first label corresponds to an entity of the first subset of entities, the entity corresponds to a vertex in the first subset of entities that is perceptually uniform, and the neural network is used to concurrently output the one or more room labels and the one or more wall labels as an output using the first subset of entities as an input.

Assignments (2)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073388/0027 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2025
From: PHALAK, AMEYA PRAMOD
To: MAGIC LEAP, INC.
Reel/Frame 072113/0709 →
Continuity (3)
Continuation 17190889 · Mar 3, 2021
Provisional Application 62985263 · Mar 4, 2020
Related Publication 20240203138A1 · Jun 20, 2024
References Cited (83)
US 20160055268A1 · Bell et al. · 2016 [cited by applicant]
US 20180315162A1 · Sturm et al. · 2018 [cited by applicant]
US 20180330184A1 · Mehr et al. · 2018 [cited by applicant]
US 20190026957A1 · Gausebeck · 2019 [cited by applicant]
US 20190035099A1 · Ebrahimi Afrouzi et al. · 2019 [cited by applicant]
US 20190147220A1 · Mccormac · 2019 [cited by examiner]
US 20190197777A1 · Steinbrucker et al. · 2019 [cited by applicant]
US 20210225090A1 · Tang et al. · 2021 [cited by applicant]
US 20220319106A1 · Totty · 2022 [cited by examiner]
CN 110633640 · 2019 [cited by applicant]
Foreign NOA for JP Patent Appln. No. 2022-552796 dated Feb. 21, 2025. [cited by applicant]
Foreign OA for JP Patent Appln. No. 2022-552796 dated Nov. 1, 2024. [cited by applicant]
Foreign OA for CN Patent Appln. No. 202180032804.8 dated Jul. 22, 2025. [cited by applicant]
English Translation of Foreign OA for CN Patent Appln. No. 202180032804.8 dated Jul. 22, 2025. [cited by applicant]
Foreign OA Response for JP Patent Appln. No. 2022-552796 dated Jan. 27, 2025. [cited by applicant]
Foreign OA for JP Patent Appln. No. 2022-552796 dated Jun. 27, 2024 (with English translation). [cited by applicant]
Foreign Response for EP Patent Appln. No. 21763943.4 dated Feb. 20, 2024. [cited by applicant]
Foreign OA Response for JP Patent Appln. No. 2022-552796 dated Sep. 20, 2024. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/190,889 dated Nov. 3, 2022. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/190,889 dated Mar. 13, 2023. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/190,889 dated Jun. 27, 2022. [cited by applicant]
Amendment Response to FOA for U.S. Appl. No. 17/190,889 dated Jan. 3, 2023. [cited by applicant]
Amendment Response to FOA for U.S. Appl. No. 17/190,889 dated Jun. 13, 2023. [cited by applicant]
Amendment Response to FOA for U.S. Appl. No. 17/190,889 dated Sep. 27, 2022. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/190,889 dated Oct. 17, 2023. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/190,889 dated Aug. 9, 2023. [cited by applicant]
Richter et al., “Constructing Hierarchical Representations of Indoor Spaces”, 2009 Tenth International Conference on Mobile Data Management: Systems, Services and Middleware, IEEE Computer Society, p. 686-691. (Year: 20… [cited by applicant]
Arayici et al., “Building Information Modelling (BIM) for Facilities Management (FM): The Mediacity Case Study Approach”, International Journal of3-D Information Modeling, 1(1) 55-73, Jan.-Mar. 2012. (Year: 2012). [cited by applicant]
PCT International Search Report and Written Opinion for International Appln. No. PCT/US21/20668, Applicant Magic Leap, Inc., dated May 24, 2021 (15 pages). [cited by applicant]
Ankerst, M., et al., “OPTICS: Ordering Points To Identify the Clustering Structure,” Proc. ACM SIGMOD'99 Int. Conf. on Management of Data, Philadelphia PA, 1999. [cited by applicant]
Avetisya, A., et al., “Scan2CAD: Learning CAD Model Alignment in RGB-D Scans,” dated Nov. 27, 2018. [cited by applicant]
Ayad, H., et al., “Cumulative Voting Consensus Method for Partitions with a Variable Number of Clusters,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, No. 1, Jan. 2008. [cited by applicant]
Caron, M., et al., “Deep Clustering for Unsupervised Learning of Visual Features,” Facebook AI Research, dated Mar. 18, 2019. [cited by applicant]
Choy, C., et al., “4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks,” dated Jun. 13, 2019. [cited by applicant]
G. A. Croes, “A Method for Solving Traveling-Salesman Problems,” Operations Research 6(6):791-812. https://doi.org/10.1287/opre.6.6.791 (1958). [cited by applicant]
Dai, A., et al., “ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans,” dated Mar. 28, 2018. [cited by applicant]
Dasgupta, S., et al., “DeLay: Robust Spatial Layout Estimation for Cluttered Indoor Scenes,” Stanford University, dated 2016. [cited by applicant]
Derpanis, K., “Mean Shift Clustering,” dated Aug. 15, 2005. [cited by applicant]
Weingessel, A., et al., “An Ensemble Method for Clustering,” DSC 2003 Working Papers, dated 2003. [cited by applicant]
Ester, M., et al., “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise,” Institute for Computer Science, University of Munich, copyright 1996. [cited by applicant]
Fischler, M., et al., “Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography,” Communications of the ACM, Jun. 1981, vol. 24, No. 6. [cited by applicant]
Gkioxari, G., et al., “Mesh R-CNN,” dated Jan. 25, 2020. [cited by applicant]
Graham, B., et al., “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks,” dated Nov. 29, 2017. [cited by applicant]
He, K., et al., “Mask R-CNN,” Facebook AI Research (FAIR), dated Jan. 24, 2018. [cited by applicant]
Hsiao, C., et al., “Flat2Layout: Flat Representation for Estimating Layout of General Room Types,” dated May 29, 2019. [cited by applicant]
Hirschmuller, H., “Stereo Processing by Semi-Global Matching and Mutual Information,” IEEE Transactions on Pattern Analysis and Machine Intelligence, copyright 2007. [cited by applicant]
“3D-SIC: 3D Semantic Instance Completion for RGB-D Scans,” ICLR, 2020. [cited by applicant]
Chen, J. et al., “Floor-SP: Inverse CAD for Floorplans by Sequential Room-wise Shortest Path,” dated Aug. 19, 2019. [cited by applicant]
Hou, Y., Multi-View Stereo by Temporal Nonparametric Fusion, Department of Computer Science, Aalto University, Finland, Aug. 16, 2019. [cited by applicant]
Kruzhilov, I., et al., “Double Refinement Network for Room Layout Estimation,” Preprints, dated May 22, 2019. [cited by applicant]
Lee, C., et al. “RoomNet: End-to-End Room Layout Estimation,” Magic Leap, Inc., Aug. 7, 2017. [cited by applicant]
Huang, P., et al., “DeepMVS: Learning Multi-view Stereopsis,” dated Apr. 2, 2018. [cited by applicant]
Li, Y., et al., “PointCNN: Convolution On X-Transformed Points,” Preprint, dated Nov. 5, 2018. [cited by applicant]
Liu, C., et al., “FloorNet: A Unied Framework for Floorplan Reconstruction from 3D Scans,” dated Mar. 31, 2018. [cited by applicant]
Liu, W., et al., “SSD: Single Shot MultiBox Detector,” Dec. 29, 2016. [cited by applicant]
Murali, S., et al., “Indoor Scan2BIM: Building Information Models of House Interiors,” dated 2017. [cited by applicant]
Ng, A., et al., “On Spectral Clustering: Analysis and an algorithm,” dated 2001. [cited by applicant]
Phalak, A., et al., “DeepPerimeter: Indoor Boundary Estimation from Posed Monocular Sequences,” Magic Leap, Inc., dated Jul. 1, 2019. [cited by applicant]
Wu, B., et al., “DGCNN: Disordered Graph Convolutional Neural Network Based on the Gaussian Mixture Model,” State Key Laboratory of Software Development Environment, Beihang University, P.R.China, dated Dec. 10, 2017. [cited by applicant]
Qi, C., et al., “Deep Hough Voting for 3D Object Detection in Point Clouds,” dated Aug. 22, 2019. [cited by applicant]
Qi, C., et al., “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” dated Apr. 10, 2017. [cited by applicant]
Qi, C., et al., “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space,” Stanford University, dated Jun. 7, 2017. [cited by applicant]
Qi, X., et al., “3D Graph Neural Networks for RGBD Semantic Segmentation,” dated 2017. [cited by applicant]
Redmon, J., et al., “You Only Look Once: Unified, Real-Time Object Detection,” dated May 9, 2016. [cited by applicant]
Ren, S., et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” dated Jan. 6, 2016. [cited by applicant]
Schubert, E., et al., “DBSCAN Revisited, Revisited: Why and How You Should (Still) Use DBSCAN,” ACM Transactions on Database Systems, vol. 42, No. 3, Article 19. Publication date: Jul. 2017. [cited by applicant]
Shen, Y., et al., “Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling,” dated Apr. 3, 2018. [cited by applicant]
Shukla, A., et al., “Semi-Supervised Clustering with Neural Networks,” IIIT-Delhi, India, dated Jul. 10, 2018. [cited by applicant]
Tatarchenko, M., et al., “Octree Generating Networks: Efficient Convolutional Architectures for High-resolution 3D Outputs,” dated Aug. 7, 2017. [cited by applicant]
Wang, P., et al., “O-CNN: Octree-based Convolutional Neural Networks for 3D Shape Analysis,” ACM Transactions on Graphics, vol. 36, No. 4, Article 72. Publication date: Jul. 2017. [cited by applicant]
Wang, Y., et al., “Dynamic Graph CNN for Learning on Point Clouds,” ACM Trans. Graph., vol. 1, No. 1, Article 1. Publication date: Jan. 2019. [cited by applicant]
Xie, J., et al., “Unsupervised Deep Embedding for Clustering Analysis,” Proceedings of the 33 rd International Conference on Machine Learning, New York, NY, USA, 2016. MLR: W&CP vol. 48. Copyright 2016. [cited by applicant]
Zhang, J., et al., “Estimating the 3D Layout of Indoor Scenes and its Clutter from Depth Sensors,” Stanford University, dated 2013. [cited by applicant]
Zhang, T., et al., “BIRCH: A New Data Clustering Algorithm and Its Applications,” Data Mining and Knowledge Discovery, 1, 141-182 (1997), copyright 1997 Kluwer Academic Publishers. [cited by applicant]
Zhang, W., et al., “Edge-Semantic Learning Strategy for Layout Estimation in Indoor Environment,” School of Control Science and Engineering, Shandong University, Jinan, Shandong, dated Jan. 3, 2019. [cited by applicant]
Zhao, H., et al., “Pyramid Scene Parsing Network,” dated Apr. 27, 2017. [cited by applicant]
Zheng, J., et al., “Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling,” dated Jul. 17, 2020. [cited by applicant]
Zhou, Y., et al., “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection,” dated Nov. 17, 2017. [cited by applicant]
Zou, C., et al., “LayoutNet: Reconstructing the 3D Room Layout from a Single RGB Image,” dated Mar. 23, 2018. [cited by applicant]
Extended European Search Report for EP Patent Appln. No. 21763943.4 dated Jul. 25, 2023. [cited by applicant]
Ameya Phalak et al: “DeepPerimeter: Indoor Boundary Estimation from Posed Monocular Sequences”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 25, 2019 (Apr. 25, 2019),… [cited by applicant]
Cui Yang et al: “Automatic 3-D Reconstruction of Indoor Environment With Mobile Laser Scanning Point Clouds”, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, IEEE, USA, vol. 12 , No. 8,… [cited by applicant]
Foreign Exam Report for EP Patent Appln. No. 21 763 943.4 dated Jul. 27, 2026. [cited by applicant]