IP Library › Granted Patent US 12,505,618
Granted Patent B2
US 12,505,618 · App. 17/981,156 · Granted Dec 23, 2025

Manhattan layout estimation using geometric and semantic information

Inventors: Haichao Zhu (Los Angeles, CA); Bing Jian (Cupertino, CA); Weiwei Feng (Mountain View, CA); Lu He (Palo Alto, CA); Kelin Liu (Thornhill, CA); Shan Liu (San Jose, CA)
Assignee: Tencent America LLC
G06T17/20G06T7/33G06T7/55G06T7/60G06T11/203G06T15/06G06V10/764G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,618
App. No.
17/981,156
Filed
Nov 4, 2022
Granted
Dec 23, 2025
Kind
B2
Art Unit
2611
USPC
345/423
Abstract

A plurality of two-dimensional (2D) images of the scene is received. Geometric information and semantic information of each of the plurality of 2D images is determined. The geometric information indicates a detected line and a reference direction in the respective 2D image. The semantic information includes classification information of pixels in the respective 2D image. A layout estimation associated with the respective 2D image of the scene is determined based on the geometric information and the semantic information of the respective 2D image. A combined layout estimation associated with the scene is determined based on a plurality of the determined layout estimations associated with the plurality of 2D images of the scene. The Manhattan layout associated with the scene is generated based on the combined layout estimation. The Manhattan layout includes at least a three-dimensional (3D) shape of the scene that includes wall faces orthogonal with respect to each other.

Claims (83)

1 . A method for estimating a Manhattan layout associated with a scene, the method comprising:

receiving a plurality of two-dimensional (2D) images of the scene;

determining geometric information and semantic information of each of the plurality of 2D images, the geometric information indicating a detected line and a reference direction in the respective 2D image, the semantic information including classification information of pixels in the respective 2D image;

determining a plurality of layout estimations associated with the plurality of 2D images of the scene, each layout estimation being associated with the respective 2D image of the scene based on the geometric information and the semantic information of the respective 2D image;

determining a combined layout estimation associated with the scene based on a shrunk polygon that is generated based on a plurality of candidate edges, the shrunk polygon being a portion of a base polygon, the base polygon being formed by combining the plurality of the determined layout estimations associated with the plurality of 2D images of the scene, which one of the plurality of candidate edges is selected for the shrunk polygon being determined based on whether (i) the one of the plurality of candidate edges is parallel to a corresponding edge of the base polygon, (ii) a projected overlapping portion between the one of the plurality of candidate edges and the corresponding edge of the base polygon is larger than a threshold, and (iii) the one of the plurality of candidate edges is closer to an original view position associated with a corresponding 2D image than the corresponding edge of the base polygon; and

generating the Manhattan layout associated with the scene based on the combined layout estimation, the Manhattan layout including at least a three-dimensional (3D) shape of the scene that includes wall faces orthogonal with respect to each other.

2 . The method of claim 1 , wherein the determining the geometric information and the semantic information further comprises:

extracting first geometric information of a first 2D image of the plurality of 2D images, the first geometric information including at least one of detected lines, reference directions of the first 2D image, a ratio of a first distance from a ceiling to a ground and a second distance from a camera to the ground, or a relative pose between the first 2D image and a second 2D image of the plurality of 2D images; and

labeling pixels of the first 2D image to generate first semantic information, the first semantic information indicating first structure information of the pixels in the first 2D image.

3 . The method of claim 2 , wherein the determining the layout estimation associated with the respective 2D image of the scene further comprises:

determining a first layout estimation of the plurality of the determined layout estimations associated with the scene based on the first geometric information and the first semantic information of the first 2D image; and

the determining the first layout estimation further comprises:

determining whether each of the detected lines is a borderline that corresponds to a wall border in the scene;

aligning the borderlines of the detected lines with the reference directions of the first 2D image; and

generating a first polygon that indicates the first layout estimation based on the aligned borderlines with one of a 2D polygon denoising and a staircase removal.

4 . The method of claim 3 , wherein the generating the first polygon further comprises completing a plurality of incomplete borderlines of the borderlines based on one of:

estimating the plurality of incomplete borderlines based on a combination of a ceiling borderline and a floor borderline of the borderlines; and

connecting a pair of incomplete borderlines of the plurality of incomplete borderlines based on one of (i) adding a perpendicular line to the pair of incomplete borderlines when the pair of incomplete borderlines are parallel and (ii) extending at least one of the pair of incomplete borderlines such that an intersection of the pair of incomplete borderlines is positioned on the extended pair of incomplete borderlines.

5 . The method of claim 3 , wherein the determining the combined layout estimation associated with the scene further comprises:

determining the base polygon by combining a plurality of polygons via a polygon union algorithm, each of the plurality of polygons corresponding to a respective layout estimation of the plurality of the determined layout estimations;

determining the shrunk polygon based on the base polygon, the shrunk polygon including updated edges that are updated from edges of the base polygon; and

determining a final polygon based on the shrunk polygon with one of the 2D polygon denoising and the staircase removal, the final polygon corresponding to the combined layout estimation associated with the scene.

6 . The method of claim 5 , wherein the determining the shrunk polygon further comprises:

determining the plurality of candidate edges from the plurality of polygons for the edges of the base polygon, each of the plurality of candidate edges corresponding to a respective edge of the base polygon; and

generating the updated edges of the shrunk polygon by replacing one or more edges of the base polygon with the corresponding one or more candidate edges when the one or more candidate edges are closer to original view positions in the plurality of 2D images than the corresponding one or more edges of the base polygon.

7 . The method of claim 5 , wherein the determining the combined layout estimation associated with the scene further comprises:

determining an edge set that includes edges of the final polygon;

generating a plurality of edge groups based on the edge set; and

generating a plurality of internal edges of the final polygon that is indicated by a plurality of average edges of one or more edge groups of the edge set, each of the one or more edge groups of the plurality of edge groups including a respective number of edges that is greater than a target value, each of the plurality of average edges being obtained by averaging edges of a respective one of the one or more edge groups.

8 . The method of claim 7 , wherein:

the plurality of edge groups includes a first edge group, and

the first edge group further includes a first edge and a second edge, the first edge and the second edge being parallel, a distance between the first edge and the second edge being less than a first threshold, and a projected overlapping region between the first edge and the second edge being greater than a second threshold.

9 . The method of claim 1 , wherein the generating the Manhattan layout associated with the scene further comprises:

generating the Manhattan layout associated with the scene based on one of triangle meshes triangulated from the combined layout estimation, quadrilateral meshes quadrangulated from the combined layout estimation, sampling points sampled from one of the triangle meshes and the quadrilateral meshes, or discrete grids generated from one of the triangle meshes and the quadrilateral meshes via voxelization.

10 . The method of claim 9 , wherein:

the Manhattan layout associated with the scene is generated based on the triangle meshes triangulated from the combined layout estimation, and

the generating the Manhattan layout associated with the scene further comprises:

generating a ceiling face and a floor face in the scene by triangulating the combined layout estimation;

generating the wall faces in the scene by triangulating rectangles that surround a ceiling borderline and a floor borderline in the scene; and

generating textures of the Manhattan layout associated with the scene via a ray-casting based process.

11 . An apparatus for estimating a Manhattan layout associated with a scene, the apparatus comprising:

processing circuitry configured to:

receive a plurality of two-dimensional (2D) images of the scene;

determine geometric information and semantic information of each of the plurality of 2D images, the geometric information indicating a detected line and a reference direction in the respective 2D image, the semantic information including classification information of pixels in the respective 2D image;

determine a plurality of layout estimations associated with the plurality of 2D images of the scene, each layout estimation being associated with the respective 2D image of the scene based on the geometric information and the semantic information of the respective 2D image;

determine a combined layout estimation associated with the scene based on a shrunk polygon that is generated based on a plurality of candidate edges, the shrunk polygon being a portion of a base polygon, the base polygon being formed by combining the plurality of the determined layout estimations associated with the plurality of 2D images of the scene, which one of the plurality of candidate edges is selected for the shrunk polygon being determined based on whether (i) the one of the plurality of candidate edges is parallel to a corresponding edge of the base polygon, (ii) a projected overlapping portion between the one of the plurality of candidate edges and the corresponding edge of the base polygon is larger than a threshold, and (iii) the one of the plurality of candidate edges is closer to an original view position associated with a corresponding 2D image than the corresponding edge of the base polygon; and

generate the Manhattan layout associated with the scene based on the combined layout estimation, the Manhattan layout including at least a three-dimensional (3D) shape of the scene that includes wall faces orthogonal with respect to each other.

12 . The apparatus of claim 11 , wherein the processing circuitry is configured to:

extract first geometric information of a first 2D image of the plurality of 2D images, the first geometric information including at least one of detected lines, reference directions of the first 2D image, a ratio of a first distance from a ceiling to a ground and a second distance from a camera to the ground, or a relative pose between the first 2D image and a second 2D image of the plurality of 2D images; and

label pixels of the first 2D image to generate first semantic information, the first semantic information indicating first structure information of the pixels in the first 2D image.

13 . The apparatus of claim 12 , wherein

the processing circuitry is configured to:

determine a first layout estimation of the plurality of the determined layout estimations associated with the scene based on the first geometric information and the first semantic information of the first 2D image; and

to determine the first layout estimation, the processing circuitry is further configured to:

determine whether each of the detected lines is a borderline that corresponds to a wall border in the scene;

align the borderlines of the detected lines with the reference directions of the first 2D image; and

generate a first polygon that indicates the first layout estimation based on the aligned borderlines with one of a 2D polygon denoising and a staircase removal.

14 . The apparatus of claim 13 , wherein the processing circuitry is configured to:

complete a plurality of incomplete borderlines of the borderlines based on one of:

estimating the plurality of incomplete borderlines based on a combination of a ceiling borderline and a floor borderline of the borderlines; and

connecting a pair of incomplete borderlines of the plurality of incomplete borderlines based on one of (i) adding a perpendicular line to the pair of incomplete borderlines when the pair of incomplete borderlines are parallel and (ii) extending at least one of the pair of incomplete borderlines such that an intersection of the pair of incomplete borderlines is positioned on the extended pair of incomplete borderlines.

15 . The apparatus of claim 13 , wherein the processing circuitry is configured to:

determine the base polygon by combining a plurality of polygons via a polygon union algorithm, each of the plurality of polygons corresponding to a respective layout estimation of the plurality of the determined layout estimations;

determine the shrunk polygon based on the base polygon, the shrunk polygon including updated edges that are updated from edges of the base polygon; and

determine a final polygon based on the shrunk polygon with one of the 2D polygon denoising and the staircase removal, the final polygon corresponding to the combined layout estimation associated with the scene.

16 . The apparatus of claim 15 , wherein the processing circuitry is configured to:

determine the plurality of candidate edges from the plurality of polygons for the edges of the base polygon, each of the plurality of candidate edges corresponding to a respective edge of the base polygon; and

generate the updated edges of the shrunk polygon by replacing one or more edges of the base polygon with the corresponding one or more candidate edges when the one or more candidate edges are closer to original view positions in the plurality of 2D images than the corresponding one or more edges of the base polygon.

17 . The apparatus of claim 15 , wherein the processing circuitry is configured to:

determine an edge set that includes edges of the final polygon;

generate a plurality of edge groups based on the edge set; and

generate a plurality of internal edges of the final polygon that is indicated by a plurality of average edges of one or more edge groups of the edge set, each of the one or more edge groups of the plurality of edge groups including a respective number of edges that is greater than a target value, each of the plurality of average edges being obtained by averaging edges of a respective one of the one or more edge groups.

18 . The apparatus of claim 17 , wherein:

the plurality of edge groups includes a first edge group, and

the first edge group further includes a first edge and a second edge, the first edge and the second edge being parallel, a distance between the first edge and the second edge being less than a first threshold, and a projected overlapping region between the first edge and the second edge being greater than a second threshold.

19 . The apparatus of claim 18 , wherein the processing circuitry is configured to:

generate the Manhattan layout associated with the scene based on one of triangle meshes triangulated from the combined layout estimation, quadrilateral meshes quadrangulated from the combined layout estimation, sampling points sampled from one of the triangle meshes and the quadrilateral meshes, or discrete grids generated from one of the triangle meshes and the quadrilateral meshes via voxelization.

20 . The apparatus of claim 19 , wherein:

the Manhattan layout associated with the scene is generated based on the triangle meshes triangulated from the combined layout estimation, and

to generate the Manhattan layout associated with the scene, the processing circuitry is configured to:

generate a ceiling face and a floor face in the scene by triangulating the combined layout estimation;

generate the wall faces in the scene by triangulating rectangles that surround a ceiling borderline and a floor borderline in the scene; and

generate textures of the Manhattan layout associated with the scene via a ray-casting based process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: ZHU, HAICHAO; JIAN, BING; FENG, WEIWEI; HE, LU; LIU, KELIN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 061662/0195 →
Continuity (2)
Provisional Application 63306001 · Feb 2, 2022
Related Publication 20230245390A1 · Aug 3, 2023
References Cited (111)
US 7302359B2 · McKitterick · 2007 [cited by examiner]
US 10026218B1 · Mertens · 2018 [cited by examiner]
US 10809066B2 · Colburn · 2020 [cited by examiner]
US 11004202B2 · Tchapmi · 2021 [cited by examiner]
US 11087479B1 · Geraghty · 2021 [cited by examiner]
US 11164368B2 · Vincent · 2021 [cited by examiner]
US 11188787B1 · Ulbricht · 2021 [cited by examiner]
US 11252329B1 · Cier · 2022 [cited by examiner]
US 11274929B1 · Afrouzi · 2022 [cited by examiner]
US 11393114B1 · Ebrahimi Afrouzi · 2022 [cited by examiner]
US 11481925B1 · Li · 2022 [cited by examiner]
US 11481970B1 · Mertens · 2022 [cited by examiner]
US 11501492B1 · Li · 2022 [cited by examiner]
US 11507713B2 · Tran · 2022 [cited by examiner]
US 11514633B1 · Çetintaş · 2022 [cited by examiner]
US 11561102B1 · Ebrahimi Afrouzi · 2023 [cited by examiner]
US 11600049B2 · Chen · 2023 [cited by examiner]
US 11830135B1 · Khosravan · 2023 [cited by examiner]
US 11989834B1 · Agarwal · 2024 [cited by examiner]
US 12140419B2 · Pershing · 2024 [cited by examiner]
US 12147239B2 · Lai · 2024 [cited by examiner]
US 12175562B2 · Hutchcroft · 2024 [cited by examiner]
US 12190581B2 · Stoeva · 2025 [cited by examiner]
US 20040196480A1 · Foster · 2004 [cited by examiner]
US 20060053399A1 · Honda · 2006 [cited by examiner]
US 20070185681A1 · McKitterick · 2007 [cited by examiner]
US 20130259308A1 · Klusza · 2013 [cited by examiner]
US 20140301633A1 · Furukawa · 2014 [cited by examiner]
US 20150116509A1 · Birkler · 2015 [cited by examiner]
US 20150228095A1 · Chen · 2015 [cited by examiner]
US 20150286893A1 · Straub · 2015 [cited by examiner]
US 20160110916A1 · Eikhoff · 2016 [cited by examiner]
US 20180032643A1 · Wright · 2018 [cited by examiner]
US 20180046733A1 · Wong · 2018 [cited by examiner]
US 20180149481A1 · Hirai · 2018 [cited by examiner]
US 20180268220A1 · Lee · 2018 [cited by examiner]
US 20180315162A1 · Sturm · 2018 [cited by examiner]
US 20180330184A1 · Mehr · 2018 [cited by examiner]
US 20180374225A1 · Edge · 2018 [cited by examiner]
US 20190120633A1 · Afrouzi · 2019 [cited by examiner]
US 20190205485A1 · Rejeb Sfar · 2019 [cited by examiner]
US 20190243928A1 · Rejeb Sfar · 2019 [cited by examiner]
US 20190266293A1 · Ishida · 2019 [cited by examiner]
US 20200007841A1 · Sedeffow · 2020 [cited by examiner]
US 20200082198A1 · Yao · 2020 [cited by examiner]
US 20200211284A1 · Lin · 2020 [cited by examiner]
US 20200312013A1 · Dougherty · 2020 [cited by examiner]
US 20200394849A1 · Barker · 2020 [cited by examiner]
US 20210049784A1 · Török · 2021 [cited by examiner]
US 20210064216A1 · Li · 2021 [cited by examiner]
US 20210073449A1 · Segev · 2021 [cited by examiner]
US 20210125397A1 · Moulon · 2021 [cited by examiner]
US 20210127060A1 · Ji · 2021 [cited by examiner]
US 20210150805A1 · Stekovic · 2021 [cited by examiner]
US 20210158609A1 · Raskob · 2021 [cited by examiner]
US 20210279950A1 · Phalak · 2021 [cited by examiner]
US 20210385378A1 · Cier · 2021 [cited by examiner]
US 20220003555A1 · Colburn · 2022 [cited by examiner]
US 20220027656A1 · Jia · 2022 [cited by examiner]
US 20220114291A1 · Li · 2022 [cited by examiner]
US 20220148327A1 · Fu · 2022 [cited by examiner]
US 20220156426A1 · Hampali · 2022 [cited by examiner]
US 20220164493A1 · Li · 2022 [cited by examiner]
US 20220224833A1 · Cier · 2022 [cited by examiner]
US 20220269885A1 · Wixson · 2022 [cited by examiner]
US 20220269888A1 · Stoeva · 2022 [cited by examiner]
US 20220358694A1 · Zhang · 2022 [cited by examiner]
US 20220358716A1 · Zhang · 2022 [cited by examiner]
US 20220383572A1 · Hu · 2022 [cited by examiner]
US 20220406007A1 · Ton-That · 2022 [cited by examiner]
US 20230035601A1 · Alimo · 2023 [cited by examiner]
US 20230071446A1 · Narayana · 2023 [cited by examiner]
US 20230095173A1 · Khosravan · 2023 [cited by examiner]
US 20230106339A1 · Goyal · 2023 [cited by examiner]
US 20230128740A1 · Hong · 2023 [cited by examiner]
US 20230138762A1 · Lambert · 2023 [cited by examiner]
US 20230154110A1 · Stout · 2023 [cited by examiner]
US 20230184949A1 · Huang · 2023 [cited by examiner]
US 20230196670A1 · Su · 2023 [cited by examiner]
US 20230206393A1 · Hutchcroft · 2023 [cited by examiner]
US 20230259667A1 · Palmer · 2023 [cited by examiner]
US 20230306539A1 · Frei · 2023 [cited by examiner]
US 20230306629A1 · Cho · 2023 [cited by examiner]
US 20230334803A1 · Odamaki · 2023 [cited by examiner]
US 20230368458A1 · Dryer · 2023 [cited by examiner]
US 20230409766A1 · Narayana · 2023 [cited by examiner]
US 20230419526A1 · Lianos · 2023 [cited by examiner]
US 20240029352A1 · Wan · 2024 [cited by examiner]
US 20240096097A1 · Penner · 2024 [cited by examiner]
US 20240118103A1 · Shin · 2024 [cited by examiner]
US 20240161348A1 · Hutchcroft · 2024 [cited by examiner]
US 20240312136A1 · Narayana · 2024 [cited by examiner]
US 20240371061A1 · Wei · 2024 [cited by examiner]
JP 2022501684A · 2022 [cited by applicant]
Zou, Chuhang, et al. “Layoutnet: Reconstructing the 3d room layout from a single rgb image.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. (Year: 2018). [cited by examiner]
Liu, Chenxi, et al. “Rent3d: Floor-plan priors for monocular layout estimation.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. (Year: 2015). [cited by examiner]
Zou, Chuhang, et al. “Manhattan Room Layout Reconstruction from a Single 360 [cited by examiner]
Fernandes, L.A. and Oliveira, M.M., 2008. Real-time line detection through an improved Hough transform voting scheme. Pattern recognition, 41(1), pp. 299-314. [cited by applicant]
Zhang, Y., Song, S., Tan, P. and Xiao, J., Sep. 2014. Panocontext: A whole-room 3d context model for panoramic scene understanding. In European conference on computer vision (pp. 668-686). Springer, Cham. [cited by applicant]
Long, J., Shelhamer, E. and Darrell, T., 2015. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3431-3440). [cited by applicant]
Zitova, B. and Flusser, J., 2003. Image registration methods: a survey. Image and vision computing, 21(11), pp. 977-1000. [cited by applicant]
Maillot, Patrick-Gilles. “A new, fast method for 2D polygon clipping: analysis and software implementation.” ACM Transactions on Graphics (TOG) 11, No. 3 (1992):276-290. [cited by applicant]
Narkhede, A. and Manocha, D., 1995. Fast polygon triangulation based on seidel's algorithm. In Graphics Gems V (pp. 394-397). Academic Press. [cited by applicant]
Schmitt, A., Muller, H. and Leister, W., 1988. Ray tracing algorithms—theory and practice. In Theoretical foundations of Computer Graphics and CAD (pp. 997-1030). Springer, Berlin, Heidelberg. [cited by applicant]
Nakamoto A, Kawatani G, Matsumoto N, Urrutia J. Geometric quadrangulations of a polygon. Electronic Notes in Discrete Mathematics. Jul. 1, 2018;68:59-64. [cited by applicant]
Cohen-Or D, Kaufman A. Fundamentals of surface voxelization. Graphical models and image processing. Nov. 1, 1995;57(6):453-61. [cited by applicant]
Sik, Martin, and Jaroslav Krivanek. “Fast Random Sampling of Triangular Meshes.” (2013), pp. 1-6. [cited by applicant]
Office Action received for Japanese Patent Application No. 2023-560170, mailed on Aug. 5, 2024, 9 pages (4 pages of English Translation and 5 pages of Original Document). [cited by applicant]
Cabral et al., “Piecewise Planar and Compact Floorplan Reconstruction from Images”, 2014 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2014, pp. 628-635. [cited by applicant]
Extended European Search Report received for European Patent Application No. 22925212.7, mailed on May 2, 2025, 13 pages. [cited by applicant]
Pintore et al., “Recovering 3D existing-conditions of indoor structures from spherical images”, Computers and Graphics, Elsevier, vol. 77, Sep. 2018, pp. 16-29. [cited by applicant]