IP Library › Granted Patent US 12,675,950
Granted Patent B2
US 12,675,950 · App. 18/629,613 · Granted Jul 7, 2026

Method and an electronic device for 3D scene reconstruction and visualization

Inventors: Anna Ilyinichna Sokolova (Moscow, RU); Anna Borisovna Vorontsova (Moscow, RU); Alexander Georgievich Limonov (Moscow, RU)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T17/20G06T7/11G06V10/761G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,950
App. No.
18/629,613
Filed
Apr 8, 2024
Granted
Jul 7, 2026
Kind
B2
Art Unit
2616
USPC
345/420
Abstract

A method for 3D scene reconstruction and visualization, may include, using at least one processor: obtaining a trained base neural network by training the base neural network for obtaining distance information for voxels of a real scene; operating the trained base neural network for obtaining the distance information of an input sequence of frames of the real scene; inputting the distance information to an algorithm that outputs a 3D reconstruction of the real scene; obtaining a 3D visualization of the real scene by rendering the 3D reconstruction of the real scene; and instructing at least one display to display the 3D visualization of the real scene.

Claims (68)

1 . A method for 3D scene reconstruction and visualization, the method comprising, using at least one processor:

obtaining a trained base neural network by training the base neural network for obtaining distance information for voxels of a real scene;

operating the trained base neural network for obtaining the distance information of an input sequence of frames of the real scene;

inputting the distance information to an algorithm that outputs a 3D reconstruction of the real scene;

obtaining a 3D visualization of the real scene by rendering the 3D reconstruction of the real scene; and

instructing at least one display to display the 3D visualization of the real scene,

wherein the training the base neural network includes:

inputting training data of a training scene, including a training sequence of RGB (Red, Green, Blue) frames with corresponding camera pose data, into a backbone of the base neural network;

computing a common loss function as a sum of Truncated Signed Distance Function (TSDF) losses, segmentation losses, and normal losses; and

minimizing the common loss function to obtain a minimized common loss function.

2 . The method of claim 1 , wherein the training the base neural network includes:

computing the TSDF losses between a TSDF prediction, obtained by a TSDF head of the base neural network from data output from the backbone, and a ground truth scan;

deriving, by a segmentation head of the base neural network from the data output from the backbone, a segmentation prediction defining floor areas, wall areas, and other areas for each training voxel in the training scene;

computing the segmentation losses between the segmentation prediction and a segmentation ground truth;

computing normal coordinates as gradients of the TSDF prediction of TSDF values in each training voxel;

computing the normal losses for normals for the wall areas and the floor areas based on gradients of TSDF for the ground truth scan and gradients of the TSDF prediction.

3 . The method of claim 2 , wherein the segmentation head and TSDF head are used in parallel while training.

4 . The method of claim 2 , wherein the minimizing the common loss function comprises:

minimizing the common loss function by computing a gradient of the common loss function over parameters of the base neural network.

5 . The method of claim 2 , further comprising updating parameters of the base neural network by providing backpropagation of the minimized common loss function and updating the parameters of the base neural network according to the minimized common loss function.

6 . The method of claim 2 , wherein the minimizing the common loss function comprises:

minimizing the common loss function repeatedly until the common loss function stops decreasing.

7 . The method of claim 2 , wherein the minimizing the common loss function comprises:

minimizing the common loss function repeatedly until the common loss function reaches a preset threshold value.

8 . The method of claim 2 , further comprising, after the computing normal coordinates:

selecting wall normal vectors in the wall areas and computing vertical components of the wall normal vectors; and

selecting floor normal vectors in the floor areas and computing horizontal components of the floor normal vectors.

9 . The method of claim 8 , wherein the computing the normal losses comprises:

obtaining, at each point in the wall areas, the vertical components of wall normal vectors,

obtaining, at each point in the floor areas, the horizontal components of the floor normal vectors and x- and y-components of the floor normal vectors, and

computing the normal losses by using lengths of z-components of the wall normal vectors, normals of two-dimensional vectors composed of the x- and y-components.

10 . The method of claim 1 , wherein the 3D reconstruction of the real scene comprises a point cloud or triangle mesh.

11 . A non-transitory computer-readable medium storing program instructions that when executed by at least one processor, cause the at least one processor of an electronic device to implement the method of claim 1 .

12 . An electronic device for 3D scene reconstruction and visualization, the electronic device comprising:

at least one display;

at least one memory configured to store instructions; and

at least one processor configured to execute the instructions to:

obtain a trained base neural network by training the base neural network for obtaining distance information for voxels of a real scene;

operate the trained base neural network for obtaining the distance information of an input sequence of frames of the real scene;

input the distance information to an algorithm that outputs a 3D reconstruction of the real scene;

obtain a 3D visualization of the real scene by rendering the 3D reconstruction of the real scene; and

instruct the at least one display to display the 3D visualization of the real scene,

wherein the at least one processor is further configured to execute the instructions to:

input training data of a training scene, including a training sequence of RGB (Red, Green, Blue) frames with corresponding camera pose data, into a backbone of the base neural network;

compute a common loss function as a sum of Truncated Signed Distance Function (TSDF) losses, segmentation losses, and normal losses; and

minimize the common loss function to obtain a minimized common loss function.

13 . The electronic device of claim 12 , wherein the at least one processor is further configured to:

compute the TSDF losses between a TSDF prediction, obtained by a TSDF head of the base neural network from data output from the backbone, and a ground truth scan,

derive, by a segmentation head of the base neural network from the data output from the backbone, a segmentation prediction defining floor areas, wall areas, and other areas for each training voxel in the training scene,

compute the segmentation losses between the segmentation prediction and a segmentation ground truth,

compute normal coordinates as gradients of the TSDF prediction of TSDF values in each training voxel,

compute the normal losses for normals for the wall areas and floor areas based on gradients of TSDF for the ground truth scan and gradients of the TSDF prediction.

14 . The electronic device of claim 13 , wherein the segmentation head and TSDF head are used in parallel while training.

15 . The electronic device of claim 13 , wherein the at least one processor is further configured to:

minimize the common loss function by computing a gradient of the common loss function over parameters of the base neural network.

16 . The electronic device of claim 13 , wherein the at least one processor is further configured to:

update parameters of the base neural network by providing backpropagation of the minimized common loss function, and

update the parameters of the base neural network according to the minimized common loss function.

17 . The electronic device of claim 13 , wherein the at least one processor is further configured to:

minimize the common loss function repeatedly until the common loss function stops decreasing.

18 . The electronic device of claim 12 , wherein the at least one processor is further configured to:

select the wall normal vectors in the wall areas and compute vertical components of the wall normal vectors, and

select floor normal vectors in the floor areas and compute horizontal components of the floor normal vectors.

19 . The electronic device of claim 18 , wherein the at least one processor is further configured to:

obtain, at each point in the wall areas, the vertical components of wall normal vectors,

obtain, at each point in the floor areas, the horizontal components of the floor normal vectors and x- and y-components of the floor normal vectors, and

compute the normal losses by using lengths of z-components of the wall normal vectors, normals of two-dimensional vectors composed of the x- and y-components.

20 . The electronic device of claim 12 , wherein the 3D reconstruction of the real scene comprises a point cloud or triangle mesh.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2024
From: SOKOLOVA, ANNA ILYINICHNA; VORONTSOVA, ANNA BORISOVNA; LIMONOV, ALEXANDER GEORGIEVICH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 067037/0777 →
Priority Claims (2)
RU 2023109346 · Apr 13, 2023 · national
RU 2023122047 · Aug 24, 2023 · national
Continuity (2)
Continuation PCTKR2024004183 · Apr 1, 2024
Related Publication 20240346765A1 · Oct 17, 2024
References Cited (42)
US 8587583B2 · Newcombe et al. · 2013 [cited by applicant]
US 9076250B2 · Kim et al. · 2015 [cited by applicant]
US 9978177B2 · Mehr · 2018 [cited by examiner]
US 11170246B2 · Ono et al. · 2021 [cited by applicant]
US 11188787B1 · Ulbricht et al. · 2021 [cited by applicant]
US 20180018805A1 · Kutliroff · 2018 [cited by examiner]
US 20190043203A1 · Fleishman et al. · 2019 [cited by applicant]
US 20200342674A1 · Chen · 2020 [cited by examiner]
US 20200356899A1 · Rejeb Sfar et al. · 2020 [cited by applicant]
US 20210279950A1 · Phalak · 2021 [cited by examiner]
US 20210390741A1 · Zhang et al. · 2021 [cited by applicant]
US 20220366635A1 · Murez · 2022 [cited by applicant]
US 20220375164A1 · Chen · 2022 [cited by applicant]
US 20230086928A1 · Fang et al. · 2023 [cited by applicant]
US 20230094308A1 · Yang et al. · 2023 [cited by applicant]
US 20230360241A1 · Sayed · 2023 [cited by examiner]
CN 108230337A · 2018 [cited by applicant]
JP 7553266B2 · 2024 [cited by applicant]
KR 101839035B1 · 2018 [cited by applicant]
KR 102483354B1 · 2022 [cited by applicant]
KR 102757806B1 · 2025 [cited by applicant]
RU 2693267C1 · 2019 [cited by applicant]
RU 2776814C1 · 2022 [cited by applicant]
WO 2019089822A1 · 2019 [cited by applicant]
Zhou L, Wu G, Zuo Y, Chen X, Hu H. A comprehensive review of vision-based 3d reconstruction methods. Sensors. Apr. 5, 2024;24(7):2314. [cited by examiner]
Newcombe RA, Izadi S, Hilliges O, Molyneaux D, Kim D, Davison AJ, Kohi P, Shotton J, Hodges S, Fitzgibbon A. Kinectfusion: Real-time dense surface mapping and tracking. In2011 10th IEEE international symposium on mixed … [cited by examiner]
Dias P, Matos M, Santos V. 3D reconstruction of real world scenes using a low-cost 3D range scanner. Computer-Aided Civil and Infrastructure Engineering. Oct. 2006;21(7):486-97. [cited by examiner]
Son H, Kim C. Automatic segmentation and 3D modeling of pipelines into constituent parts from laser-scan data of the built environment. Automation in Construction. Aug. 1, 2016;68:203-11. [cited by examiner]
Lehtola VV, Nikoohemat S, Nüchter A. Indoor 3D: Overview on scanning and reconstruction methods. Handbook of big geospatial data. Dec. 17, 2020:55-97. [cited by examiner]
Wang H, Li M. A New Era of Indoor Scene Reconstruction: A Survey. IEEE Access. 2024; 12:110160-92. [cited by examiner]
Kang Z, Yang J, Yang Z, Cheng S. A review of techniques for 3d reconstruction of indoor environments. ISPRS International Journal of Geo-Information. May 19, 2020;9(5):330. [cited by examiner]
Ning X, Ma J, Lv Z, Xu Q, Wang Y. Structure reconstruction of indoor scene from terrestrial laser scanner. InInternational Conference on E-Learning and Games Jun. 28, 2018 (pp. 91-98). Cham: Springer International Publi… [cited by examiner]
Liu C, Wu J, Furukawa Y. FloorNet: A Unified Framework for Floorplan Reconstruction from 3D Scans. arXiv preprint arXiv: 1804.00090. Mar. 31, 2018. [cited by examiner]
Li J, Gao W, Wu Y, Liu Y, Shen Y. High-quality indoor scene 3D reconstruction with RGB-D cameras: A brief review. Computational Visual Media. Sep. 2022;8(3):369-93. [cited by examiner]
Xiao J, Owens A, Torralba A. Sun3d: A database of big spaces reconstructed using sfm and object labels. InProceedings of the IEEE international conference on computer vision 2013 (pp. 1625-1632). [cited by examiner]
Noah Stier et al., “VoRTX: Volumetric 3D Reconstruction with Transformers for Voxelwise View Selection and Fusion”, Proceedings of the 2021 International Conference on 3D Vision (3DV), Dec. 2021, pp. 320-330, DOI: 10.11… [cited by applicant]
Communication issued on Apr. 18, 2024 by the Russian Patent Office for Russian Patent Application No. 2023122047. [cited by applicant]
Communication issued on Mar. 27, 2024 by the Russian Patent Office for Russian Patent Application No. 2023122047. [cited by applicant]
International Communications (PCT/ISA/210 & 237) dated Jul. 3, 2024, issued by the International Searching Authority counterpart in International Application No. PCT/KR2024/004183. [cited by applicant]
Communication dated Jul. 12, 2024, issued by Russian Patent Office counterpart in Russian Application No. 2023122047. [cited by applicant]
Wikipedia, “Signed distance function”, Jun. 22, 2006 , 2 pages, https://en.wikipedia.org/wiki/Signed_distance_function. [cited by applicant]
Brian Curless et al., “A Volumetric Method for Building Complex Models from Range Images”, Proceedings of SIGGRAPH '96, Aug. 1, 1996, 10 pages. [cited by applicant]