IP Library Granted Patent US 12,475,666
Granted Patent B2
US 12,475,666 · App. 18/742,947 · Granted Nov 18, 2025

Systems and methods for defurnishing and furnishing spaces, and removing objects from spaces

Inventors: David Alan Gausebeck (Sunnyvale, CA); Japjit Tulsi (Sunnyvale, CA); Rj Pittman (Sunnyvale, CA); David Lippman (Sunnyvale, CA); Vivek Tanna (Sunnyvale, CA); Nicole Guernsey (Sunnyvale, CA); Olaf Brandt (Sunnyvale, CA)
Assignee: CoStar Realty Information, Inc.
G06T19/20G06T17/20G06V10/82G06V20/70G06T2200/24G06T2210/04G06T2219/004G06T2219/2024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,666
App. No.
18/742,947
Granted
Nov 18, 2025
Kind
B2
Abstract

An example system may access data of a multidimensional space representing a physical environment and identify interior elements within the multidimensional space using a first machine learning model. The interior element may represent furniture in the physical environment. The system may mask one or more of the interior elements with masks and fill each of the masks with imagery of the physical environment to create an appearance of a defurnished space, the defurnished space having the one or more interior elements appearing as missing from the multidimensional space representing the physical environment. The system may provide all or some of the defurnished space for display.

Claims (56)

1 . A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method, the method comprising:

accessing data of a multidimensional space representing a physical environment, the data including input 2D images of the physical environment;

identifying interior elements within the input 2D images using a machine learning model, at least one of the interior elements representing furniture in the physical environment;

masking one or more of the interior elements within the input 2D images with masks;

filling each of the masks with imagery of the physical environment to generate filled 2D images, wherein filling includes utilizing a latent diffusion model to fill each of the masks with imagery of the physical environment to generate the filled 2D images, wherein the latent diffusion model was trained using first 2D images of unfurnished spaces, and wherein the latent diffusion model was fine-tuned using second 2D images of unfurnished spaces modified to include artifacts not present in the second 2D images of unfurnished spaces;

generating, based on the input 2D images and the filled 2D images, output 2D images that create an appearance of a defurnished space, the defurnished space having the one or more interior elements appearing as missing from the multidimensional space representing the physical environment;

providing, using the output 2D images, all or some of the defurnished space for display;

generating one or more additional interior elements within the defurnished space based at least on characteristics of the one or more additional interior elements;

positioning the one or more additional interior elements within the defurnished space based on a geometry of the defurnished space and a layout rule, the layout rule being applied to a position or orientation of the one or more additional interior elements within the defurnished space; and

providing all or some of the defurnished space with the one or more additional interior elements positioned within the defurnished space for display.

2 . The non-transitory computer-readable medium of claim 1 , wherein the physical environment is a furnished room.

3 . The non-transitory computer-readable medium of claim 1 , wherein the multidimensional space is a 2D representation of the physical environment.

4 . The non-transitory computer-readable medium of claim 3 , wherein the 2D representation of the physical environment is used to generate a corresponding 3D representation of the physical environment.

5 . The non-transitory computer-readable medium of claim 1 , wherein the multidimensional space is a 3D representation of the physical environment.

6 . The non-transitory computer-readable medium of claim 1 , wherein the data of the multidimensional space further includes a textured 3D mesh.

7 . The non-transitory computer-readable medium of claim 1 , wherein the interior elements within the multidimensional space further include at least one wall within the physical environment.

8 . The non-transitory computer-readable medium of claim 1 , wherein the machine learning model includes a semantic segmentation neural network.

9 . The non-transitory computer-readable medium of claim 1 , wherein filling each of the masks with the imagery of the physical environment comprising applying inpainting to the masks.

10 . The non-transitory computer-readable medium of claim 1 , the method further comprising:

receiving a selection of at least one design style type from a user; and

selecting interior elements within the multidimensional space that were previously identified based on the at least one design style type from the user, wherein masking the one or more interior elements comprises masking the one or more interior elements that are of the at least one design style type.

11 . The non-transitory computer-readable medium of claim 1 , the method further comprising:

receiving a selection of at least one design style type from a user; and

selecting interior elements within the multidimensional space that were previously identified based on the at least one design style type from the user, wherein masking the one or more interior elements comprises masking the one or more interior elements that are not of the at least one design style type.

12 . The non-transitory computer-readable medium of claim 1 , wherein the one or more additional interior elements to be added to the defurnished space are requested by a user.

13 . The non-transitory computer-readable medium of claim 12 , the method further comprising:

receiving an additions request from the user, comprising receiving a prompt from a user; and

applying the prompt to a large language model to receive a response from the large language model, the response identifying the one or more additional interior elements.

14 . A system comprising at least one processor and memory containing executable instructions, the executable instructions being executable by the at least one processor to:

access data of a multidimensional space representing a physical environment, the data including input 2D images of the physical environment;

identify interior elements within the input 2D images using a machine learning model, at least one of the interior elements representing furniture in the physical environment;

mask at least some of the interior elements within the input 2D images with masks;

fill each of the masks with imagery of the physical environment to generate filled 2D images, wherein filling includes utilizing a latent diffusion model to fill each of the masks with imagery of the physical environment to generate the filled 2D images, wherein the latent diffusion model was trained using first 2D images of unfurnished spaces, and wherein the latent diffusion model was fine-tuned using second 2D images of unfurnished spaces modified to include artifacts not present in the second 2D images of unfurnished spaces;

generating, based on the input 2D images and the filled 2D images, output 2D images that create an appearance of a defurnished space, the defurnished space having at least some of the interior elements appearing as missing from the multidimensional space representing the physical environment;

provide, using the output 2D images, all or some of the defurnished space for display;

generate one or more additional interior elements within the defurnished space based at least on characteristics of the one or more additional interior elements;

position the one or more additional interior elements within the defurnished space based on a geometry of the defurnished space and a layout rule, the layout rule being applied to a position or orientation of the one or more additional interior elements within the defurnished space; and

provide all or some of the defurnished space with the one or more additional interior elements positioned within the defurnished space for display.

15 . The system of claim 14 , wherein the physical environment is a furnished room.

16 . The system of claim 14 , wherein the multidimensional space is a 2D representation of the physical environment.

17 . The system of claim 16 , wherein the 2D representation of the physical environment is used to generate a corresponding 3D representation of the physical environment.

18 . The system of claim 14 , wherein the multidimensional space is a 3D representation of the physical environment.

19 . The system of claim 14 , wherein the data of the multidimensional space further includes a textured 3D mesh.

20 . The system of claim 14 , wherein the interior elements within the multidimensional space further include at least one wall within the physical environment.

21 . The system of claim 14 , wherein the machine learning model includes a semantic segmentation neural network.

22 . The system of claim 14 , wherein filling each of the masks with the imagery of the physical environment comprising applying inpainting to the masks.

23 . The system of claim 14 , the executable instructions being further executable by the at least one processor to:

receive a selection of at least one design style type from a user; and

select interior elements within the multidimensional space that were previously identified based on the at least one design style type from the user, wherein masking at least some of the interior elements comprises masking at least some of the interior elements that are of the at least one design style type.

24 . The system of claim 14 , the executable instructions being further executable by the at least one processor to:

receive a selection of at least one design style type from a user; and

select interior elements within the multidimensional space that were previously identified based on the at least one design style type from the user, wherein masking at least some of the interior elements comprises masking at least some of the interior elements that are not of the at least one design style type.

25 . The system of claim 14 , wherein the one or more additional interior elements to be added to the defurnished space are requested by a user.

26 . The system of claim 25 , wherein the executable instructions being further configured by the at least one processor to:

receive an additions request from the user, comprising to receive a prompt from a user; and

apply the prompt to a large language model to receive a response from the large language model, the response identifying the one or more additional interior elements.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2025
From: MATTERPORT, LLC
To: COSTAR REALTY INFORMATION, INC.
Reel/Frame 072275/0724 →
MERGER AND CHANGE OF NAME Recorded Sep 10, 2025
From: MATTERPORT, INC.; MATRIX MERGER SUB II LLC
To: MATTERPORT, LLC
Reel/Frame 072212/0401 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2024
From: GAUSEBECK, DAVID ALAN; TULSI, JAPJIT; PITTMAN, RJ; LIPPMAN, DAVID; TANNA, VIVEK; GUERNSEY, NICOLE; BRANDT, OLAF
To: MATTERPORT, INC.
Reel/Frame 068087/0151 →
Continuity (3)
Provisional Application 63567185 · Mar 19, 2024
Provisional Application 63472795 · Jun 13, 2023
Related Publication 20240420437A1 · Dec 19, 2024
References Cited (71)
US 20200242201A1 · Lafreniere · 2020 [cited by examiner]
US 20210027539A1 · Huang et al. · 2021 [cited by applicant]
US 20210142497A1 · Pugh · 2021 [cited by examiner]
US 20220084296A1 · Sadalgi · 2022 [cited by examiner]
US 20220366153A1 · Li et al. · 2022 [cited by applicant]
US 20230177594A1 · Besecker · 2023 [cited by examiner]
US 20230177832A1 · Besecker · 2023 [cited by examiner]
US 20230333644A1 · Cazamias · 2023 [cited by examiner]
US 20240029279A1 · Maschmeyer · 2024 [cited by examiner]
Lortz et al., “Enhancing Product Search with Large Language Models (LLMs)”, Apr. 26, 2023, https://www.databricks.com/blog/enhancing-product-search-large-language-models-llms.html (Year: 2023). [cited by examiner]
Boussaha et al., “Large scale textured mesh reconstruction from mobile mapping images and LiDAR scans”, ISPRS Annals of Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. IV-2, 2018 (Year: 2018). [cited by examiner]
CAD in black, “Revit—How to Hide and Isolate Elements”, Sep. 27, 2021, https://www.youtube.com/watch?v=gs07XpRiK8Y (Year: 2021). [cited by examiner]
Halmetoja, Esa. The role of digital twins and their application for the built environment. Industry 4.0 for the built environment: Methodologies, technologies and skills, pp. 415-442, 2022. [cited by applicant]
Yu, Jiahui et al., Free-form image inpainting with gated convolution, 2019. [cited by applicant]
Zeng, Yu, et al., High-resolution image inpainting with iterative confidence feedback and guided upsampling, 2020. [cited by applicant]
Zhang, Edward et al., No Shadow Left Behind: Removing Objects and their Shadows using Approximate Lighting and Geometry. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. [cited by applicant]
Zhang, Lvmin et al., Adding conditional control to text-to-image diffusion models, 2023. [cited by applicant]
Zhang, Richard, et al., The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR, 2018. [cited by applicant]
Zhang, Yinda, et al., PanoContext: A whole-room 3D context model for panoramic scene understanding. In European Conference on Computer Vision (ECCV) oral presentation, 2014. [cited by applicant]
Zhao, Hengshuang, et al., Pyramid scene parsing network, 2017. [cited by applicant]
Zhou, Bolei, Zhao, Hang, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene Parsing through ADE20K Dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. [cited by applicant]
Zhou, Yiyang, et al., Analyzing and Mitigating Object Hallucination in Large Vision-Language Models. In The Twelfth International Conference on Learning Representations, 2024. [cited by applicant]
International Application No. PCT/US2024/033913, International Search Report and the Written Opinion, dated Nov. 7, 2024, 8 pages. [cited by applicant]
Barnes, Connelly, et al., Patchmatch: a randomized correspondence algorithm for structural image editing. ACM SIG-GRAPH 2009 papers, 2009. [cited by applicant]
Bertalmio, Marcelo, et al., Image inpainting. In Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, p. 417-424, USA, 2000. ACM Press/Addison-Wesley Publishing Co. [cited by applicant]
Chang, Angel, Dai, Angela, Funkhouser, Thomas, Halber, Maciej, Niessner, Matthias, Savva, Manolis, Song, Shuran, Zeng, Andy, and Zhang, Yinda. Matterport3d: Learning from rgb-d data in indoor environments. International… [cited by applicant]
Chen, Zhe, et al., Vision Transformer Adapter for Dense Predictions. In The Eleventh International Conference on Learning Representations (ICLR), 2023. [cited by applicant]
Chi, Lu, et al., Fast Fourier Convolution, In Advances in Neural Information Processing Systems (NeurIPS), 2020. [cited by applicant]
Daniotti, Bruno et al., The development of a bim-based interoperable toolkit for efficient renovation in buildings: From bim to digital twin. Buildings, 12(2):231, 2022. [cited by applicant]
Deitke, Matt, et al., Objaverse: A Universe of Annotated 3D Objects, 2022. [cited by applicant]
Delgado, Juan Manuel Davila and Oyedele, Lukumon. Digital twins for the built environment: learning from conceptual and process models in manufacturing. Advanced Engineering Informatics, 49:101332, 2021. [cited by applicant]
Dhariwal, Prafulla and Nichol, Alexander. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, pp. 8780-8794. Curran Associates, Inc., 2021. [cited by applicant]
Dosovitskiy, Alexey, et al., An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [cited by applicant]
Eyre, Jonathan, and Freeman, Chris. Immersive applications of industrial digital twins. The Industrial Track of EuroVR 2018, pp. 11-13, 2018. [cited by applicant]
Gao, Chao-Chen et al., Layout-guided Indoor Panorama Inpainting with Plane-aware Normalization. In Asian Conference on Computer Vision (ACCV), 2022. [cited by applicant]
Gkitsas, V et al., PanoDR: Spherical panorama diminished reality for indoor scenes, 2021. [cited by applicant]
Gkitsas, Vasileios, et al., Towards full-to-empty room generation with structure-aware feature encoding and soft semantic region-adaptive normalization, 2021. [cited by applicant]
Hays, James and Efros, Alexei A., Scene completion using millions of photographs. ACM Trans. Graph., 26(3):4-es, 2007. [cited by applicant]
Hu, Edward J., et al., LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations, 2022. [cited by applicant]
Iizuka, Satoshi, et al., Globally and locally consistent image completion, ACM Transactions on Graphics, 36:1-14, 2017. [cited by applicant]
Isogawa, Mariko and Mikami, Dan and Iwai, Daisuke and Kimata, Hideaki and Sato, Kosuke. Mask Optimization for Image Inpainting. IEEE Access, 2018. 2. [cited by applicant]
Ji, Guanzhou et al., Virtual Home Staging: Inverse Rendering and Editing an Indoor Panorama under Natural Illumination. In International Symposium on Visual Computing, 2023. [cited by applicant]
Kalantari, Saleh and Neo, Jun Rong Jeffrey. Virtual environments for design research: Lessons learned from use of fully immersive virtual reality in interior design research. Journal of Interior Design, 45(3):27-42, 202… [cited by applicant]
Kirillov, Alexander et al., Segment Anything. arXiv:2304.02643, 2023. [cited by applicant]
Kuehn, Wolfgang, Digital twins for decision making in complex production and logistic enterprises. International Journal of Design & Nature and Ecodynamics, 13(3):260-271, 2018. [cited by applicant]
Kulshreshtha, Prakhar et al., Layout Aware Inpainting for Automated Furniture Removal in Indoor Scenes. In IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR Adjunct), 2022. [cited by applicant]
Liu, Chen, et al., PlanerRCNN: 3D plane detection and reconstruction from a single image, 2019. [cited by applicant]
Liu, Guilin, et al., Image inpainting for irregular holes using partial convolutions, 2018. [cited by applicant]
Liu, Hanchao et al., A Survey on Hallucination in Large Vision-Language Models, 2024. [cited by applicant]
Long, Jonathan, et al., Fully convolutional networks for semantic segmentation, 2015. [cited by applicant]
Lugmayr, Andreas, et al., RePaint: Inpainting using Denoising Diffusion Probabilistic Models, 2022. [cited by applicant]
Mantiuk, Rafal K., et al., FovVideoVDP: a visible difference predictor for wide field-of-view video. ACM Transactions on Graphics (SIGGRAPH), 40(4), 2021. [cited by applicant]
Nazeri, Kamyar, et al., Edgeconnect: Structure guided image inpainting using edge prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2019. [cited by applicant]
Osher, Stanley, et al., An iterative regularization method for total variation-based image restoration. Multiscale Modeling & Simulation, 4(2):460-489, 2005. [cited by applicant]
Pathak, Deepak, et al., Context encoders: Feature learning by inpainting, 2016. [cited by applicant]
Ramakrishnan, Santhosh Kumar, et al., Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks T… [cited by applicant]
Rombach, Robin et al., High-resolution image synthesis with latent diffusion models, 2022. [cited by applicant]
Ronneberger, Olaf et al., U-net: Convolutional networks for biomedical image segmentation, 2015. [cited by applicant]
Ruiz, Nataniel, et al., DreamBooth: Fine Tuning Text-to-image Diffusion Models for Subject-Driven Generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. [cited by applicant]
Schuhmann, Christoph, et al., LAION-5B: An open large-scale dataset for training next generation image-text models. In 36th Conference on Neural Information Processing Systems (NeurIPS), 2022. [cited by applicant]
Shahzad, Muhammad, et al., Digital twins in built environments: an investigation of the characteristics, applications, and challenges. Buildings, 12(2):120, 2022. [cited by applicant]
Song, Yuhang, et al., Spg-net: Segmentation prediction and guidance network for image inpainting, 2018. [cited by applicant]
Sun, Cheng, et al., HorizonNet: Learning Room Layout With 1D Representation and Pano Stretch Data Augmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. [cited by applicant]
Suvorov, Roman, et al., Resolution-robust Large Mask Inpainting with Fourier Convolutions. Winter Conference on Applications of Computer Vision (WACV), 2022. [cited by applicant]
Telea, Alexandru. An image inpainting technique based on the fast marching method. Journal of Graphics Tools, 9, 2004. [cited by applicant]
Tonmoy, S. M. Towhidul Islam et al., A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models, 2024. [cited by applicant]
Wang, Xintao, et al., Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data. In International Conference on Computer Vision Workshops (ICCVW), 2021. [cited by applicant]
Wang, Y. T., et al., A survey of personalized interior design. In Computer Graphics Forum. Wiley Online Library, 2023. [cited by applicant]
Wang, Zhou, et al., Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 13 (4), 2004. [cited by applicant]
Xu, Ziwei, et al., Hallucination is Inevitable: An Innate Limitation of Large Language Models, 2024. [cited by applicant]
Yu, Jiahu et al., Generative image inpainting with contextual attention, 2018. [cited by applicant]