IP Library › Granted Patent US 12,664,621
Granted Patent B2
US 12,664,621 · App. 18/228,472 · Granted Jun 23, 2026

Multi-view segmentation and perceptual inpainting with neural radiance fields

Inventors: Ashkan Mirzaei (Toronto, CA); Tristan TY Aumentado-Armstrong (Toronto, CA); Konstantinos G. Derpanis (Toronto, CA); Marcus A. Brubaker (Toronto, CA); Igor Gilitschenski (Toronto, CA); Aleksai Levinshtein (Thornhill, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T5/77G06T7/11G06V10/774G06V10/82G06V10/945G06V20/49G06T2200/04G06T2200/24G06T2207/10021G06T2207/10028G06T2207/20021G06T2207/20081G06T2207/20084G06T2207/20092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,621
App. No.
18/228,472
Filed
Jul 31, 2023
Granted
Jun 23, 2026
Kind
B2
Art Unit
2668
USPC
382/157
Abstract

A computer-implemented method of configuring an electronic device for inpainting source three-dimensional (3D) scenes, includes: receiving the source 3D scenes and a user's input about a first object of the source 3D scenes; generating accurate object masks about the first object of the source 3D scenes; and generating inpainted 3D scenes of the source 3D scenes by using an inpainting neural radiance field (NeRF) based on the accurate object masks.

Claims (38)

1 . A computer-implemented method of configuring an electronic device for inpainting source three-dimensional (3D) scenes, the computer-implemented method comprising:

receiving the source 3D scenes and a user's input about a first object of the source 3D scenes;

generating accurate object masks about the first object of the source 3D scenes; and

training an inpainting neural radiance field (NeRF) by using the source 3D scenes and the accurate object masks;

generating inpainted 3D scenes of the source 3D scenes by using the inpainting NeRF based on the accurate object masks; and

wherein the training the inpainting NeRF by using the source 3D scenes and the accurate object masks comprises training the inpainting NeRF by using at least a perceptual loss about the source 3D scenes, and

wherein the perceptual loss is used to guide the inpainting NeRF in regions identified by the accurate object masks.

2 . The computer-implemented method of claim 1 , wherein the inpainted 3D scenes are consistent when a set of two-dimensional (2D) images of the inpainted 3D scenes corresponds to 2D projections of the inpainted 3D scenes.

3 . The computer-implemented method of claim 1 , wherein the training the inpainting NeRF by using the source 3D scenes and the accurate object masks as inputs of the inpainting NeRF comprises training the inpainting NeRF by using at least depth priors about the source 3D scenes.

4 . The computer-implemented method of claim 3 , further comprising generating, by the inpainting NeRF, depths about the source 3D scenes, based on point cloud data of the source 3D scenes.

5 . The computer-implemented method of claim 1 , further comprising generating a first segmentation mask about the first object and the first view of the source 3D scenes.

6 . The computer-implemented method of claim 5 , further comprising obtaining coarse 2D object masks at least by propagating the first segmentation mask to other views of the source 3D scenes.

7 . The computer-implemented method of claim 6 , wherein the obtaining coarse 2D object masks at least by propagating the first segmentation mask to other views of the source 3D scenes comprises obtaining coarse 2D object masks about the first object by propagating the first segmentation mask to other views of the source 3D scenes by using a video segmentation method.

8 . The computer-implemented method of claim 1 , wherein the generating the accurate object masks about the first object of the source 3D scenes comprises generating the accurate object masks about the first object of the source 3D scenes by using a semantic segmentation NeRF.

9 . The computer-implemented method of claim 8 , further comprising training the semantic segmentation NeRF based on the first object and the source 3D scenes.

10 . The computer-implemented method of claim 9 , further comprising training the inpainting NeRF based on the source 3D scenes and the accurate object masks.

11 . The computer-implemented method of claim 1 , wherein receiving the source 3D scenes and the user's input about the first object of the source 3D scenes comprises:

selecting a first icon on a display, the first icon indicating that the first object is selected; and

selecting a second icon on the display, the second icon indicating an object other than the first object of the source 3D scenes is not selected.

12 . The computer-implemented method of claim 6 , further comprising:

receiving a user's another input about a second object of the source 3D scenes; and

obtaining a second segmentation mask about the second object of the source 3D scenes,

wherein the obtaining coarse 2D object masks at least by propagating the first segmentation mask to other views of the source 3D scenes comprises obtaining coarse 2D object masks by propagating the first segmentation mask and the second segmentation mask to other views of the source 3D scenes.

13 . The computer-implemented method of claim 1 , wherein the receiving the source 3D scenes and the user's input about the first object of the source 3D scenes comprises:

receiving the user's command about the first object;

recognizing the first object by analyzing the user's command based on a language model; and

detecting the recognized first object on the source 3D scenes by using a scene analysis model.

14 . An electronic device for inpainting source three-dimensional (3D) scenes, the electronic device comprising:

an input component configured to receive the source 3D scenes and a user's input about a first object of the source 3D scenes;

a memory storing computer-readable instructions and configured to store the source 3D scenes and the user's input about the first object of the source 3D scenes;

a processor operatively connected to the input component, the memory, and a 3D scene component, the processor being configured to execute the computer-readable instructions to instruct the 3D scene component to:

generate accurate object masks about the first object of the source 3D scenes,

train an inpainting neural radiance field (NeRF) by using the source 3D scenes and the accurate object masks, and

generate inpainted 3D scenes of the source 3D scenes by using the inpainting neural radiance field (NeRF) based on the accurate object masks,

wherein the processor is further configured to execute the computer-readable instructions to instruct the 3D scene component to train the inpainting NeRF by using at least a perceptual loss about the source 3D scenes, and

wherein the perceptual loss is used to guide the inpainting NeRF in regions identified by the accurate object masks.

15 . The electronic device of claim 14 , wherein the inpainted 3D scenes are consistent when a set of two-dimensional (2D) images of the inpainted 3D scenes corresponds to 2D projections of the inpainted 3D scenes.

16 . The electronic device of claim 14 , wherein the processor is further configured to execute the computer-readable instructions to instruct the 3D scene component to train the inpainting NeRF by using at least depth priors about the source 3D scenes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: MIRZAEI, ASHKAN; AUMENTADO-ARMSTRONG, TRISTAN TY; DERPANIS, KONSTANTINOS G.; BRUBAKER, MARCUS A.; GILITSCHENSKI, IGOR; LEVINSHTEIN, ALEKSAI
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064440/0379 →
Continuity (2)
Provisional Application 63420275 · Oct 28, 2022
Related Publication 20240153046A1 · May 9, 2024
References Cited (31)
US 20210125313A1 · Bai et al. · 2021 [cited by applicant]
US 20210375435A1 · O'Connor · 2021 [cited by examiner]
US 20220139036A1 · Bertel · 2022 [cited by examiner]
US 20220157028A1 · Huang et al. · 2022 [cited by applicant]
US 20220198731A1 · Lombardi et al. · 2022 [cited by applicant]
US 20220237750A1 · Pan et al. · 2022 [cited by applicant]
US 20220301252A1 · Wang et al. · 2022 [cited by applicant]
US 20230281913A1 · Rematas · 2023 [cited by examiner]
US 20240161388A1 · Luo · 2024 [cited by examiner]
Liu, H. K., Shen, I., & Chen, B. Y. (2022). Nerf-in: Free-form nerf inpainting with rgb-d priors. arXiv preprint arXiv:2206.04901. (Year: 2022). [cited by examiner]
Fan, Z., Wang, P., Jiang, Y., Gong, X., Xu, D., & Wang, Z. (2022). Nerf-sos: Any-view self-supervised object segmentation on complex scenes. arXiv preprint arXiv:2209.08776. (Year: 2022). [cited by examiner]
Zhi, S., Laidlow, T., Leutenegger, S., & Davison, A. J. (2021). In-place scene labelling and understanding with implicit scene representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (… [cited by examiner]
Kacper Kania, Kwang Moo Yi, Marek Kowalski, Tomasz Trzcinski, and Andrea Tagliasacchi. CoNeRF: Controllable neural radiance fields. In CVPR, 2022. (Year: 2022). [cited by examiner]
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. CLIP-NeRF: Text-and-image driven manipulation of neural radiance fields. CVPR, 2022. (Year: 2022). [cited by examiner]
V. Tschernezki, I. Laina, D. Larlus and A. Vedaldi, “Neural Feature Fusion Fields: 3D Distillation of Self-Supervised 2D Image Representations,” 2022 International Conference on 3D Vision (3DV), Prague, Czech Republic, … [cited by examiner]
Liu, H., et al., “NeRF-In: Free-Form NeRF Inpainting with RGB-D Priors” arXiv:2206.04901v1 [cs.CV] (Jun. 10, 2022), 10 pages. [cited by applicant]
International Search Report dated Jan. 30, 2024 in International Application No. PCT/KR2023/016656. [cited by applicant]
Written Opinion dated Jan. 30, 2024 in International Application No. PCT/KR2023/016656. [cited by applicant]
Roessle et al., “Dense Depth Priors for Neural Radiance Fields from Sparse Input Views”, arXiv:2112.03288v2, Apr. 7, 2022, <URL: https://arxiv.org/pdf/2112.03288v2.pdf>, pp. 1-12 (12 pages total). [cited by applicant]
Zhi et al., “In-Place Scene Labelling and Understanding with Implicit Scene Representation”, arXiv:2103.15875v2, Aug. 21, 2021, <URL: https://arxiv.org/pdf/2103.15875v2.pdf> (14 pages total). [cited by applicant]
Communication dated Sep. 1, 2025 issued by the European Patent Office in European Patent Application No. 23883094.7. [cited by applicant]
Ben Mildenhall et al., “NeRF: Representing scenes as neural radiance fields for view synthesis”, ECCV, Aug. 3, 2020, arXiv:2003.08934v2 [cs.CV], pp. 1-25 (25 pages total). [cited by applicant]
Yuying Hao et al., “Edgeflow: Achieving practical interactive segmentation with edge-guided flow”, ICCV Workshops, Oct. 26, 2021, arXiv:2109.09406v2 [cs.CV] (10 pages total). [cited by applicant]
Mathilde Caron et al., “Emerging properties in self-supervised vision transformers”, ICCV, May 24, 2021, arXiv:2104.14294v2 [cs.CV] (21 pages total). [cited by applicant]
Hao-Kang Liu et al., “NeRF-In: Free-Form NeRF Inpainting with RGB-D Priors”, arxiv.org, Cornell University Library, Jun. 10, 2022, arXiv:2206.04901v1 [cs.CV], XP091244480, pp. 1-10 (10 pages total). [cited by applicant]
Barbara Roessle et al., “Dense Depth Priors for Neural Radiance Fields from Sparse Input Views”, arXiv (Cornell University), Apr. 7, 2022, XP093165681, DOI: 10.48550/arxiv.2112.03288, https://arxiv.org/abs/2112.03288, p… [cited by applicant]
Shuaifeng Zhi et al., “In-Place Scene Labelling and Understanding with Implicit Scene Representation”, arxiv.org, Cornell University Library, Aug. 21, 2021, arXiv:2103.15875v2 [cs.CV], XP091023971 (14 pages total). [cited by applicant]
Jiaxin Li et al., “NeMI: Unifying Neural Radiance Fields with Multiplane Images for Novel View Synthesis”, Arxiv.Org, Cornell University Library, Apr. 8, 2021, arXiv:2013.14910v2 [cs.Cv], XP081930403 (16 pages total). [cited by applicant]
Sagie Benaim et al., “Volumetric Disentanglement for 3D Scene Manipulation”, arxiv.org, Cornell University Library, Jun. 6, 2022, XP091240410, pp. 1-19 (19 pages total). [cited by applicant]
Giannis Daras et al., “Solving Inverse Problems with NerfGANs”, arxiv.org, Cornell University Library, Dec. 16, 2021, arXiv:2112.09061v1 [cs.CV], XP091117575, pp. 1-16 (16 pages total). [cited by applicant]
Communication dated Jan. 26, 2026 issued by the European Patent Office in European Patent Application No. 23883094.7. [cited by applicant]