IP Library › Granted Patent US 12,705,820
Granted Patent B2
US 12,705,820 · App. 18/332,273 · Granted Aug 11, 2026

Method and appratus with neural rendering based on view augmentation

Inventors: Young Chun Ahn (Suwon-si, KR); Nahyup Kang (Suwon-si, KR); Seokhwan Jang (Suwon-si, KR); Jiyeon Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T15/08G06T3/18G06T7/194G06T7/50G06T15/06G06T2207/20081G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,820
App. No.
18/332,273
Filed
Jun 9, 2023
Granted
Aug 11, 2026
Kind
B2
Examiner
TAHA, AHMED
Art Unit
2613
USPC
345/419
Abstract

A method and apparatus for neural rendering based on view augmentation are provided. A method of training a neural scene representation (NSR) model includes: receiving original training images of a target scene, the original training images respectively corresponding to base views of the target scene; generating augmented images of the target scene by warping the original training images, the augmented images respectively corresponding to new views of the target scene; performing background-foreground segmentation on the original training images and the augmented images to generate segmentation masks; and training a neural scene representation (NSR) model to be configured for volume rendering of the target scene by using the original training images, the augmented images, and the segmentation masks.

Claims (75)

1 . A method of training a neural scene representation (NSR) model, the method comprising:

receiving original training images of a target scene, the original training images respectively corresponding to base views of the target scene;

generating augmented images of the target scene by warping the original training images, the augmented images respectively corresponding to new views of the target scene, wherein the generating the augmented images comprises:

determining a transformation function for transforming a camera pose of a first of the base views of a first of the original training images into a camera pose of a first of the new views of a first of the augmented images; and

generating a first augmented image by warping a first original training image using an intrinsic camera parameter of the first original training image, an original depth map corresponding to the first original training image, and the transformation function;

performing background-foreground segmentation on the original training images and the augmented images to generate segmentation masks; and

training the NSR model to be configured for volume rendering of the target scene by using the original training images, the augmented images, and the segmentation masks,

wherein the training of the NSR model comprises:

performing primary training of the NSR model using the original training images, the augmented images, the segmentation masks, and a first loss function; and

performing secondary training of the NSR model using the original training images and a second loss function,

wherein the first loss function is based on a pixel error between (i) an actual pixel value from the original training images and the augmented images and (ii) a pixel value estimated by the NSR model, and

wherein the second loss function is based on a pixel error between the original training images and a synthesized image estimated by the NSR model, semantic consistency between the original training images and the synthesized image, and uncertainty of transmittance based on a ray.

2 . The method of claim 1 , wherein the performing of the primary training comprises:

selecting a first sample image from the original training images and the augmented images;

determining a first query output of the NSR model according to a first query input defining a first ray;

determining a target area to which the first ray belongs from among a foreground area of the first sample image and a background area of the first sample image, based on the segmentation masks; and

determining a loss value of the first loss function based on an actual pixel value of a first pixel of the target area specified by the first ray and an estimated pixel value according to the first query input.

3 . The method of claim 2 , wherein the determining of the target area comprises:

dividing the foreground area of the first sample image and the background area of the first sample image by applying a first of the segmentation masks corresponding to the first sample image to the first sample image;

when the first ray indicates the foreground area of the first sample image, determining the foreground area of the first sample image to be the target area; and

when the first ray indicates the background area of the first sample image, determining the background area of the first sample image to be the target area.

4 . The method of claim 1 , wherein the performing of the secondary training comprises:

generating a first synthesized image according to a first ray set of a first of the original training images by using the NSR model;

estimating first semantic characteristics of patches of the first original training image and second semantic characteristics of patches of the first synthesized image;

determining semantic consistency between the first original training image and the first synthesized image based on a difference between the first semantic characteristics and the second semantic characteristics; and

determining a loss value of the second loss function based on the determined semantic consistency.

5 . The method of claim 1 , wherein the performing of the secondary training comprises:

based on products of volume densities and transmittances of sample points of rays of the first original training image among the original training images, determining weights of the sample points; and

determining a loss value of the second loss function based on the weights of the rays.

6 . The method of claim 1 , wherein a number of original training images is limited to a predetermined number.

7 . An apparatus comprising:

one or more processors; and

a memory storing instructions configured to cause the one or more processors to:

receive original training images of a target scene, the original training images respectively corresponding to base views of the target scene,

generate augmented images of the target scene by warping the original training images, the augmented images respectively corresponding to new views of the target scene,

determine foreground-background segmentation masks of the original training images and the augmented images by performing foreground-background segmentation on the original training images and the augmented images, and

train a neural scene representation (NSR) model to be configured for volume rendering of the target scene by using the original training images, the augmented images, and the foreground-background segmentation masks,

wherein, to generate the augmented images, the instructions stored in the memory are further configured to cause the one or more processors to:

determine a transformation function for transforming a base camera pose of a first original training image of the original training images into a new camera pose of a first augmented image of the augmented images, and

generate the first augmented image by warping the first original training image using a camera intrinsic parameter of the first original training image, an original depth map corresponding to the first original training image, and the transformation function,

wherein, to train the NSR model, the instructions are further configured to cause the one or more processors to:

perform primary training of the NSR model using the original training images, the augmented images, the foreground-background segmentation masks, and a first loss function, and

perform secondary training of the NSR model using the original training images and a second loss function,

wherein the first loss function is based on a pixel error between an actual pixel value of the original training images and the augmented images and a pixel value estimated by the NSR model, and

wherein the second loss function is based on a pixel error between the original training images and a synthesized image estimated by the NSR model, semantic consistency between the original training images and the synthesized image, and uncertainty of transmittance based on a ray.

8 . The apparatus of claim 7 , wherein the original training images are respectively associated with base camera poses, the augmented images are respectively associated with new camera poses, and wherein the training of the NSR model also uses the base camera poses and the new camera poses.

9 . The apparatus of claim 7 , wherein, to perform the primary training, the instructions are further configured to cause the one or more processors to:

select a first sample image from the original training images and the augmented images, determine a first query output of the NSR model according to a first query input indicating a first ray,

determine a target area to which the first ray belongs among a foreground area of the first sample image and a background area of the first sample image, based on the foreground-background segmentation masks, and

determine a loss value of the first loss function based on an actual pixel value of a first pixel of the target area specified by the first ray and an estimated pixel value according to the first query output.

10 . The apparatus of claim 9 , wherein, to determine the target area, the instructions are further configured to cause the one or more processors to:

divide the foreground area of the first sample image and the background area of the first sample image by applying a first of the foreground-background segmentation masks corresponding to the first sample image to the first sample image,

when the first ray indicates the foreground area of the first sample image, determine the foreground area of the first sample image to be the target area, and

when the first ray indicates the background area of the first sample image, determine the background area of the first sample image to be the target area.

11 . The apparatus of claim 7 , wherein, to perform the secondary training, the instructions are further configured to cause the one or more processors to:

generate a first synthesized image according to a first ray set of a first of the original training images by using the NSR model,

estimate first semantic characteristics of multi-level patches of the first original training image and second semantic characteristics of multi-level patches of the first synthesized image, determine semantic consistency between the first original training image and the first synthesized image based on a difference between the first semantic characteristics and the second semantic characteristics, and

determine a loss value of the second loss function based on the determined semantic consistency.

12 . An electronic device comprising:

a camera generating original training images of respective original camera poses of a target scene, the original training images respectively corresponding to base views of the target scene; and

one or more processors;

a memory storing instructions configured to cause the one or more processors to:

generate augmented images of respective augmentation-image camera

poses for the target scene by warping the original training images, the augmented images respectively corresponding to new views of the target scene,

determine segmentation masks for dividing areas of the original training

images and the augmented images by performing segmentation on the original training images and the augmented images, and

train a neural scene representation (NSR) model used for volume rendering for the target scene by using the original training images and their respective original camera poses, the augmented images and their respective augmentation-image camera poses, and the segmentation masks,

wherein, to generate the augmented images, the instructions are further configured to cause the one or more processors to:

determine a transformation function for transform a first original camera pose of a first of the original training images into a first of the augmentation-image camera poses of a first of the augmented images, and

generate the first of the augmented images by warping a first original training image using a camera intrinsic parameter of the first of the original training images, an original depth map corresponding to the first of the original training images, and the transformation function,

wherein, to train the NSR model, the instructions are further configured to cause the one or more processors to:

perform primary training of the NSR model using the original training images, the augmented images, the segmentation masks, and a first loss function, and

perform secondary training of the NSR model using the original training images and a second loss function,

wherein the first loss function is based on a pixel error between an actual pixel value of the original training images and the augmented images and a pixel value estimated by the NSR model, and

wherein the second loss function is based on a pixel error between the original training images and a synthesized image estimated by the NSR model, semantic consistency between the original training images and the synthesized image, and uncertainty of transmittance based on a ray.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2023
From: AHN, YOUNG CHUN; KANG, NAHYUP; JANG, SEOKHWAN; KIM, JIYEON
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 063910/0500 →
Priority Claims (2)
KR 10-2022-0128898 · Oct 7, 2022 · national
KR 10-2022-0178564 · Dec 19, 2022 · national
Continuity (1)
Related Publication 20240135632A1 · Apr 25, 2024
References Cited (25)
US 11164394B2 · Gausebeck · 2021 [cited by applicant]
US 20210279952A1 · Chen et al. · 2021 [cited by applicant]
US 20220101047A1 · Puri et al. · 2022 [cited by applicant]
US 20220130111A1 · Martin Brualla et al. · 2022 [cited by applicant]
US 20220156886A1 · Petrangeli · 2022 [cited by examiner]
KR 102395123B1 · 2022 [cited by applicant]
(C. Xie, K. Park, R. Martin-Brualla, M. Brown “FiG-NeRF: Figure-Ground Neural Radiance Fields for 3D Object Category Modelling,” arXiv:2104.08418 [cs.CV], published Apr. 2021, 12 pages) (Year: 2021). [cited by examiner]
Kisantal et al (M. Kisantal, Z. Wojna, J. Murawski, J. Naruniec, K. Cho “Augmentation for small object detection” arXiv: 1902.07296v1 [cs.CV], published Feb. 2019, 15 pages). (Year: 2019). [cited by examiner]
(E. Guo, Z. Chen, Y. Zhou, D. Wu “Unsupervised Learning of Depth and Camera Pose with Feature Map Warping”, published Jan. 30, 2021, 15 pages) (Year: 2021). [cited by examiner]
(A. Gunawan, S. Rizky, H. Madjid “CISRNet: Compressed Image Super-Resolution Network” arXiv:2201.06045 [cs.CV], published Jan. 2022, 10 pages). (Year: 2022). [cited by examiner]
(N. Sunderhauf , J, Abou-Chakara, D. Miller “Density-aware NeRF Ensembles: Quantifying Predictive Uncertainty in Neural Radiance Fields” arXiv:2209.08718 [cs.CV], published Sep. 19, 2022, 7 pages). (Year: 2022). [cited by examiner]
Yu, Alex, et al., “Pixelnerf: Neural Radiance Fields from One or Few Images,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, (p. 4578-4587). [cited by applicant]
Jain, Ajay, et al., “Putting Nerf on a Diet: Semantically Consistent Few-Shot View Synthesis,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, (p. 5885-5894). [cited by applicant]
Mildenhall, Ben, et al., “Nerf: Representing Scenes as Neural Radiance Fields for View Synthesis,” Communications of the ACM, 2021, (17 Pages in English). [cited by applicant]
Radford, Alec, et al., “Learning Transferable Visual Models from Natural Language Supervision,” International conference on machine learning, 2021, (47 Pages in English). [cited by applicant]
Caron, Mathilde, et al., “Emerging Properties in Self-Supervised Vision Transformers,” Proceedings of the IEEE/CVF international conference on computer vision, 2021, (p. 9650-9660). [cited by applicant]
Dosovitskiy, Alexey, et al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv:2010.11929v2, Jun. 3, 2021, (22 Pages in English). [cited by applicant]
Kim, Mijeong, et al., “Infonerf: Ray Entropy Minimization for Few-Shot Neural Volume Rendering,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, (p. 12912-12921). [cited by applicant]
Niemeyer, Michael, et al., “Regnerf: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, (p. 5480-5490). [cited by applicant]
Deng, Kangle, et al., “Depth-Supervised Nerf: Fewer Views and Faster Training for Free,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, (p. 12882-12891). [cited by applicant]
Xu, Dejia, et al., “Sinnerf: Training Neural Radiance Fields on Complex Scenes from a Single Image,” Computer Vision-ECCV, Oct. 23, 2022, (18 Pages in English). [cited by applicant]
Ahn, Young Chun, et al., “Panerf: Pseudo-View Augmentation for Improved Neural Radiance Fields Based on Few-Shot Inputs,” arXiv:2211.12758v1, Nov. 23, 2022, (10 Pages in English). [cited by applicant]
Zhang, Cha, et al., “A survey on image-based rendering-representation, sampling and compression.” Signal Processing: Image Communication 19.1 (2004): 1-28. [cited by applicant]
Kundu, Abhijit, et al. “Panoptic Neural Fields: a Semantic Object-Aware Neural Scene Representation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, (11 pages). [cited by applicant]
Extended European search report issued on Apr. 10, 2024, in counterpart European Patent Application No. 23187904.0 (7 pages). [cited by applicant]