IP Library › Granted Patent US 12,229,880
Granted Patent B2
US 12,229,880 · App. 18/513,716 · Granted Feb 18, 2025

Method for rendering relighted 3D portrait of person and computing device for the same

Inventors: Artem Mikhailovich Sevastopolskiy (Moscow, RU); Victor Sergeevich Lempitsky (Moscow, RU)
Assignee: Samsung Electronics Co., Ltd.
G06T15/60G06F18/2148G06N3/084G06T7/194G06T15/04G06T15/10G06T17/20G06T2207/10016G06T2207/10152G06T2207/20081G06T2207/20084G06T2207/30201G06T2207/30244G06T2215/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,880
App. No.
18/513,716
Granted
Feb 18, 2025
Kind
B2
Abstract

The disclosure provides a method for generating relightable 3D portrait using a deep neural network and a computing device implementing the method. A possibility of obtaining, in real time and on computing devices having limited processing resources, realistically relighted 3D portraits having quality higher or at least comparable to quality achieved by prior art solutions, but without utilizing complex and costly equipment is provided. A method for rendering a relighted 3D portrait of a person, the method including: receiving an input defining a camera viewpoint and lighting conditions, rasterizing latent descriptors of a 3D point cloud at different resolutions based on the camera viewpoint to obtain rasterized images, wherein the 3D point cloud is generated based on a sequence of images captured by a camera with a blinking flash while moving the camera at least partly around an upper body, the sequence of images comprising a set of flash images and a set of no-flash images, processing the rasterized images with a deep neural network to predict albedo, normals, environmental shadow maps, and segmentation mask for the received camera viewpoint, and fusing the predicted albedo, normals, environmental shadow maps, and segmentation mask into the relighted 3D portrait based on the lighting conditions.

Claims (35)

1. A method for rendering a relighted 3D portrait, the method comprising:

capturing a sequence of frames by a camera with a blinking flash while moving the camera around a person;

generating a 3D point cloud based on the sequence of frames;

receiving an input defining a camera viewpoint and lighting conditions;

obtaining feature information regarding the 3D point cloud for the camera viewpoint;

obtaining, via a neural network, at least one 2D image for the camera viewpoint based on the feature information; and

fusing the at least one 2D image into the relighted 3D portrait according to the lighting conditions.

2. The method of claim 1 , wherein the at least one 2D image includes albedo map, normal map, environmental lighting map, and segmentation mask.

3. The method of claim 2 , wherein the obtaining feature information comprises obtaining rasterized images based on the feature information, and

wherein the obtaining the at least one 2D image comprises processing the rasterized images with the neural network to the albedo map, the normal map, the environmental lighting map, and the segmentation mask for the camera viewpoint.

4. The method of claim 3 , wherein the 3D point cloud is generated based on the sequence of images captured by the camera while moving the camera around an upper body of the person.

5. The method of claim 4 , wherein the sequence of images comprising a set of flash images and a set of no-flash images.

6. The method of claim 1 , wherein the generating the 3D point cloud further comprises estimating camera viewpoints with which the sequence of images is captured.

7. The method of claim 6 , wherein the received camera viewpoint and the lighting conditions differ from camera viewpoints and lighting conditions with which the sequence of images is captured.

8. The method of claim 6 , wherein the generating the 3D point cloud and the estimating camera viewpoints are performed using Structure-from-Motion (SfM), wherein the 3D point cloud comprises points corresponding to an upper body.

9. The method of claim 1 , wherein the generating the 3D point cloud further comprises:

processing each image of the sequence by segmenting a foreground; and

filtering the 3D point cloud based on the segmented foreground by a segmentation neural network to obtain the filtered 3D point cloud.

10. The method of claim 1 , wherein the obtaining feature information is performed using Z-buffering.

11. The method of claim 5 , wherein flash images of the set of flash images are alternated with no-flash images of the set of no-flash images in said sequence of images.

12. The method of claim 1 , wherein the obtaining feature information comprises rasterizing latent descriptors of the 3D point cloud at different resolutions according to the camera viewpoint, and

wherein the latent descriptor includes a multi-dimensional latent vector characterizing properties of a corresponding point in the 3D point cloud.

13. The method of claim 1 , wherein the method further comprises training the neural network, and

wherein the training the neural network comprises:

randomly sampling an image from the captured sequence of images, the camera viewpoint corresponding to the image, and

obtaining a predicted image by the neural network for the camera viewpoint, wherein the training stage is carried out iteratively.

14. The method of claim 13 , wherein the neural network is trained by backpropagation of a loss to weights of the neural network, latent descriptors, and auxiliary parameters,

wherein the loss is calculated based on one or more of: main loss, segmentation loss, room shading loss, symmetry loss, albedo color matching loss, normal loss.

15. The method of claim 14 , wherein the auxiliary parameters include one or more of room lighting color temperature, flashlight color temperature, and albedo half-texture.

16. The method of claim 15 , further comprising predicting auxiliary face meshes with corresponding texture mapping by a 3D face mesh reconstruction network for each of the images in the captured sequence,

wherein the corresponding texture mapping includes a specification of two-dimensional coordinates in the fixed, predefined texture space for every vertex of the mesh.

17. The method of claim 14 , wherein the main loss is calculated as a mismatch between predicted image and the sampled image by a combination of non-perceptual and perceptual loss functions.

18. The method of claim 14 , wherein the segmentation loss is calculated as a mismatch between a predicted segmentation mask and a segmented foreground of an image of the sequence of frames.

19. The method of claim 14 , wherein the room shading loss is calculated as a penalty for a sharpness of predicted environmental shadow maps, wherein the penalty increases as the sharpness increases.

20. A computer-readable non-transitory storage medium having stored therein instructions that, when executed, cause the electronic device to perform the method of claim 1 .

Priority Claims (2)
RU 2020137990 · Nov 19, 2020 · national
RU 2021104328 · Feb 19, 2021 · national
Continuity (3)
Continuation 17457078 · Dec 1, 2021
Continuation PCTKR2021012284 · Sep 9, 2021
Related Publication 20240096011A1 · Mar 21, 2024
References Cited (32)
US 9123116B2 · Debevec et al. · 2015 [cited by applicant]
US 11823327B2 · Sevastopolskiy et al. · 2023 [cited by applicant]
US 20130286161A1 · Lv · 2013 [cited by examiner]
US 20160314619A1 · Luo · 2016 [cited by applicant]
US 20170161938A1 · Imber et al. · 2017 [cited by applicant]
US 20180365874A1 · Hadap · 2018 [cited by applicant]
US 20190035149A1 · Chen et al. · 2019 [cited by applicant]
US 20190213778A1 · Du et al. · 2019 [cited by applicant]
US 20200201165A1 · Lock · 2020 [cited by applicant]
US 20200205723A1 · Massey et al. · 2020 [cited by applicant]
US 20200219301A1 · Yu · 2020 [cited by applicant]
US 20200258247A1 · Lasserre et al. · 2020 [cited by applicant]
EP 3382645 · 2020 [cited by applicant]
KR 101707939 · 2017 [cited by applicant]
KR 1020180004635 · 2018 [cited by applicant]
KR 1020200057077 · 2020 [cited by applicant]
KR 1020200104068 · 2020 [cited by applicant]
KR 1020200129657 · 2020 [cited by applicant]
RU 2721180 · 2020 [cited by applicant]
WO 2020037676 · 2020 [cited by applicant]
Nam, Giljoo, et al. “Practical svbrdf acquisition of 3d objects with unstructured flash photography.” ACM Transactions on Graphics (TOG) 37.6 (2018): 1-12. (Year: 2018). [cited by examiner]
Ha, Hyunho, et al. “Progressive acquisition of svbrdf and shape in motion.” Computer Graphics Forum. vol. 39. No. 6. 2020. First published Aug. 8, 2020. (Year: 2020). [cited by examiner]
U.S. Appl. No. 17/457,078, filed Dec. 1, 2021; Sevastopolskiy et al. [cited by applicant]
Zhou et al., “Deep single-image portrait relighting”, In Proceedings of the IEEE International Conference on Computer Vision, 2019, 9 pages. [cited by applicant]
Zhang et al., “Neural Light Transport for Relighting and View Synthesis”, In Proceedings of ACM Transactions on Graphics, Aug. 9, 2020, 16 pages. [cited by applicant]
Search Report for RU2021104328/28 dated Feb. 19, 2021, 4 pages. [cited by applicant]
Decision to Grant for RU2021104328/28 dated Sep. 20, 2021, 18 pages, with its English Translation. [cited by applicant]
International Search Report and Written Opinion dated Dec. 24, 2021 in corresponding International Application No. PCT/KR2021/012284. [cited by applicant]
Kara-Ali Aliev et al., “Neural point-based graphics”, In: Computer Vision, ECCV 2020, 16th European Conference, Aug. 23-28, 2020, Springer International Publishing. [cited by applicant]
M. Sabbadin et al., “High Dynamic Range Point Clouds for Real-Time Relighting” Pacific Graphics 2019, vol. 38 (2019), No. 7, Nov. 14, 2019, pp. 513-525. [cited by applicant]
Murmariri, Lukas, et al. A Datasetof Multi-Illumination Images in the Wild. 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE. (Year: 2019). [cited by applicant]
Yu, Ye, et al. Self-supervised outdoor scene relighting. Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part )(XII 16. Springer International Publishing, 2020. (Year: 20… [cited by applicant]