IP Library Granted Patent US 12711698
Granted Patent B2
US 12711698 · App. 18/712,636 · Granted Aug 18, 2026

Method, apparatus, electronic device, and storage medium for rendering three-dimensional view

Inventor: Guangwei Wang (Beijing, CN)
Assignee: Beijing Bytedance Network Technology Co., Ltd.
G06T15/506G06T5/50G06T7/90G06T15/08G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711698
App. No.
18/712,636
Granted
Aug 18, 2026
Kind
B2
Abstract

The disclosure provides a method and apparatus for rendering a three-dimensional view, an electronic device, and a storage medium. The method for rendering a three-dimensional view includes: obtaining a plurality of images to be processed and photographing attribute information of the images to be processed; processing, for each image to be processed, a current image to be processed based on a pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed; processing, for each image to be processed, photographing attribute information of the current image to be processed based on a pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed; and determining a target image corresponding to each image to be processed according to target spherical harmonic illumination and target object attribute information of each image to be processed.

Claims (72)

1 . A method for rendering a three-dimensional view, comprising:

obtaining a plurality of images to be processed and photographing attribute information of the images to be processed, wherein obtaining photographing attribute information of the images to be processed comprises: determining a camera viewing angle corresponding to the image to be processed, and using the camera viewing angle as the photographing attribute information of the image to be processed;

processing, for each image to be processed, a current image to be processed based on a pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed;

processing, for each image to be processed, photographing attribute information of the current image to be processed based on a pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed;

determining a target image corresponding to each image to be processed according to target spherical harmonic illumination and target object attribute information of each image to be processed by:

determining, for each image to be processed, target normal information corresponding to each piece of voxel position information according to voxel position information in the target object attribute information corresponding to the current image to be processed; and

rendering the target image of each image to be processed according to the target spherical harmonic illumination, target normal information, color information and material parameter information of each image to be processed; and

determining a target three-dimensional view based on a plurality of target images.

2 . The method of claim 1 , wherein processing a current image to be processed based on the pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed comprises:

obtaining the target spherical harmonic illumination output by the pre-trained target illumination estimation model corresponding to the current image to be processed by using the current image to be processed as an input parameter of the pre-trained target illumination estimation model.

3 . The method of claim 1 , wherein the target object attribute information at least comprises the voxel position information, the color information and the material parameter information, and processing photographing attribute information of the current image to be processed based on the pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed comprises:

obtaining at least voxel position information, color information and material parameter information of a target object in the current image to be processed output by the pre-trained target object attribute determination model by using the camera viewing angle in the photographing attribute information of the current image to be processed as an input parameter of the pre-trained target object attribute determination model.

4 . The method of claim 1 , wherein determining a target three-dimensional view based on a plurality of target images comprises:

obtaining a target three-dimensional view corresponding to a target object in the image to be processed by fusing the plurality of target images.

5 . The method of claim 1 , further comprising:

obtaining the pre-trained target illumination estimation model by conducting training, wherein

obtaining the pre-trained target illumination estimation model by training comprises:

determining a plurality of first images to be trained at a plurality of camera viewing angles according to at least one three-dimensional model, and obtaining a plurality of first training samples in a training sample set based on the plurality of first images to be trained and the corresponding camera viewing angles;

inputting, for each first training sample, the first image to be trained in a first current training sample into an illumination estimation model to be trained to obtain actual spherical harmonic illumination output by the illumination estimation model to be trained corresponding to the first current training sample;

conducting loss processing on the actual spherical harmonic illumination and the camera viewing angle of the first current training sample based on a first preset loss function in the illumination estimation model to be trained to correct a model parameter in the illumination estimation model to be trained according to a loss value obtained; and

using convergence of the first preset loss function as a training target to obtain the pre-trained target illumination estimation model.

6 . The method of claim 1 , further comprising: obtaining the pre-trained target object attribute determination model by training;

obtaining the pre-trained target object attribute determination model conducting by training comprises:

obtaining a plurality of second images to be trained at a plurality of camera viewing angles, and determining a plurality of second training samples based on the plurality of second images to be trained and the corresponding camera viewing angles;

using, for each second training sample, the camera viewing angle in a second current training sample as an input parameter of an object attribute determination model to be trained to obtain actual voxel position information, actual color information and actual material parameter information output by the object attribute determination model to be trained corresponding to the second current training sample;

inputting the second image to be trained in the second current training sample into the pre-trained target illumination estimation model to obtain spherical harmonic illumination to be used corresponding to the second current training sample;

correcting a model parameter in the object attribute determination model to be trained according to the second image to be trained, the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample; and

using convergence of a second preset loss function in the object attribute determination model to be trained as a training target to obtain the pre-trained target object attribute determination model.

7 . The method of claim 6 , wherein correcting a model parameter in the object attribute determination model to be trained according to the second image to be trained, the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample comprises:

rendering an actual image corresponding to the second current training sample according to the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample; and

conducting loss processing on the second image to be trained and the actual image of the second current training sample based on the second preset loss function in the object attribute determination model to be trained to correct the model parameter in the object attribute determination model to be trained according to a loss result obtained.

8 . An electronic device, comprising:

at least one processor; and

a storage apparatus configured to store at least one program;

wherein when the at least one processor executes the at least one program, the at least one processor is caused to:

obtain a plurality of images to be processed and photographing attribute information of the images to be processed, wherein the electronic device is further caused to obtain photographing attribute information of the images to be processed by: determine a camera viewing angle corresponding to the image to be processed, and using the camera viewing angle as the photographing attribute information of the image to be processed;

process, for each image to be processed, a current image to be processed based on a pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed;

process, for each image to be processed, photographing attribute information of the current image to be processed based on a pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed;

determine a target image corresponding to each image to be processed according to target spherical harmonic illumination and target object attribute information of each image to be processed by:

determining, for each image to be processed, target normal information corresponding to each piece of voxel position information according to voxel position information in the target object attribute information corresponding to the current image to be processed; and

rendering the target image of each image to be processed according to the target spherical harmonic illumination, target normal information, color information and material parameter information of each image to be processed; and

determining a target three-dimensional view based on a plurality of target images.

9 . The electronic device of claim 8 , wherein the electronic device is further caused to process a current image to be processed based on the pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed by:

obtaining the target spherical harmonic illumination output by the pre-trained target illumination estimation model corresponding to the current image to be processed by using the current image to be processed as an input parameter of the pre-trained target illumination estimation model.

10 . The electronic device of claim 8 , wherein the target object attribute information at least comprises the voxel position information, the color information and the material parameter information, and processing photographing attribute information of the current image to be processed based on the pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed comprises:

obtaining at least voxel position information, color information and material parameter information of a target object in the current image to be processed output by the pre-trained target object attribute determination model by using the camera viewing angle in the photographing attribute information of the current image to be processed as an input parameter of the pre-trained target object attribute determination model.

11 . The electronic device of claim 8 , wherein the electronic device is further caused to determining a target three-dimensional view based on a plurality of target images by:

obtaining a target three-dimensional view corresponding to a target object in the image to be processed by fusing the plurality of target images.

12 . The electronic device of claim 8 , wherein the electronic device is further caused to:

obtain the pre-trained target illumination estimation model by conducting training, wherein obtain the pre-trained target illumination estimation model by training comprises:

determine a plurality of first images to be trained at a plurality of camera viewing angles according to at least one three-dimensional model, and obtaining a plurality of first training samples in a training sample set based on the plurality of first images to be trained and the corresponding camera viewing angles;

input, for each first training sample, the first image to be trained in a first current training sample into an illumination estimation model to be trained to obtain actual spherical harmonic illumination output by the illumination estimation model to be trained corresponding to the first current training sample;

conduct loss processing on the actual spherical harmonic illumination and the camera viewing angle of the first current training sample based on a first preset loss function in the illumination estimation model to be trained to correct a model parameter in the illumination estimation model to be trained according to a loss value obtained; and

use convergence of the first preset loss function as a training target to obtain the pre-trained target illumination estimation model.

13 . The electronic device of claim 8 , wherein the electronic device is further caused to obtain the pre-trained target object attribute determination model by training;

obtaining the pre-trained target object attribute determination model conducting by training comprises:

obtaining a plurality of second images to be trained at a plurality of camera viewing angles, and determining a plurality of second training samples based on the plurality of second images to be trained and the corresponding camera viewing angles;

using, for each second training sample, the camera viewing angle in a second current training sample as an input parameter of an object attribute determination model to be trained to obtain actual voxel position information, actual color information and actual material parameter information output by the object attribute determination model to be trained corresponding to the second current training sample;

inputting the second image to be trained in the second current training sample into the pre-trained target illumination estimation model to obtain spherical harmonic illumination to be used corresponding to the second current training sample;

correcting a model parameter in the object attribute determination model to be trained according to the second image to be trained, the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample; and

using convergence of a second preset loss function in the object attribute determination model to be trained as a training target to obtain the pre-trained target object attribute determination model.

14 . The electronic device of claim 12 , wherein the electronic device is further caused to correct a model parameter in the object attribute determination model to be trained according to the second image to be trained, the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample by:

rendering an actual image corresponding to the second current training sample according to the spherical harmonic illumination to be used, the actual voxel position information, the actual color information and the actual material parameter information of the second current training sample; and

conducting loss processing on the second image to be trained and the actual image of the second current training sample based on the second preset loss function in the object attribute determination model to be trained to correct the model parameter in the object attribute determination model to be trained according to a loss result obtained.

15 . A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program when being executed by a processor, causing the processor to:

obtain a plurality of images to be processed and photographing attribute information of the images to be processed, wherein the non-transitory computer-readable storage medium is further caused to obtain photographing attribute information of the images to be processed by: determine a camera viewing angle corresponding to the image to be processed, and using the camera viewing angle as the photographing attribute information of the image to be processed;

process, for each image to be processed, a current image to be processed based on a pre-trained target illumination estimation model to obtain target spherical harmonic illumination corresponding to the current image to be processed;

process, for each image to be processed, photographing attribute information of the current image to be processed based on a pre-trained target object attribute determination model to obtain target object attribute information corresponding to the current image to be processed;

determine a target image corresponding to each image to be processed according to target spherical harmonic illumination and target object attribute information of each image to be processed by:

determining, for each image to be processed, target normal information corresponding to each piece of voxel position information according to voxel position information in the target object attribute information corresponding to the current image to be processed; and

rendering the target image of each image to be processed according to the target spherical harmonic illumination, target normal information, color information and material parameter information of each image to be processed; and

determine a target three-dimensional view based on a plurality of target images.