IP Library › Granted Patent US 12,217,443
Granted Patent B2
US 12,217,443 · App. 17/714,654 · Granted Feb 4, 2025

Depth image generation method, apparatus, and storage medium and electronic device

Inventors: Runze Zhang (Shenzhen, CN); Hongwei Yi (Shenzhen, CN); Ying Chen (Shenzhen, CN); Shang Xu (Shenzhen, CN); Yu Wing Tai (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LTD
G06T7/55G06F18/253G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,443
App. No.
17/714,654
Granted
Feb 4, 2025
Kind
B2
Abstract

A depth image generation method, apparatus, and storage medium and electronic device. The method includes: acquiring a plurality of target images; performing multi-stage convolution processing on the plurality of target images through a plurality of convolutional layers in a convolution model to obtain feature map sets respectively outputted by the plurality of convolutional layers; performing view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set; and performing fusion processing on the plurality of obtained aggregated features to obtain a depth image. The plurality of acquired target images are obtained by photographing the target object from different views respectively, so that the plurality of obtained target images include information from different angles, which enriches information content of the acquired target images.

Claims (90)

1. A depth image generation method, performed by a computer device, comprising:

acquiring a plurality of target images;

performing multi-stage convolution processing on the plurality of target images through a plurality of convolutional layers in a convolution model to obtain feature map sets respectively outputted by the plurality of convolutional layers, each feature map set comprising feature maps corresponding to the plurality of target images;

performing view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set; and

performing fusion processing on the plurality of obtained aggregated features to obtain a depth image,

wherein the performing view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set comprises:

regarding any one of the target images as a reference image, and regarding other target images in the plurality of target images as a first image;

performing the following processing on a feature map set:

determining, in the feature map set, a reference feature map corresponding to the reference image and a first feature map corresponding to the first image;

performing, according to a difference between photographing views of the first image and the reference image, view conversion on the first feature map to obtain a second feature map after conversion; and

performing fusion processing on the reference feature map and the second feature map to obtain the aggregated feature.

2. The depth image generation method according to claim 1 , wherein the performing multi-stage convolution processing on the plurality of target images through a plurality of convolutional layers in a convolution model to obtain feature map sets respectively outputted by the plurality of convolutional layers comprises:

performing convolution processing on the plurality of target images through a first convolutional layer in first block convolution model to obtain a feature map set outputted by the first convolutional layer; and

performing, through a next convolutional layer in the convolution model, convolution processing on each feature map in the feature map set outputted by the previous convolutional layer to obtain a feature map set outputted by the next convolutional layer, until feature map sets outputted respectively by the plurality of convolutional layers are obtained.

3. The depth image generation method according to claim 1 , wherein the performing, according to a difference between photographing views of the first image and the reference image, view conversion on the first feature map to obtain a second feature map after conversion comprises:

acquiring a first photographing parameter corresponding to the first image and a reference photographing parameter corresponding to the reference image;

determining a plurality of depth values corresponding to the convolutional layer that outputs the feature map set;

determining, according to a difference between the first photographing parameter and the second photographing parameter as well as the plurality of depth values, a plurality of view conversion matrices corresponding to the plurality of depth values; and

performing, according to the plurality of view conversion matrices, view conversion on the first feature map respectively to obtain a plurality of second feature maps after conversion.

4. The depth image generation method according to claim 3 , wherein the determining a plurality of depth values corresponding to the convolutional layer that outputs the feature map set comprises:

determining a depth layer number corresponding to the convolutional layer that outputs the feature map set; and

dividing a preset depth range according to the depth layer number to obtain the plurality of depth values.

5. The depth image generation method according to claim 3 , wherein the first image comprises a plurality of first images, and the performing fusion processing on the reference feature map and the second feature map to obtain the aggregated feature comprises:

performing fusion processing on a first quantity of reference feature maps to obtain a reference feature volume corresponding to the reference image, the first quantity being equal to the quantity of the plurality of depth values;

performing, for each first image, fusion processing on a plurality of second feature maps converted from the first feature maps corresponding to the first image to obtain first feature volumes, and determining differences between the first feature volumes and the reference feature volume as second feature volumes; and

performing fusion processing on the plurality of determined second feature volumes to obtain the aggregated feature.

6. The depth image generation method according to claim 5 , wherein the performing fusion processing on the plurality of determined second feature volumes to obtain the aggregated feature comprises:

acquiring a weight matrix corresponding to the convolutional layer that outputs the feature map set, the weight matrix comprising a weight corresponding to each pixel position in the feature map outputted by the convolutional layer; and

performing weighted fusion processing on the plurality of second feature volumes according to the weight matrix to obtain the aggregated feature.

7. The depth image generation method according to claim 1 , wherein scales of the feature maps outputted by the plurality of convolutional layers decrease sequentially; and the performing fusion processing on the plurality of obtained aggregated features to obtain a depth image comprises:

regarding an aggregated feature of the maximum scale in the plurality of aggregated features as a first aggregated feature, and regarding a plurality of other aggregated features in the plurality of aggregated features as second aggregated features;

performing multi-stage convolution processing on the first aggregated feature to obtain a plurality of third aggregated features, scales of the plurality of third aggregated features one-to-one corresponding to scales of the plurality of second aggregated features;

performing fusion processing on a second aggregated feature of a first scale and a third aggregated feature of the first scale, and performing deconvolution processing on the fused feature to obtain a fourth aggregated feature of a second scale, the first scale being a minimum scale of the plurality of second aggregated features, and the second scale being a previous-level scale of the first scale;

performing fusion processing continuously on the currently obtained fourth aggregated feature as well as the second aggregated feature and the third aggregated feature that are of scales equal to that of the fourth aggregated feature, and performing deconvolution processing on the fused feature to obtain a fourth aggregated feature of a previous-level scale, until a fourth aggregated feature with a scale equal to the scale of the first aggregated feature is obtained;

performing fusion processing on the fourth aggregated feature of the scale equal to that of the first aggregated feature and the first aggregated feature to obtain a fifth aggregated feature; and

performing, according to a probability map corresponding to the first aggregated feature, convolution processing on the fifth aggregated feature to obtain the depth image.

8. The depth image generation method according to claim 7 , wherein the continuously performing fusion processing on the currently obtained fourth aggregated feature as well as the second aggregated feature and the third aggregated feature that are of scales equal to that of the fourth aggregated feature, and performing deconvolution processing on the fused feature to obtain a fourth aggregated feature of a previous-level scale comprises:

performing fusion processing continuously on the currently obtained fourth aggregated feature as well as the second aggregated feature and the third aggregated feature that are of scales equal to that of the fourth aggregated feature, and a probability map of the second aggregated feature, and performing deconvolution processing on the fused feature to obtain the fourth aggregated feature of the previous-level scale.

9. The depth image generation method according to claim 1 , wherein the acquiring a plurality of target images comprises:

photographing a target object from a plurality of different views to obtain the plurality of target images, or;

photographing the target object from a plurality of different views to obtain a plurality of original images, and

performing scale adjustment on the plurality of original images to obtain the plurality of target images, the plurality of target images are of equal scales.

10. The depth image generation method according to claim 9 , wherein the performing scale adjustment on the plurality of original images to obtain the plurality of target images comprises:

performing a plurality of rounds of scale adjustment on the plurality of original images to obtain a plurality of target image sets, each target image set comprising a plurality of target images of a same scale, and target images in different target image sets being of different scales; and

the method further comprises performing fusion processing on depth images corresponding to the plurality of target image sets to obtain a fused depth image.

11. The depth image generation method according to claim 10 , wherein the performing fusion processing on depth images corresponding to the plurality of target image sets to obtain a fused depth image comprises:

replacing, starting from a depth image of a minimum scale, a depth value of a second pixel corresponding to a first pixel in a depth image of a previous scale with a depth value of the first pixel meeting a preset condition in a current depth image, until a depth value in a depth image of a maximum scale is replaced, for obtaining a depth image after replacing the depth value of the depth image of the maximum scale.

12. The depth image generation method according to claim 11 , further comprising:

mapping, for a first depth image and a second depth image of adjacent scales, any second pixel in the second depth image into the first depth image according to a pixel mapping relationship between the first depth image and the second depth image to obtain a first pixel, a scale of the second depth image being greater than a scale of the first depth image;

inversely mapping, according to the pixel mapping relationship, the first pixel into the second depth image to obtain a third pixel; and

determining, in response to a distance between the first pixel and the third pixel being less than a first preset threshold, that the first pixel corresponds to the second pixel.

13. A depth image generation apparatus, the apparatus comprising:

at least one memory configured to store computer program code; and

at least one processor configured to read the computer program code and operate as instructed by the computer program code, the computer program code comprising:

image acquisition code configured to cause the at least one processor to acquire a plurality of target images, the plurality of target images being obtained respectively by photographing a target object from different views;

convolution processing code configured to cause the at least one processor to perform multi-stage convolution processing on the plurality of target images through a plurality of convolutional layers in a convolution model to obtain feature map sets respectively outputted by the plurality of convolutional layers, each feature map set comprising feature maps corresponding to the plurality of target images;

view aggregation code configured to cause the at least one processor to perform view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set; and

feature fusion code configured to cause the at least one processor to perform fusion processing on the plurality of obtained aggregated features to obtain a depth image,

wherein the view aggregation code further comprises:

image determining code configured to cause the at least one processor to regard one of the target images as a reference image, and regard other target images in the plurality of target images as a first image; and

wherein the view aggregation code is further configured to cause the at least one processor to perform the following processing on any feature map set:

determine, in the feature map set, a reference feature map corresponding to the reference image and a first feature map corresponding to the first image;

perform, according to a difference between photographing views of the first image and the reference image, view conversion on the first feature map to obtain a second feature map after conversion; and

perform fusion processing on the reference feature map and the second feature map to obtain the aggregated feature.

14. The depth image generation apparatus according to claim 13 , wherein the convolution processing code is further configured to cause the at least one processor to:

perform convolution processing on the plurality of target images through a first convolutional layer in first block convolution model to obtain a feature map set outputted by the first convolutional layer; and

perform, through a next convolutional layer in the convolution model, convolution processing on each feature map in the feature map set outputted by the previous convolutional layer to obtain a feature map set outputted by the next convolutional layer, until feature map sets outputted respectively by the plurality of convolutional layers are obtained.

15. The depth image generation apparatus according to claim 13 ,

wherein the view aggregation code is further configured to cause the at least one processor to:

acquire a first photographing parameter corresponding to the first image and a reference photographing parameter corresponding to the reference image;

determine a plurality of depth values corresponding to the convolutional layer that outputs the feature map set;

determine, according to a difference between the first photographing parameter and the second photographing parameter as well as the plurality of depth values, a plurality of view conversion matrices corresponding to the plurality of depth values; and

perform, according to the plurality of view conversion matrices, view conversion on the first feature map respectively to obtain a plurality of second feature maps after conversion.

16. The depth image generation apparatus according to claim 15 , wherein the view aggregation code is further configured to cause the at least one processor to:

determine a depth layer number corresponding to the convolutional layer that outputs the feature map set; and

divide a preset depth range according to the depth layer number to obtain the plurality of depth values.

17. The depth image generation apparatus according to claim 15 ,

wherein the view aggregation code is further configured to cause the at least one processor to:

perform fusion processing on a first quantity of reference feature maps to obtain a reference feature volume corresponding to the reference image, the first quantity being equal to the quantity of the plurality of depth values;

perform, for each first image, fusion processing on a plurality of second feature maps converted from the first feature maps corresponding to the first image to obtain first feature volumes, and determining differences between the first feature volumes and the reference feature volume as second feature volumes; and

perform fusion processing on the plurality of determined second feature volumes to obtain the aggregated feature.

18. A non-transitory computer-readable storage medium, storing computer program code that when executed by at least one processor causes the at least one processor to:

acquire a plurality of target images, the plurality of target images being obtained respectively by photographing a target object from different views;

perform multi-stage convolution processing on the plurality of target images through a plurality of convolutional layers in a convolution model to obtain feature map sets respectively outputted by the plurality of convolutional layers, each feature map set comprising feature maps corresponding to the plurality of target images;

perform view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set; and

perform fusion processing on the plurality of obtained aggregated features to obtain a depth image,

wherein the perform view aggregation on a plurality of feature maps in each feature map set respectively to obtain an aggregated feature corresponding to each feature map set comprises:

regarding any one of the target images as a reference image, and regarding other target images in the plurality of target images as a first image;

performing the following processing on a feature map set:

determining, in the feature map set, a reference feature map corresponding to the reference image and a first feature map corresponding to the first image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: ZHANG, RUNZE; YI, HONGWEI; CHEN, YING; XU, SHANG; TAI, YU WING
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LTD
Reel/Frame 059520/0230 →
Priority Claims (1)
CN 202010119713.5 · Feb 26, 2020 · national
Continuity (2)
Continuation PCTCN2020127891 · Nov 10, 2020
Related Publication 20220230338A1 · Jul 21, 2022
References Cited (16)
US 10482334B1 · Chen et al. · 2019 [cited by applicant]
US 11017586B2 · Long · 2021 [cited by examiner]
US 20200086879A1 · Lakshmi Narayanan · 2020 [cited by examiner]
CN 109461180A · 2019 [cited by examiner]
CN 109905691A · 2019 [cited by applicant]
CN 110021069A · 2019 [cited by applicant]
CN 110378943A · 2019 [cited by applicant]
CN 110457515A · 2019 [cited by applicant]
CN 110543581A · 2019 [cited by applicant]
CN 111340866A · 2020 [cited by applicant]
Hongwei Yi, et al., “Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation”, arXiv:1912.03001v2 [cs.CV], Jul. 21, 2020, (1-16 pages total). [cited by applicant]
Yao Yao, et al., “MVSNet: Depth Inference for Unstructured Multi-view Stereo”, arXiv:1804.02505v2 [cs.CV], Jul. 17, 2018, (1-17 pages total). [cited by applicant]
Translation of the Written Opinion issued Feb. 10, 2021 in International Application No. PCT/CN2020/127891. [cited by applicant]
Po-Han Huang et al., “DeepMVS: Learning Multi-view Stereopsis,” Apr. 2, 2018, arXiv: 1804.00650v1 (10 pages). [cited by applicant]
International Search Report for PCT/CN2020/127891, dated Feb. 10, 2021. [cited by applicant]
Written Opinion for PCT/CN2020/127891, dated Feb. 10, 2021. [cited by applicant]