Attention-based refinement for depth completion
A processor-implemented method for attention-based depth completion includes receiving, by an artificial neural network (ANN), an input. The input includes an image and a sparse depth measurement. The ANN extracts multi-scale visual features of the input. The ANN applies a self-attention mechanism to the multi-scale visual features to generate a set of attended multi-scale visual features. The ANN generates a dense depth map based on the set of attended multi-scale visual features.
1 . A processor-implemented method performed by at least one processor, the processor-implemented method comprising:
receiving, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;
extracting, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;
applying, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and
generating, by the ANN, a dense depth map based on the set of attended multi-scale visual features.
2 . The processor-implemented method of claim 1 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.
3 . The processor-implemented method of claim 1 , wherein the processor-implemented method is performed by at least one processor of a mobile device.
4 . The processor-implemented method of claim 1 , further comprising implementing the dense depth map in an extended reality (XR) application, an autonomous driving application, a robotics application, or an image processing application.
5 . The processor-implemented method of claim 1 , further comprising processing, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.
6 . The processor-implemented method of claim 1 , in which the ANN comprises a sparse-to-dense (S2D) network.
7 . The processor-implemented method of claim 1 , in which the ANN comprises a convolutional neural network (CNN).
8 . The processor-implemented method of claim 1 , in which the image is captured by a single camera.
9 . The processor-implemented method of claim 8 , wherein the processor-implemented method is performed by at least one processor of a mobile device, wherein the single camera is included in the mobile device.
10 . An apparatus, comprising:
at least one memory; and
at least one processor coupled to the at least one memory, the at least one processor configured to:
receive, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;
extract, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;
apply, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and
generate, by the ANN, a dense depth map based on the set of attended multi-scale visual features.
11 . The apparatus of claim 10 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.
12 . The apparatus of claim 10 , in which the at least one processor is included in a mobile device.
13 . The apparatus of claim 10 , in which the at least one processor is further configured to implement the dense depth map in an extended reality (XR) application, an autonomous driving application, a robotics application, or an image processing application.
14 . The apparatus of claim 10 , in which the at least one processor is further configured to process, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.
15 . The apparatus of claim 10 , in which the ANN comprises a sparse-to-dense (S2D) network.
16 . The apparatus of claim 10 , in which the ANN comprises a convolutional neural network (CNN).
17 . The apparatus of claim 10 , in which the image is captured by a single camera.
18 . The apparatus of claim 17 , in which the at least one processor and the single camera are included in a mobile device.
19 . A non-transitory computer-readable medium having program code recorded thereon, the program code executed by at least one processor and comprising:
program code to receive, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;
program code to extract, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;
program code to apply, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and
program code to generate, by the ANN, a dense depth map based on the set of attended multi-scale visual features.
20 . The non-transitory computer-readable medium of claim 19 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.
21 . The non-transitory computer-readable medium of claim 19 , in which the program code comprises program code to process, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.
22 . The non-transitory computer-readable medium of claim 19 , in which the ANN comprises a sparse-to-dense (S2D) network.
23 . The non-transitory computer-readable medium of claim 19 , in which the image is captured by a single camera.
24 . An apparatus, comprising:
means for receiving, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;
means for extracting, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;
means for applying, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and
means for generating, by the ANN, a dense depth map based on the set of attended multi-scale visual features.
25 . The apparatus of claim 24 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.
26 . The apparatus of claim 24 , further comprising means for processing, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.
27 . The apparatus of claim 24 , in which the ANN comprises a sparse-to-dense (S2D) network.
28 . The apparatus of claim 24 , in which the image is captured by a single camera.