IP Library › Granted Patent US 12,737,900
Granted Patent B2
US 12,737,900 · App. 18/448,845 · Granted Sep 15, 2026

Attention-based refinement for depth completion

Inventors: Yunxiao Shi (San Diego, CA); Hong Cai (San Diego, CA); Fatih Murat Porikli (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T7/50G06T2207/10024G06T2207/10028G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,900
App. No.
18/448,845
Filed
Aug 11, 2023
Granted
Sep 15, 2026
Kind
B2
Art Unit
2677
USPC
382/106
Abstract

A processor-implemented method for attention-based depth completion includes receiving, by an artificial neural network (ANN), an input. The input includes an image and a sparse depth measurement. The ANN extracts multi-scale visual features of the input. The ANN applies a self-attention mechanism to the multi-scale visual features to generate a set of attended multi-scale visual features. The ANN generates a dense depth map based on the set of attended multi-scale visual features.

Claims (46)

1 . A processor-implemented method performed by at least one processor, the processor-implemented method comprising:

receiving, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;

extracting, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;

applying, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and

generating, by the ANN, a dense depth map based on the set of attended multi-scale visual features.

2 . The processor-implemented method of claim 1 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.

3 . The processor-implemented method of claim 1 , wherein the processor-implemented method is performed by at least one processor of a mobile device.

4 . The processor-implemented method of claim 1 , further comprising implementing the dense depth map in an extended reality (XR) application, an autonomous driving application, a robotics application, or an image processing application.

5 . The processor-implemented method of claim 1 , further comprising processing, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.

6 . The processor-implemented method of claim 1 , in which the ANN comprises a sparse-to-dense (S2D) network.

7 . The processor-implemented method of claim 1 , in which the ANN comprises a convolutional neural network (CNN).

8 . The processor-implemented method of claim 1 , in which the image is captured by a single camera.

9 . The processor-implemented method of claim 8 , wherein the processor-implemented method is performed by at least one processor of a mobile device, wherein the single camera is included in the mobile device.

10 . An apparatus, comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

receive, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;

extract, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;

apply, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and

generate, by the ANN, a dense depth map based on the set of attended multi-scale visual features.

11 . The apparatus of claim 10 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.

12 . The apparatus of claim 10 , in which the at least one processor is included in a mobile device.

13 . The apparatus of claim 10 , in which the at least one processor is further configured to implement the dense depth map in an extended reality (XR) application, an autonomous driving application, a robotics application, or an image processing application.

14 . The apparatus of claim 10 , in which the at least one processor is further configured to process, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.

15 . The apparatus of claim 10 , in which the ANN comprises a sparse-to-dense (S2D) network.

16 . The apparatus of claim 10 , in which the ANN comprises a convolutional neural network (CNN).

17 . The apparatus of claim 10 , in which the image is captured by a single camera.

18 . The apparatus of claim 17 , in which the at least one processor and the single camera are included in a mobile device.

19 . A non-transitory computer-readable medium having program code recorded thereon, the program code executed by at least one processor and comprising:

program code to receive, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;

program code to extract, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;

program code to apply, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and

program code to generate, by the ANN, a dense depth map based on the set of attended multi-scale visual features.

20 . The non-transitory computer-readable medium of claim 19 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.

21 . The non-transitory computer-readable medium of claim 19 , in which the program code comprises program code to process, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.

22 . The non-transitory computer-readable medium of claim 19 , in which the ANN comprises a sparse-to-dense (S2D) network.

23 . The non-transitory computer-readable medium of claim 19 , in which the image is captured by a single camera.

24 . An apparatus, comprising:

means for receiving, by an artificial neural network (ANN), an input comprising an image and a sparse depth measurement;

means for extracting, by the ANN, visual features of the input at multiple different scales to produce multi-scale visual features of the input, the multi-scale visual features comprising visual features at each of the multiple different scales;

means for applying, by the ANN, a self-attention mechanism to a subset of the multi-scale visual features to generate a set of attended multi-scale visual features, the subset comprising visual features at fewer than all of the multiple different scales; and

means for generating, by the ANN, a dense depth map based on the set of attended multi-scale visual features.

25 . The apparatus of claim 24 , in which the sparse depth measurement comprises a light detection and ranging (LiDAR) measurement, a red green blue depth (RGBD) measurement, or a time-of-flight (ToF) measurement.

26 . The apparatus of claim 24 , further comprising means for processing, by the ANN, the multi-scale visual features by applying a depth-separable convolution to the subset of the multi-scale visual features.

27 . The apparatus of claim 24 , in which the ANN comprises a sparse-to-dense (S2D) network.

28 . The apparatus of claim 24 , in which the image is captured by a single camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: SHI, YUNXIAO; CAI, HONG; PORIKLI, FATIH MURAT
To: QUALCOMM INCORPORATED
Reel/Frame 065093/0972 →
Continuity (1)
Related Publication 20250054168A1 · Feb 13, 2025
References Cited (28)
US 10832487B1 · Ulbricht · 2020 [cited by examiner]
US 11494937B2 · Urtasun · 2022 [cited by examiner]
US 11579629B2 · Wu · 2023 [cited by examiner]
US 12333730B2 · Xiong · 2025 [cited by examiner]
US 20190289281A1 · Badrinarayanan · 2019 [cited by examiner]
US 20220026918A1 · Guizilini · 2022 [cited by examiner]
US 20220067950A1 · Lv · 2022 [cited by examiner]
US 20230092248A1 · Xiong · 2023 [cited by examiner]
US 20230230269A1 · Yoo · 2023 [cited by examiner]
US 20240127596A1 · Singh · 2024 [cited by examiner]
US 20240161319A1 · Ingle · 2024 [cited by examiner]
CN 108830192A · 2018 [cited by examiner]
CN 110097589A · 2019 [cited by examiner]
CN 114004754B · 2022 [cited by examiner]
CN 116245930A · 2023 [cited by examiner]
CN 114743079B · 2025 [cited by examiner]
GB 2576548B · 2020 [cited by examiner]
WO WO2021013334A1 · 2021 [cited by examiner]
WO WO2022103171A1 · 2022 [cited by examiner]
WO WO2022265347A1 · 2022 [cited by examiner]
SDformer: Efficient End-to-End Transformer for Depth Completion, Jian Qian et al., IEEE, 2022, pp. 56-61 (Year: 200). [cited by examiner]
Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian Loss, Lam Huynh et al., IEEE, 2021, pp. 228-238 (Year: 2021). [cited by examiner]
Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks, Md Fahim Faysal Khan et al., ACM, 2022, pp. 3559-3568 (Year: 2022). [cited by examiner]
HMS-Net: Hierarchical Multi-Scale Sparsity-Invariant Network for Sparse Depth Completion, Zixuan Huang et al., IEEE, 2019, pp. 3429-3441 (Year: 2019). [cited by examiner]
Huynh L., et al., “Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian Loss”, 2021 International Conference on 3D Vision (3DV), IEEE, Dec. 1, 2021, pp. 228-238. [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/030604—ISA/EPO—Sep. 18, 2024. [cited by applicant]
Khan F. F., et al., “Robust Multimodal Depth Estimation Using Transformer Based Generative Adversarial Networks”, Proceedings of the Genetic and Evolutionary Computation Conference, ACMPUB27, New York, NY, USA, Oct. 10,… [cited by applicant]
Qian J., et al., “SDformer: Efficient End-to-End Transformer for Depth Completion”, 2022 International Conference on Industrial Automation, Robotics and Control Engineering (IARCE), IEEE, Jun. 10, 2022, pp. 56-61. [cited by applicant]