IP Library › Granted Patent US 12,601,831
Granted Patent B2
US 12,601,831 · App. 18/463,049 · Granted Apr 14, 2026

Radar and camera fusion for vehicle applications

Inventors: Senthil Kumar Yogamani (Headford, IE); Varun Ravi Kumar (San Diego, CA)
Assignee: QUALCOMM Incorporated
G01S13/867G01S7/354G01S7/417G01S13/931
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,601,831
App. No.
18/463,049
Granted
Apr 14, 2026
Kind
B2
Abstract

This disclosure provides systems, methods, and devices for vehicle driving assistance systems that support image processing. In a first aspect, a method of image processing includes receiving image BEV features and receiving first radio detection and ranging (RADAR) BEV features. The first RADAR BEV features that are received are determined based on first RADAR data associated with a first data type. First normalized RADAR BEV features are determined, which includes rescaling the first RADAR BEV features using a first attention mechanism based on the image BEV features and the first RADAR BEV features. Fused data is determined that combines the first normalized RADAR BEV features and the image BEV features. Other aspects and features are also claimed and described.

Claims (73)

1 . A method for image processing for use in a vehicle assistance system, comprising:

receiving image Bird's Eye View (BEV) features;

receiving first radio detection and ranging (RADAR) BEV features that are determined based on first RADAR data associated with a first data type;

determining first normalized RADAR BEV features, which includes rescaling the first RADAR BEV features using a first attention mechanism based on the image BEV features and the first RADAR BEV features; and

determining fused data that combines the first normalized RADAR BEV features and the image BEV features.

2 . The method of claim 1 , wherein the first attention mechanism is a self-attention mechanism.

3 . The method of claim 1 , wherein the first attention mechanism is a cross-attention mechanism.

4 . The method of claim 1 , wherein the first data type is one of the data types in the group consisting of: a range-doppler map, a range-azimuth map, a point cloud, a 2D RADAR image, a list of detected objects based on the first RADAR data, and a set of 3D bounding boxes based on the first RADAR data.

5 . The method of claim 4 , wherein

if the first data type is the range-doppler map data type, the range-azimuth map data type, or the point cloud data type, then the first attention mechanism is a cross-attention mechanism, and

if the first data type is the 2D RADAR image type, the list of detected objects based on the first RADAR data type, or the set of 3D bounding boxes based on the first RADAR data type, then the attention mechanism is a self-attention mechanism.

6 . The method of claim 1 , further comprising:

receiving second RADAR BEV features that are determined based on second RADAR data associated with a second data type different than the first data type; and

determining second normalized RADAR BEV features, which includes rescaling the second RADAR BEV features using a second attention mechanism based on the image BEV features and the second RADAR BEV features,

wherein the fused data combines the first normalized RADAR BEV features, the second normalized RADAR BEV features, and the image BEV features.

7 . The method of claim 6 , wherein the first attention mechanism is a cross-attention mechanism and the second attention mechanism is a self-attention mechanism.

8 . The method of claim 1 , wherein the image BEV features are determined based on image data received from an image sensor.

9 . The method of claim 1 , wherein the attention mechanism includes learned attention weights based on ground-truth annotations of a camera that captured image data on which the image BEV features are based.

10 . The method of claim 1 , further comprising:

detecting an object based on the fused data; and

controlling a function of a vehicle based on the object detected.

11 . An apparatus, comprising:

a memory storing processor-readable code; and

at least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations including:

receiving image Bird's Eye View (BEV) features;

receiving first radio detection and ranging (RADAR) BEV features that are determined based on first RADAR data associated with a first data type;

determining normalized RADAR BEV features, which includes rescaling the first RADAR BEV features using a first attention mechanism based on the image BEV features and the first RADAR BEV features; and

determining fused data that combines the normalized RADAR BEV features and the image BEV features.

12 . The apparatus of claim 11 , wherein the first attention mechanism is a self-attention mechanism.

13 . The apparatus of claim 11 , wherein the first attention mechanism is a cross-attention mechanism.

14 . The apparatus of claim 11 , wherein the first data type is one of the data types in the group consisting of: a range-doppler map, a range-azimuth map, a point cloud, a 2D RADAR image, a list of detected objects, and a set of 3D bounding boxes.

15 . The apparatus of claim 14 , wherein

if the first data type is the range-doppler map data type, the range-azimuth map data type, or the point cloud data type, then the first attention mechanism is a cross-attention mechanism, and

if the first data type is the 2D RADAR image type, the list of detected objects based on the first RADAR data type, or the set of 3D bounding boxes based on the first RADAR data type, then the attention mechanism is a self-attention mechanism.

16 . The apparatus of claim 11 , wherein the operations further include:

receiving second RADAR BEV features that are determined based on second RADAR data associated with a second data type different than the first data type; and

determining second normalized RADAR BEV features, which includes rescaling the second RADAR BEV features using a second attention mechanism based on the image BEV features and the second RADAR BEV features,

wherein the fused data combines the first normalized RADAR BEV features, the second normalized RADAR BEV features, and the image BEV features.

17 . The apparatus of claim 16 , wherein the first attention mechanism is a cross-attention mechanism and the second attention mechanism is a self-attention mechanism.

18 . The apparatus of claim 11 , wherein the image BEV features are determined based on image data received from an image sensor.

19 . The apparatus of claim 11 , wherein the attention mechanism includes learned attention weights based on ground-truth annotations of a camera that captured image data on which the image BEV features are based.

20 . The apparatus of claim 11 , wherein the operations further include:

detecting an object based on the fused data; and

controlling a function of a vehicle based on the object detected.

21 . A non-transitory computer-readable medium storing instructions that, when executed by a processing system that includes one or more processors, cause the processing system to perform operations comprising:

receiving image Bird's Eye View (BEV) features;

receiving first radio detection and ranging (RADAR) BEV features that are determined based on first RADAR data associated with a first data type;

determining normalized RADAR BEV features, which includes rescaling the first RADAR BEV features using a first attention mechanism based on the image BEV features and the first RADAR BEV features; and

determining fused data that combines the normalized RADAR BEV features and the image BEV features.

22 . The non-transitory, computer-readable medium of claim 21 , wherein the first attention mechanism is a self-attention mechanism or a cross-attention mechanism.

23 . The non-transitory, computer-readable medium of claim 21 , wherein the first data type is one of the data types in the group consisting of: a range-doppler map, a range-azimuth map, a point cloud, a 2D RADAR image, a list of detected objects, and a set of 3D bounding boxes.

24 . The non-transitory, computer-readable medium of claim 21 , wherein the operations further include:

receiving second RADAR BEV features that are determined based on second RADAR data associated with a second data type different than the first data type; and

determining second normalized RADAR BEV features, which includes rescaling the second RADAR BEV features using a second attention mechanism based on the image BEV features and the second RADAR BEV features,

wherein the fused data combines the first normalized RADAR BEV features, the second normalized RADAR BEV features, and the image BEV features.

25 . The non-transitory, computer-readable medium of claim 21 , wherein the attention mechanism includes learned attention weights based on ground-truth annotations of a camera that captured the image data.

26 . A vehicle, comprising:

an image sensor;

a radio detection and ranging (RADAR) system; and

a memory storing processor-readable code; and

at least one processor coupled to the memory, the at least one processor in communication with the image sensor and the RADAR system and configured to execute the processor-readable code to cause the at least one processor to perform operations including:

receiving image Bird's Eye View (BEV) features that are determined based on image data received from the image sensor;

receiving first RADAR BEV features that are determined based on first RADAR data associated with a first data type, the first RADAR data received from the RADAR system;

determining normalized RADAR BEV features, which includes rescaling the first RADAR BEV features using a first attention mechanism based on the image BEV features and the first RADAR BEV features; and

determining fused data that combines the normalized RADAR BEV features and the image BEV features.

27 . The vehicle of claim 26 , wherein the first attention mechanism is a self-attention mechanism or a cross-attention mechanism.

28 . The vehicle of claim 26 , wherein the RADAR system is a first RADAR system, the vehicle further comprising a second RADAR system in communication with the processing system, the operations including:

receiving second RADAR BEV features that are determined based on second RADAR data associated with a second data type different than the first data type, the second RADAR data received from the second RADAR system; and

determining second normalized RADAR BEV features, which includes rescaling the second RADAR BEV features using a second attention mechanism based on the image BEV features and the second RADAR BEV features,

wherein the fused data combines the first normalized RADAR BEV features, the second normalized RADAR BEV features, and the image BEV features.

29 . The vehicle of claim 28 , wherein the first data type is one of the data types in the group consisting of: a range-doppler map, a range-azimuth map, and a point cloud, and

wherein the second data type is one of the data types in the group consisting of: a 2D RADAR image, a list of detected objects, and a set of 3D bounding boxes.

30 . The vehicle of claim 26 , wherein the attention mechanism includes learned attention weights based on ground-truth annotations of a camera that captured the image data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: YOGAMANI, SENTHIL KUMAR; RAVI KUMAR, VARUN
To: QUALCOMM INCORPORATED
Reel/Frame 065014/0028 →
Continuity (1)
Related Publication 20250085413A1 · Mar 13, 2025
References Cited (15)
US 10176405B1 · Zhou · 2019 [cited by examiner]
US 11222217B1 · Zhang · 2022 [cited by examiner]
US 20110169957A1 · Bartz · 2011 [cited by examiner]
US 20190100144A1 · Asayama · 2019 [cited by examiner]
US 20190122103A1 · Gao · 2019 [cited by examiner]
US 20190311223A1 · Wang · 2019 [cited by examiner]
US 20210358137A1 · Lee · 2021 [cited by examiner]
US 20220303505A1 · Itoh · 2022 [cited by examiner]
US 20230071437A1 · Kim · 2023 [cited by examiner]
US 20230260266A1 · Karasev · 2023 [cited by examiner]
A. Vaswani et al; “Attention Is All You Need”; Proceedings of the 31st Conference on Neural Information Processing Systems; Long Beach, CA, USA; published in the year 2017. (Year: 2017). [cited by examiner]
“Brand Guide for Bluetooth Trademarks”; no author given; published by the Bluetooth Special Interest Group; Kirkland, Washington, USA; posted on the Internet at bluetooth.com; dated Jun. 2022. (Year: 2022). [cited by examiner]
G. Brauwers et al., “A General Survey on Attention Mechanisms in Deep Learning”; arXiv:2203.14263v1; Mar. 27, 2022. (Year: 2022). [cited by examiner]
“Who We Are: Our Brands”; no author given; published by the Wi-Fi Alliance; Austin, TX, USA; posted on the Internet at wi-fi.org; copyright in the year 2024. (Year: 2024). [cited by examiner]
“Guidance for use of the LTE logo”; no author given; published by 3GPP Partners; Sophia Antipolis, France; posted on the Internet at 3gpp.org; accessed in the year 2024. (Year: 2024). [cited by examiner]