IP Library › Granted Patent US 12,602,915
Granted Patent B2
US 12,602,915 · App. 18/313,287 · Granted Apr 14, 2026

Feature fusion for near field and far field images for vehicle applications

Inventors: Varun Ravi Kumar (San Diego, CA); Senthil Kumar Yogamani (Headford, IE)
Assignee: QUALCOMM Incorporated
G06V10/806G06V10/7715G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,915
App. No.
18/313,287
Granted
Apr 14, 2026
Kind
B2
Abstract

This disclosure provides systems, methods, and devices for vehicle driving assistance systems that support image processing. In a first aspect, a method of fusing features from near-field images and far-field images is provided that includes determining feature vectors and spatial locations for received images from near-field and far-field image sensors. A first set of weighted feature vectors may be determined based on spatial locations of the features and a second set of weighted feature vectors may be determined based on corresponding features between the feature vectors. Fused feature vectors may then be determined based on the weighted feature vectors, such as using a transformer attention process trained to select and combine features from both sets of weighted feature vectors. Vehicle control instructions may be determined based on the fused feature vectors. Other aspects and features are also claimed and described.

Claims (72)

1 . A method for image processing for use in a vehicle assistance system, comprising:

receiving a first plurality of images captured by near-field image sensors and a second plurality of images captured by far-field image sensors;

determining a first set of feature vectors from the first plurality of images and a second set of feature vectors from the second plurality of images;

determining a first set of weighted feature vectors based on spatial locations for features from the first set of feature vectors and the second set of feature vectors;

determining a second set of weighted feature vectors based on corresponding features between the first set of feature vectors and the second set of feature vectors;

determining fused feature vectors based on the first set of weighted feature vectors and the second set of weighted feature vectors; and

controlling a vehicle based on the fused feature vectors.

2 . The method of claim 1 , wherein determining the first set of weighted feature vectors comprises:

determining, for feature values from the first set of feature vectors, higher weights to features that are located closer to the vehicle; and

determining, for feature values from the second set of feature vectors, higher weights to features that are located further from the vehicle.

3 . The method of claim 1 , wherein determining the second set of weighted feature vectors comprises determining sets of corresponding features between the feature vectors, wherein each set of corresponding features identifies at least two feature values from at least two feature vectors located in the same or similar location.

4 . The method of claim 3 , wherein each of at least a subset of the sets of corresponding features is used to determine a corresponding feature value of the second set of weighted feature vectors.

5 . The method of claim 1 , wherein determining the second set of weighted feature vectors comprises determining weighted feature values based on feature values for corresponding features.

6 . The method of claim 5 , wherein the weighted feature values are determined using a pixel adaptive convolution process.

7 . The method of claim 1 , wherein the fused feature vectors are determined based on a transformer attention process that receives the first set of weighted feature vectors and the second set of weighted feature vectors.

8 . The method of claim 7 , wherein the second set of weighted feature vectors are provided as query vectors to the transformer attention process and the first set of weighted feature vectors are provided as key vectors to the transformer attention process.

9 . The method of claim 8 , wherein the fused feature vectors are determined based on output values from the transformer attention process.

10 . The method of claim 1 , wherein the feature vectors from the first plurality of images and the second plurality of images are identified by an encoder model.

11 . The method of claim 1 , further comprising determining spatial locations of the first and second sets of feature vectors within a physical area surrounding a vehicle.

12 . The method of claim 11 , wherein the spatial locations are determined as part of a top view of the physical area surrounding the vehicle.

13 . The method of claim 11 , wherein the spatial locations specify a distance from the vehicle and an angular offset from a heading of the vehicle for corresponding features.

14 . The method of claim 1 , further comprising:

determining, based on the fused feature vectors, a top view segmentation map of an area surrounding the vehicle,

wherein the vehicle control instructions are determined based on the top view segmentation map.

15 . An apparatus, comprising:

a memory storing processor-readable code; and

at least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations including:

receiving a first plurality of images captured by near-field image sensors and a second plurality of images captured by far-field image sensors;

determining a first set of feature vectors from the first plurality of images and a second set of feature vectors from the second plurality of images;

determining a first set of weighted feature vectors based on spatial locations for features from the first set of feature vectors and the second set of feature vectors;

determining a second set of weighted feature vectors based on corresponding features between the first set of feature vectors and the second set of feature vectors;

determining fused feature vectors based on the first set of weighted feature vectors and the second set of weighted feature vectors; and

controlling a vehicle based on the fused feature vectors.

16 . The apparatus of claim 15 , wherein determining the first set of weighted feature vectors comprises:

determining, for feature values from the first set of feature vectors, higher weights to features that are located closer to the vehicle; and

determining, for feature values from the second set of feature vectors, higher weights to features that are located further from the vehicle.

17 . The apparatus of claim 15 , wherein determining the second set of weighted feature vectors comprises determining weighted feature values based on feature values for corresponding features.

18 . The apparatus of claim 17 , wherein the weighted feature values are determined using a pixel adaptive convolution process.

19 . The apparatus of claim 15 , wherein the fused feature vectors are determined based on a transformer attention process that receives the first set of weighted feature vectors and the second set of weighted feature vectors.

20 . The apparatus of claim 19 , wherein the second set of weighted feature vectors are provided as query vectors to the transformer attention process and the first set of weighted feature vectors are provided as key vectors to the transformer attention process.

21 . The apparatus of claim 20 , wherein the fused feature vectors are determined based on output values from the transformer attention process.

22 . The apparatus of claim 15 , wherein the operations further comprise:

determining, based on the fused feature vectors, a top view segmentation map of an area surrounding the vehicle,

wherein the vehicle control instructions are determined based on the top view segmentation map.

23 . A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising:

receiving a first plurality of images captured by near-field image sensors and a second plurality of images captured by far-field image sensors;

determining a first set of feature vectors from the first plurality of images and a second set of feature vectors from the second plurality of images;

determining a first set of weighted feature vectors based on spatial locations for features from the first set of feature vectors and the second set of feature vectors;

determining a second set of weighted feature vectors based on corresponding features between the first set of feature vectors and the second set of feature vectors;

determining fused feature vectors based on the first set of weighted feature vectors and the second set of weighted feature vectors; and

controlling a vehicle based on the fused feature vectors.

24 . The non-transitory, computer-readable medium of claim 23 , wherein determining the first set of weighted feature vectors comprises:

determining, for feature values from the first set of feature vectors, higher weights to features that are located closer to the vehicle; and

determining, for feature values from the second set of feature vectors, higher weights to features that are located further from the vehicle.

25 . The non-transitory, computer-readable medium of claim 23 , wherein the fused feature vectors are determined based on a transformer attention process that receives the first set of weighted feature vectors and the second set of weighted feature vectors.

26 . The non-transitory, computer-readable medium of claim 25 , wherein the second set of weighted feature vectors are provided as query vectors to the transformer attention process and the first set of weighted feature vectors are provided as key vectors to the transformer attention process.

27 . A vehicle, comprising:

near-field image sensors;

far-field image sensors;

a memory storing processor-readable code; and

at least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations including:

receiving a first plurality of images captured by the near-field image sensors and a second plurality of images captured by the far-field image sensors;

determining a first set of feature vectors from the first plurality of images and a second set of feature vectors from the second plurality of images;

determining a first set of weighted feature vectors based on spatial locations for features from the first set of feature vectors and the second set of feature vectors;

determining a second set of weighted feature vectors based on corresponding features between the first set of feature vectors and the second set of feature vectors;

determining fused feature vectors based on the first set of weighted feature vectors and the second set of weighted feature vectors; and

controlling the vehicle based on the fused feature vectors.

28 . The vehicle of claim 27 , wherein determining the first set of weighted feature vectors comprises:

determining, for feature values from the first set of feature vectors, higher weights to features that are located closer to the vehicle; and

determining, for feature values from the second set of feature vectors, higher weights to features that are located further from the vehicle.

29 . The vehicle of claim 27 , wherein the fused feature vectors are determined based on a transformer attention process that receives the first set of weighted feature vectors and the second set of weighted feature vectors.

30 . The vehicle of claim 29 , wherein the second set of weighted feature vectors are provided as query vectors to the transformer attention process and the first set of weighted feature vectors are provided as key vectors to the transformer attention process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: RAVI KUMAR, VARUN; YOGAMANI, SENTHIL KUMAR
To: QUALCOMM INCORPORATED
Reel/Frame 063771/0041 →
Continuity (1)
Related Publication 20240371147A1 · Nov 7, 2024
References Cited (16)
US 11921824B1 · Hester · 2024 [cited by examiner]
US 20130103257A1 · Almedia · 2013 [cited by examiner]
US 20200025931A1 · Liang · 2020 [cited by examiner]
US 20210064913A1 · Ko et al. · 2021 [cited by applicant]
US 20210182596A1 · Adams et al. · 2021 [cited by applicant]
US 20210406674A1 · Wu et al. · 2021 [cited by applicant]
US 20220114807A1 · Iancu et al. · 2022 [cited by applicant]
US 20220261590A1 · Brahma · 2022 [cited by examiner]
US 20220351526A1 · Bar Zvi et al. · 2022 [cited by applicant]
US 20220358328A1 · Wu · 2022 [cited by examiner]
US 20230213643A1 · Hwang · 2023 [cited by examiner]
Wu et al., HSTA: A Hierarchical Spatio-Temporal Attention Model for Trajectory Prediction, 2021, IEEE Transactions on Vehicular Technology 70(11); 11295-11307. (Year: 2021). [cited by examiner]
Su et al., Pixel-Adaptive Convolutional Neural Networks, 2019, arXiv:1904.05373v1, pp. 1-13. (Year: 2019). [cited by examiner]
Xu et al., ACDet: Attentive Cross-view Fusion for LiDAR-based 3D Object Detection, 2022, International Conference on 3D Vision, pp. 1-11. (Year: 2022). [cited by examiner]
Sharath et al., A dynamic two-dimensional (D2D) weight-based map-matching algorithm, Transportation Research Part C 98 (2019): 409-432. (Year: 2018). [cited by examiner]
International Search Report and Written Opinion—PCT/US2024/017687—ISA/EPO—Jun. 6, 2024. [cited by applicant]