IP Library › Granted Patent US 12,597,265
Granted Patent B2
US 12,597,265 · App. 18/157,034 · Granted Apr 7, 2026

Occlusion resolving gated mechanism for sensor fusion

Inventors: Varun Ravi Kumar (San Diego, CA); Senthil Kumar Yogamani (Headford, IE); Shubhankar Mangesh Borse (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06V20/58G06V10/80B60W30/095B60W2420/403B60W2420/408
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,265
App. No.
18/157,034
Granted
Apr 7, 2026
Kind
B2
Abstract

Techniques and systems are provided for processing sensor data. For instance a process can include obtaining first sensor data of an environment, wherein the first sensor data includes a representation of a first object occluding a second object, obtaining second sensor data of the environment, wherein the second sensor data includes points associated with the first object and points associated with the second object, generating estimated segment data from the first sensor data, wherein the estimated segment data includes a first segment corresponding to the first object; matching points associated with the first object to the first segment, and deemphasizing points associated with the second object based on matching the points associated with the first object to the first segment.

Claims (59)

1 . An apparatus for processing sensor data, comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

obtain first sensor data of an environment from a first sensor, wherein the first sensor data includes a representation of a first object occluding a second object;

obtain second sensor data of the environment from a second sensor, wherein the second sensor data includes points associated with the first object and points associated with the second object, wherein the second sensor differs from the first sensor, wherein the second object is occluded from a view of the first sensor, and wherein the second object is unoccluded, at least in part, from a view of the second sensor;

generate estimated segment data from the first sensor data, wherein the estimated segment data includes a first segment corresponding to the first object;

transform the first sensor data and the second sensor data based on at least one of intrinsic information or extrinsic information regarding the first sensor and second sensor, respectively, to a common coordinate frame;

match points associated with the first object to the first segment based on the transformed first sensor data and second sensor data; and

deemphasize points associated with the second object based on matching the points associated with the first object to the first segment.

2 . The apparatus of claim 1 , wherein the first sensor data comprises image data from an image, and wherein the second sensor data comprises point data.

3 . The apparatus of claim 2 , wherein the at least one processor is further configured to:

obtain point data features based on the point data; and

apply at least one weight to features of the point data features associated with the first object based on the matched points associated with the first object to the first segment.

4 . The apparatus of claim 3 , wherein the at least one processor is further configured to fuse the image data from the image, the point data, and associated weights to generate fused data for input to a perception machine learning algorithm.

5 . The apparatus of claim 4 , wherein the at least one processor is further configured to:

obtain, from the perception machine learning algorithm, an uncertainty map indicating uncertainty values for the fused data; and

update the at least one weight to apply to future features of the point data based on uncertainty values of the uncertainty map.

6 . The apparatus of claim 2 , wherein the point data includes distance information to points in the environment, and wherein the at least one processor is further configured to decompose the point data into layers based on the distance information.

7 . The apparatus of claim 6 , wherein, to match points, the at least one processor is further configured to compare a layer of the layers of decomposed point data to the first segment to match locations of points in the layer to locations in the first segment.

8 . The apparatus of claim 6 , wherein each layer includes respective points associated with a minimum distance and a maximum distance.

9 . The apparatus of claim 2 , wherein the point data comprises at least one of a light detection and ranging (LIDAR) point cloud or a radar point cloud.

10 . The apparatus of claim 1 , wherein, to deemphasize the points associated with the second object, the at least one processor is configured to reduce a weight associated with points of the second object.

11 . A method for processing sensor data, comprising:

obtaining first sensor data of an environment from a first sensor, wherein the first sensor data includes a representation of a first object occluding a second object;

obtaining second sensor data of the environment from a second sensor, wherein the second sensor data includes points associated with the first object and points associated with the second object, wherein the second sensor differs from the first sensor, wherein the second object is occluded from a view of the first sensor, and wherein the second object is unoccluded, at least in part, from a view of the second sensor;

generating estimated segment data from the first sensor data, wherein the estimated segment data includes a first segment corresponding to the first object;

transforming the first sensor data and the second sensor data based on at least one of intrinsic information or extrinsic information regarding the first sensor and second sensor, respectively, to a common coordinate frame;

matching points associated with the first object to the first segment based on the transformed first sensor data and second sensor data; and

deemphasizing points associated with the second object based on matching the points associated with the first object to the first segment.

12 . The method of claim 11 , wherein the first sensor data comprises image data from an image, and wherein the second sensor data comprises point data.

13 . The method of claim 12 , further comprising:

obtaining point data features based on the point data; and

applying at least one weight to features of the point data features associated with the first object based on the matched points associated with the first object to the first segment.

14 . The method of claim 13 , further comprising fusing the image data from the image, the point data, and associated weights to generate fused data for input to a perception machine learning algorithm.

15 . The method of claim 14 , further comprising:

obtaining, from the perception machine learning algorithm, an uncertainty map indicating uncertainty values for the fused data; and

updating the at least one weight to apply to future features of the point data based on uncertainty values of the uncertainty map.

16 . The method of claim 12 , wherein the point data includes distance information to points in the environment, and further comprising decomposing the point data into layers based on the distance information.

17 . The method of claim 16 , wherein matching points comprises comparing a layer of the layers of decomposed point data to the first segment to match locations of points in the layer to locations in the first segment.

18 . The method of claim 16 , wherein each layer includes respective points associated with a minimum distance and a maximum distance.

19 . The method of claim 12 , wherein the point data comprises at least one of a light detection and ranging (LIDAR) point cloud or a radar point cloud.

20 . The method of claim 11 , wherein deemphasizing the points associated with the second object comprises reducing a weight associated with points of the second object.

21 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:

obtain first sensor data of an environment from a first sensor, wherein the first sensor data includes a representation of a first object occluding a second object;

obtain second sensor data of the environment from a second sensor, wherein the second sensor data includes points associated with the first object and points associated with the second object, wherein the second sensor differs from the first sensor, wherein the second object is occluded from a view of the first sensor, and wherein the second object is unoccluded, at least in part, from a view of the second sensor;

generate estimated segment data from the first sensor data, wherein the estimated segment data includes a first segment corresponding to the first object;

transform the first sensor data and the second sensor data based on at least one of intrinsic information or extrinsic information regarding the first sensor and second sensor, respectively, to a common coordinate frame;

match points associated with the first object to the first segment based on the transformed first sensor data and second sensor data; and

deemphasize points associated with the second object based on matching the points associated with the first object to the first segment.

22 . The non-transitory computer-readable medium of claim 21 , wherein the first sensor data comprises image data from an image, and wherein the second sensor data comprises point data.

23 . The non-transitory computer-readable medium of claim 22 , wherein the instructions further cause the at least one processor to:

obtain point data features based on the point data; and

apply at least one weight to features of the point data features associated with the first object based on the matched points associated with the first object to the first segment.

24 . The non-transitory computer-readable medium of claim 23 , wherein the instructions further cause the at least one processor to fuse the image data from the image, the point data, and associated weights to generate fused data for input to a perception machine learning algorithm.

25 . The non-transitory computer-readable medium of claim 24 , wherein the instructions further cause the at least one processor to:

obtain, from the perception machine learning algorithm, an uncertainty map indicating uncertainty values for the fused data; and

update the at least one weight to apply to future features of the point data based on uncertainty values of the uncertainty map.

26 . The non-transitory computer-readable medium of claim 23 , wherein the point data includes distance information to points in the environment, and wherein the instructions further cause the at least one processor to decompose the point data into layers based on the distance information.

27 . The non-transitory computer-readable medium of claim 26 , wherein, to match points, the instructions cause the at least one processor to compare a layer of the layers of decomposed point data to the first segment to match locations of points in the layer to locations in the first segment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2025
From: RAVI KUMAR, VARUN; YOGAMANI, SENTHIL KUMAR; BORSE, SHUBHANKAR MANGESH
To: QUALCOMM INCORPORATED
Reel/Frame 070569/0297 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2023
From: RAVI KUMAR, VARUN; YOGAMANI, SENTHIL KUMAR; BORSE, SHUBHANKAR MANGESH
To: QUALCOMM INCORPORATED
Reel/Frame 062771/0614 →
Continuity (1)
Related Publication 20240249530A1 · Jul 25, 2024
References Cited (9)
US 20210358137A1 · Lee · 2021 [cited by examiner]
US 20220319043A1 · Chandler · 2022 [cited by examiner]
US 20240175728A1 · Dharia · 2024 [cited by examiner]
US 20240210541A1 · Bao · 2024 [cited by examiner]
International Search Report and Written Opinion—PCT/US2023/083177—ISA/EPO—Mar. 20, 2024. [cited by applicant]
Liu Y., et al., “A Multi-Sensor Fusion Based 2D-Driven 3D Object Detection Approach for Large Scene Applications”, 2019 IEEE International Conference on Robotics and Biomimetics (ROBIO), IEEE, Dec. 6, 2019, pp. 2181-218… [cited by applicant]
Wang C., et al., “PointAugmenting: Cross-Modal Augmentation for 3D Object Detection”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 20, 2021, pp. 11789-11798, XP034006655, Abstra… [cited by applicant]
Wang G., et al., “Multi-View Adaptive Fusion Network for 3D Object Detection”, arXiv:2011.00652v2 [cs.CV], Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 8, 2020, pp. 1-11, XP0818… [cited by applicant]
Xu S., et al., “FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection”, 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), IEEE, Indianapolis, USA, Sep. 19-Sep. 21,… [cited by applicant]