IP Library › Granted Patent US 12,597,262
Granted Patent B2
US 12,597,262 · App. 17/899,734 · Granted Apr 7, 2026

Detecting and identifying objects represented in sensor data generated by multiple sensor systems

Inventors: Georg Kuschk (Garmisch-Partenkirchen, DE); Marc Unzueta Canals (Munich, DE); Sven Möller (Lübbecke, DE); Michael Meyer (Munich, DE); Karl-Heinz Krachenfels (Garmisch-Partenkirchen, DE)
Assignee: GM Cruise Holdings LLC
G06V20/56G01S13/931G01S17/86G01S17/931
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,262
App. No.
17/899,734
Granted
Apr 7, 2026
Kind
B2
Abstract

A system includes a first sensor system of a first modality and a second sensor system of a second modality. The system further includes a computing system that is configured to detect and identify objects represented in sensor signals output by the first and second sensor systems. The computing system employs a hierarchical arrangement of transformers to fuse features of first sensor data output by the first sensor system and second sensor data output by the second sensor system.

Claims (55)

1 . A system comprising:

a first sensor system that generates first sensor data, the first sensor data corresponding to a first modality;

a second sensor system that generates second sensor data, the second sensor data corresponding to a second modality;

a computing system that is in communication with the first sensor system and the second sensor system, wherein the computing system comprises:

a processor; and

memory that stores computer-executable instructions that, when executed by the processor, cause the processor to perform acts comprising:

generating, by a first transformer, a first output based upon the first sensor data, wherein the first output comprises identities of objects determined by the first transformer to be represented in the first sensor data and corresponding locations of the objects in the first sensor data;

generating, by a second transformer, a second output based upon the second sensor data, wherein the second output comprises identities of objects determined by the second transformer to be represented in the second sensor data and corresponding locations of the objects in the second sensor data; and

generating, by a third transformer, a third output based upon the first output and the second output, wherein the third transformer comprises an encoder and a decoder, wherein the encoder processes the second output and the decoder, using cross-attention and self-attention, processes the first output, wherein the cross-attention correlates objects between the first output and the second output, wherein the self-attention evaluates consistency of relationships between the correlated objects, and wherein the third output comprises identities of objects determined by the third transformer to be in an environment of the system and corresponding locations of the objects in the environment of the system.

2 . The system of claim 1 , wherein the first sensor system is a camera and the second sensor system is a radar sensor system.

3 . The system of claim 1 , wherein at least one of the first sensor system or the second sensor system is a lidar system.

4 . The system of claim 1 , the acts further comprising:

extracting features from the first sensor data; and

providing the features and positional encodings to the first transformer, wherein the first transformer generates the first output based upon the extracted features and the positional encodings.

5 . The system of claim 4 , the acts further comprising:

extracting second features from the second sensor data; and

providing the second features and second positional encodings to the second transformer, wherein the second transformer generates the second output based upon the extracted second features and the second positional encodings.

6 . The system of claim 1 , wherein the first sensor data is an image and the second sensor data is a point cloud.

7 . The system of claim 1 , wherein:

the first transformer outputs the first output comprising first vectors, where each vector in the first vectors corresponds to a respective region in the first sensor data and each vector in the first vectors indicates a type of object predicted as being included in the region in the first sensor data;

the second transformer outputs the second output comprising second vectors, wherein each vector in the second vectors corresponds to a respective region in the second sensor data and each vector in the second vectors indicates a type of object predicted as being included in the region in the second sensor data; and

the third transformer receives the first vectors and the second vectors as input data and outputs the third output comprising third vectors, where each vector in the third vectors corresponds to a respective region in the environment of the system and each vector in the third vectors indicates a type of object predicted as being included in the region in the environment of the system.

8 . A method performed by a computing system, the method comprising:

generating, by a first transformer, a first output based upon first sensor data generated by a first sensor system, wherein the first output comprises identities of objects determined by the first transformer to be represented in the first sensor data and corresponding locations of the objects in the first sensor data, and further wherein the first sensor data is in a first modality;

generating, by a second transformer, a second output based upon second sensor data generated by a second sensor system, wherein the second output comprises identities of objects determined by the second transformer to be represented in the second sensor data and corresponding locations of the objects in the second sensor data, and further wherein the second sensor data is in a second modality; and

generating, by a third transformer, a third output based upon the first output and the second output, wherein the third transformer comprises an encoder and a decoder, wherein the encoder processes the second output and the decoder, using cross-attention and self-attention, processes the first output, wherein the cross-attention correlates objects between the first output and the second output, wherein the self-attention evaluates consistency of relationships between the correlated objects, and wherein the third output comprises identities of objects determined by the third transformer to be in an environment of the first sensor system and the second sensor system and corresponding locations of the objects in the environment of the first sensor system and the second sensor system.

9 . The method of claim 8 , wherein the first sensor system is a camera and the second sensor system is a radar sensor system.

10 . The method of claim 8 , wherein at least one of the first sensor system or the second sensor system is a lidar system.

11 . The method of claim 8 , further comprising:

extracting features from the first sensor data; and

providing the features and positional encodings to the first transformer, wherein the first transformer generates the first output based upon the extracted features and the positional encodings.

12 . The method of claim 11 , further comprising:

extracting second features from the second sensor data; and

providing the second features and second positional encodings to the second transformer, wherein the second transformer generates the second output based upon the extracted second features and the second positional encodings.

13 . The method of claim 8 , wherein the first sensor data is an image and the second sensor data is a point cloud.

14 . The method of claim 8 , wherein:

the first transformer outputs the first output comprising first vectors, where each vector in the first vectors corresponds to a respective region in the first sensor data and each vector in the first vectors indicates a type of object predicted as being included in the region in the first sensor data;

the second transformer outputs the second output comprising second vectors, wherein each vector in the second vectors corresponds to a respective region in the second sensor data and each vector in the second vectors indicates a type of object predicted as being included in the region in the second sensor data; and

the third transformer receives the first vectors and the second vectors as input data and outputs the third output comprising third vectors, where each vector in the third vectors corresponds to a respective region in the environment of the first sensor system and the second sensor system and each vector in the third vectors indicates a type of object predicted as being included in the region in the environment of the first sensor system and the second sensor system.

15 . A computer-readable storage device comprising instructions that, when executed by a processor, cause the processor to perform acts comprising:

generating, by a first transformer, a first output based upon first sensor data generated by a first sensor system, wherein the first output comprises identities of objects determined by the first transformer to be represented in the first sensor data and corresponding locations of the objects in the first sensor data, and further wherein the first sensor data is in a first modality;

generating, by a second transformer, a second output based upon second sensor data generated by a second sensor system, wherein the second output comprises identities of objects determined by the second transformer to be represented in the second sensor data and corresponding locations of the objects in the first second sensor data, and further wherein the second sensor data is in a second modality; and

generating, by a third transformer, a third output based upon the first output and the second output, wherein the third transformer comprises an encoder and a decoder, wherein the encoder processes the second output and the decoder, using cross-attention and self-attention, processes the first output, wherein the cross-attention correlates objects between the first output and the second output, wherein the self-attention evaluates consistency of relationships between the correlated objects, and wherein the third output comprises identities of objects determined by the third transformer to be in an environment of the first sensor system and the second sensor system and corresponding locations of the objects in the environment of the first sensor system and the second sensor system.

16 . The computer-readable storage device of claim 15 , wherein the first sensor system is a camera and the second sensor system is a radar sensor system.

17 . The computer-readable storage device of claim 15 , wherein at least one of the first sensor system or the second sensor system is a lidar system.

18 . The computer-readable storage device of claim 15 , the acts further comprising:

extracting features from the first sensor data; and

providing the features and positional encodings to the first transformer, wherein the first transformer generates the first output based upon the extracted features and the positional encodings.

19 . The computer-readable storage device of claim 18 , the acts further comprising:

extracting second features from the second sensor data; and

providing the second features and second positional encodings to the second transformer, wherein the second transformer generates the second output based upon the extracted second features and the second positional encodings.

20 . The computer-readable storage device of claim 15 , wherein:

the first transformer outputs the first output comprising first vectors, where each vector in the first vectors corresponds to a respective region in the first sensor data and each vector in the first vectors indicates a type of object predicted as being included in the region in the first sensor data;

the second transformer outputs the second output comprising second vectors, wherein each vector in the second vectors corresponds to a respective region in the second sensor data and each vector in the second vectors indicates a type of object predicted as being included in the region in the second sensor data; and

the third transformer receives the first vectors and the second vectors as input data and outputs the third output comprising third vectors, where each vector in the third vectors corresponds to a respective region in the environment of the first sensor system and the second sensor system and each vector in the third vectors indicates a type of object predicted as being included in the region in the environment of the first sensor system and the second sensor system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2022
From: MEYER, MICHAEL; CANALS, MARC UNZUETA; MÖLLER, SVEN; KUSCHK, GEORG; KRACHENFELS, KARL-HEINZ
To: GM CRUISE HOLDINGS LLC
Reel/Frame 060963/0868 →
Priority Claims (1)
EP 22184764 · Jul 13, 2022 · regional
Continuity (1)
Related Publication 20240020983A1 · Jan 18, 2024
References Cited (14)
US 10671068B1 · Xu · 2020 [cited by examiner]
US 11921824B1 · Hester · 2024 [cited by examiner]
US 20180136660A1 · Mudalige · 2018 [cited by examiner]
US 20190265714A1 · Ball · 2019 [cited by examiner]
US 20200219264A1 · Brunner · 2020 [cited by examiner]
US 20210103027A1 · Harrison · 2021 [cited by examiner]
US 20210241026A1 · Deng · 2021 [cited by examiner]
US 20220300831A1 · Friede · 2022 [cited by examiner]
US 20230206456A1 · Lee · 2023 [cited by examiner]
US 20250054286A1 · Huang · 2025 [cited by examiner]
US 20250156685A1 · Wu · 2025 [cited by examiner]
EP 4306999A1 · 2024 [cited by applicant]
“Extended European Search Report for European Patent Application No. 22184764.3”, Mailed Date: Jan. 11, 2023, 8 pages. [cited by applicant]
“Response to the Communication Pursuant to Rule 69 EPC for European Patent Application No. 22184764.3”, Filed Date: Feb. 2, 2024, 6 pages. [cited by applicant]