IP Library › Granted Patent US 12,437,554
Granted Patent B2
US 12,437,554 · App. 18/170,642 · Granted Oct 7, 2025

Detecting at least one emergency vehicle using a perception algorithm

Inventors: Arindam Das (Mumbai, IN); Sudarshan Paul (Mumbai, IN); Sanjoy Das (Mumbai, IN); Deep Doshi (Troy, MI)
Assignee: Connaught Electronics Ltd.
G06V20/58G06N3/0464B60W2420/403B60W2420/54B60W2422/95
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,554
App. No.
18/170,642
Granted
Oct 7, 2025
Kind
B2
Abstract

For training a perception algorithm to detect an emergency vehicle, respective audio datasets are received from two microphones and respective spectrograms are generated. At least one interaural difference map is generated based on the spectrograms, audio source localization data is generated, which specifies a number of audio sources in respective grid cells of a spatial grid, by applying a CRNN to first input data containing the spectrograms and the least one interaural difference map. An image is received from a camera and output data comprising a bounding box for the emergency vehicle is predicted by applying at least one further ANN to second input data containing the image and the spectrograms. Network parameters are adapted depending on the output data and the audio source localization data.

Claims (82)

1. A computer-implemented method for training a perception algorithm of a vehicle, the computer-implemented method comprising:

providing a convolutional recurrent neural network (CRNN) and at least one further artificial neural network (ANN) to detect at least one emergency vehicle;

receiving at least two time-dependent audio datasets from two microphones mounted at different positions on the vehicle;

generating at least two spectrograms based on the at least two time-dependent audio datasets;

generating at least one interaural difference map based on the at least two spectrograms, which contains at least one of an interaural phase difference map or an interaural time difference map or an interaural level difference map;

generating audio source localization data for at least one grid cell of a predefined spatial grid in an environment of the vehicle;

specifying a number of audio sources in the at least one grid cell;

applying the CRNN to a first input data containing the at least two spectrograms and the at least one interaural difference map;

receiving at least one camera image from at least one camera mounted to the vehicle;

predicting output data of at least one bounding box for the at least one emergency vehicle by applying the at least one further ANN to a second input data containing the at least one camera image and the at least two spectrograms; and

adapting network parameters of the CRNN and the at least one further ANN by depending on the output data and the audio source localization data.

2. The computer-implemented method of claim 1 , wherein adapting the network parameters further comprising:

providing a ground truth for the at least one bounding box given by or dependent upon the audio source localization data.

3. The computer-implemented method of claim 1 , further comprising:

generating at least one visual domain feature map by applying a first object detection ANN of the at least one further ANN to the at least one camera image;

generating at least one audio domain feature map by applying an audio source localization ANN of the at least one further ANN to the at least two spectrograms;

fusing the at least one visual domain feature map and the at least one audio domain feature map; and

predicting the output data by applying a second object detection ANN of the at least one further ANN to the fused feature maps.

4. The computer-implemented method of claim 3 , wherein fusing the at least one visual domain feature map and the at least one audio domain feature map further comprising:

generating at least one difference map by subtracting the at least one visual domain feature map and the at least one audio domain feature map from each other;

generating at least one re-calibration vector based on the at least one difference map;

multiplying the at least one visual domain feature map with the at least one re-calibration vector to generate at least one re-calibrated visual domain feature map;

multiplying the at least one audio domain feature map with the at least one re-calibration vector to generate at least one re-calibrated audio domain feature map; and

fusing the at least one re-calibrated video domain feature map and the at least one re-calibrated audio domain feature map.

5. The computer-implemented method of claim 4 , wherein generating the at least one re-calibration vector further comprising:

applying sum pooling, followed by a normalization, followed by a global average pooling, followed by an activation function, to the at least one difference map.

6. The computer-implemented method of claim 3 , wherein the first object detection ANN is designed for object detection according to a predefined set of object classes,

wherein the set of object classes contains a class for activated emergency vehicle lights.

7. The computer-implemented method of claim 1 , wherein each audio dataset of the at least two time-dependent audio datasets further comprising:

generating a corresponding two-dimensional single magnitude map as a function of time and audio amplitude,

wherein the first input data comprises the single magnitude maps.

8. A perception system for a vehicle, the perception system comprising:

a data processing apparatus;

a memory device configured to store a perception algorithm;

a convolutional recurrent neural network (CRNN) and at least one further artificial neural network (ANN) configured to detect at least one emergency vehicle;

at least two microphones mounted at different positions on the vehicle configured to generate at least two time-dependent audio datasets;

at least one camera mounted to the vehicle configured to generate at least one camera image; and

a computing unit configured to:

generate at least two spectrograms based on the at least two time-dependent audio datasets;

generate at least one interaural difference map based on the at least two spectrograms, which contains at least one of an interaural phase difference map or an interaural time difference map or an interaural level difference map:

generate audio source localization data for at least one grid cell of a predefined spatial grid in an environment of the vehicle;

specify a number of audio sources in the at least one grid cell:

apply the CRNN to a first input data containing the at least two spectrograms and the at least one interaural difference map;

predict output data of at least one bounding box for the at least one emergency vehicle by applying the at least one further ANN to a second input data containing the at least one camera image and the at least two spectrograms; and

adapt network parameters of the CRNN and the at least one further ANN by depending on the output data and the audio source localization data.

9. The perception system of claim 8 , wherein the at least one camera image depicts the environment of the vehicle, and

wherein the computing unit is configured to predict output data including at least one bounding box for at least one emergency vehicle in the environment of the vehicle by applying the at least one further ANN to the second input data containing the at least one camera image and the at least two spectrograms.

10. The perception system of claim 8 , wherein the at least two microphones are implemented as respective uni-directional microphones.

11. The perception system of claim 8 , wherein the at least two microphones includes a first microphone mounted in a front half of the vehicle and a second microphone is mounted in a rear half of the vehicle.

12. The perception system of claim 8 , wherein the at least two microphones are configured to have a respective direction of maximum sensitivity parallel to a longitudinal axis of the vehicle.

13. The perception system of claim 8 , wherein the data processing apparatus is configured to execute at least one instructions.

14. The perception system of claim 8 , wherein adapt the network parameters further comprises:

provide a ground truth for the at least one bounding box given by or dependent upon the audio source localization data.

15. The perception system of claim 8 , wherein the computing unit further comprises:

generate at least one visual domain feature map by applying a first object detection ANN of the at least one further ANN to the at least one camera image;

generate at least one audio domain feature map by applying an audio source localization ANN of the at least one further ANN to the at least two spectrograms;

fuse the at least one visual domain feature map and the at least one audio domain feature map; and

predict the output data by applying a second object detection ANN of the at least one further ANN to the fused feature maps.

16. The perception system of claim 15 , wherein fuse the at least one visual domain feature map and the at least one audio domain feature map further comprises:

generate at least one difference map by subtracting the at least one visual domain feature map and the at least one audio domain feature map from each other;

generate at least one re-calibration vector based on the at least one difference map;

multiply the at least one visual domain feature map with the at least one re-calibration vector to generate at least one re-calibrated visual domain feature map;

multiply the at least one audio domain feature map with the at least one re-calibration vector to generate at least one re-calibrated audio domain feature map; and

fuse the at least one re-calibrated video domain feature map and the at least one re-calibrated audio domain feature map.

17. The perception system of claim 16 , wherein generate the at least one re-calibration vector further comprises:

apply sum pooling, followed by a normalization, followed by a global average pooling, followed by an activation function, to the at least one difference map.

18. The perception system of claim 15 , wherein the first object detection ANN is designed for object detection according to a predefined set of object classes,

wherein the set of object classes contains a class for activated emergency vehicle lights.

19. The perception system of claim 8 , wherein each audio dataset of the at least two time-dependent audio datasets further comprises:

generate a corresponding two-dimensional single magnitude map as a function of time and audio amplitude,

wherein the first input data comprises the single magnitude maps.

20. A non-transitory computer-readable storage medium storing instructions, which when executed on a computer, cause the computer to perform a method for training a perception algorithm of a vehicle, the method comprising:

providing a convolutional recurrent neural network (CRNN) and at least one further artificial neural network (ANN) to detect at least one emergency vehicle;

receiving at least two time-dependent audio datasets from two microphones mounted at different positions on the vehicle;

generating at least two spectrograms based on the at least two time-dependent audio datasets;

generating at least one interaural difference map based on the at least two spectrograms, which contains at least one of an interaural phase difference map or an interaural time difference map or an interaural level difference map;

generating audio source localization data for at least one grid cell of a predefined spatial grid in an environment of the vehicle;

specifying a number of audio sources in the at least one grid cell;

applying the CRNN to a first input data containing the at least two spectrograms and the at least one interaural difference map;

receiving at least one camera image from at least one camera mounted to the vehicle;

predicting output data of at least one bounding box for the at least one emergency vehicle by applying the at least one further ANN to a second input data containing the at least one camera image and the at least two spectrograms; and

adapting network parameters of the CRNN and the at least one further ANN by depending on the output data and the audio source localization data.

Continuity (1)
Related Publication 20240282116A1 · Aug 22, 2024
References Cited (12)
US 11610599B2 · Grauman et al. · 2023 [cited by applicant]
US 11816987B2 · Dantrey et al. · 2023 [cited by applicant]
US 20180137756A1 · Moosaei et al. · 2018 [cited by applicant]
US 20210034914A1 · Bansal · 2021 [cited by applicant]
US 20210103747A1 · Moustafa et al. · 2021 [cited by applicant]
US 20220048529A1 · Ding · 2022 [cited by applicant]
US 20220157165A1 · Dantrey · 2022 [cited by examiner]
US 20220219736A1 · Xu · 2022 [cited by examiner]
EP 3965066A2 · 2022 [cited by applicant]
EP 3971770A2 · 2022 [cited by applicant]
European Patent Office, International Search Report (with English translation) and Written Opinion of corresponding International Application No. PCT/US2024/014598, dated Jun. 10, 2024. [cited by applicant]
Qian Rui et al.: “Multiple Sound Sources Localization from Coarse to Fine”, (Nov. 12, 2020) Springer, pp. 292-308, XP047588728, Section 3.3. [cited by applicant]
Cited By (1)
US 12,738,160