IP Library › Granted Patent US 12,269,511
Granted Patent B2
US 12,269,511 · App. 17/149,638 · Granted Apr 8, 2025

Emergency vehicle audio and visual detection post fusion

Inventors: Kecheng Xu (Sunnyvale, CA); Hongyi Sun (Sunnyvale, CA); Qi Luo (Sunnyvale, CA); Wei Wang (Sunnyvale, CA); Zejun Lin (Sunnyvale, CA); Wesley Reynolds (Sunnyvale, CA); Feng Liu (Sunnyvale, CA); Jiangtao Hu (Sunnyvale, CA); Jinghao Miao (Sunnyvale, CA)
Assignee: BAIDU USA LLC
B60W60/0027B60W10/18B60W10/20B60W30/18163G05B13/027G06F18/21G06N3/045G06V10/22G06V20/41G06V20/584G10L25/30G10L25/51H04N23/90H04R1/08H04R1/406H04R3/005B60W2420/40B60W2420/54B60W2554/402B60W2554/4041B60W2554/4044H04R2499/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,269,511
App. No.
17/149,638
Granted
Apr 8, 2025
Kind
B2
Abstract

In one embodiment, an emergency vehicle detection system can be provided in the ADV travelling on a road to detect the presence of an emergency vehicle in a surrounding environment of the ADV using both audio data and visual data. The emergency vehicle detection system can use a trained neutral network to independently generate a detection result from the audio data, and use another trained network to independently generate another detection result from the visual data. The emergency vehicle detection system can fuse the two detection results to determine the position and moving direction of the emergency vehicle. The ADV can take appropriate actions in response to the position and moving direction of the emergency vehicle.

Claims (49)

1. A computer-implemented method of operating an autonomous driving vehicle (ADV), the method comprising:

receiving, at an autonomous driving system (ADS) on the ADV, a stream of audio signals captured using one or more audio capturing devices and a sequence of image frames captured using one or more image capturing devices mounted on the ADV from a surrounding environment of the ADV;

determining, by the ADS using a first neural network model, a first detection result including a first probability that at least a portion of the stream of captured audio signals is from a siren sound, and a moving direction of the siren sound including a moving direction indicator indicating whether a source of the siren sound is moving towards the ADV or moving away from the ADV;

determining, by the ADS using a second neural network model, a second detection result including a second probability that at least one image frame of the sequence of image frames is from an emergency vehicle, and a distance between the ADV and the emergency vehicle, wherein the distance between the ADV and the emergency vehicle is determined based on a size of a bounding box surrounding the at least one image frame and one or more extrinsic parameters of an image capturing device used to capture the at least one image frame, wherein the size of bounding box and the one or more extrinsic parameters of the image capturing device is used as part of labeling data of the at least one image frame, and wherein the one or more extrinsic parameters include a relative rotation and translation between cameras in a multi-camera arrangement;

determining that an emergency vehicle is present in the surrounding environment in response to at least one of the first probability of the first neural network model exceeds a first predefined threshold or the second probability of the second neural network model exceeds a second predefined threshold;

determining a position of the emergency vehicle and a moving direction of the emergency vehicle by fusing the first detection result of the first neural network model and the second detection result of the second neural network model;

determining whether the emergency vehicle is moving towards the ADV based on the position of the emergency vehicle and the moving direction of the emergency vehicle; and

controlling the ADV by steering out of a current driving lane, braking to decelerate, or steering to a side of a road, in response to determining that the emergency vehicle is moving towards the ADV.

2. The method of claim 1 , further comprising:

determining, using the first neural network model, an angle between the ADV and the source of the siren sound, and the moving direction of the source.

3. The method of claim 1 ,

wherein the controlling the ADV by steering out of the current driving lane, braking to decelerate, or steering to the side of the road comprises:

controlling the ADV, by steering out of the current driving lane or braking to decelerate.

4. The method of claim 1 , wherein the first neural network model is trained with audio data representing emergency vehicle siren collected from a plurality of emergency vehicles, and wherein the second neural network model is trained with visual data collected simultaneously to the collecting of the audio data.

5. The method of claim 1 , wherein each of the first neural network model and the second neural network model is a convolutional neural network.

6. The method of claim 1 , wherein the one or more audio capturing devices include one or more microphones, and wherein the one or more image capturing devices include one or more cameras.

7. A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations of operating an autonomous driving vehicle (ADV), the operations comprising:

receiving, at an autonomous driving system (ADS) on the ADV, a stream of audio signals captured using one or more audio capturing devices and a sequence of image frames captured using one or more image capturing devices mounted on the ADV from a surrounding environment of the ADV;

determining, by the ADS using a first neural network model, a first detection result including a first probability that at least a portion of the stream of captured audio signals is from a siren sound, and a moving direction of the siren sound including a moving direction indicator indicating whether a source of the siren sound is moving towards the ADV or moving away from the ADV;

determining, by the ADS using a second neural network model, a second detection result including a second probability that at least one image frame of the sequence of image frames is from an emergency vehicle, and a distance between the ADV and the emergency vehicle, wherein the distance between the ADV and the emergency vehicle is determined based on a size of a bounding box surrounding the at least one image frame and one or more extrinsic parameters of an image capturing device used to capture the at least one image frame, wherein the size of bounding box and the one or more extrinsic parameters of the image capturing device is used as part of labeling data of the at least one image frame, and wherein the one or more extrinsic parameters include a relative rotation and translation between cameras in a multi-camera arrangement;

determining that an emergency vehicle is present in the surrounding environment in response to at least one of the first probability of the first neural network model exceeds a first predefined threshold or the second probability of the second neural network model exceeds a second predefined threshold;

determining a position of the emergency vehicle and a moving direction of the emergency vehicle by fusing the first detection result of the first neural network model and the second detection result of the second neural network model;

determining whether the emergency vehicle is moving towards the ADV based on the position of the emergency vehicle and the moving direction of the emergency vehicle; and

controlling the ADV by steering out of a current driving lane, braking to decelerate, or steering to a side of a road, in response to determining that the emergency vehicle is moving towards the ADV.

8. The machine-readable medium of claim 7 , wherein the operations further comprise:

determining, using the first neural network model, an angle between the ADV and the source of the siren sound, and the moving direction of the source.

9. The machine-readable medium of claim 7 , wherein the controlling the ADV by steering out of the current driving lane, braking to decelerate, or steering to the side of the road comprises:

controlling the ADV by steering out of the current driving lane or braking to decelerate.

10. The machine-readable medium of claim 7 , wherein the first neural network model is trained with audio data representing emergency vehicle siren collected from a plurality of emergency vehicles, and wherein the second neural network model is trained with visual data collected simultaneously to the collecting of the audio data.

11. The machine-readable medium of claim 7 , wherein each of the first neural network model and the second neural network model is a convolutional neural network.

12. The machine-readable medium of claim 7 , wherein the one or more audio capturing devices include one or more microphones, and wherein the one or more image capturing devices include one or more cameras.

13. A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations of operating an autonomous driving vehicle (ADV), the operations comprising:

receiving, at an autonomous driving system (ADS) on the ADV, a stream of audio signals captured using one or more audio capturing devices and a sequence of image frames captured using one or more image capturing devices mounted on the ADV from a surrounding environment of the ADV,

determining, by the ADS using a first neural network model, a first detection result including a first probability that at least a portion of the stream of audio signals is from a siren sound, and a moving direction of the siren sound including a moving direction indicator indicating whether a source of the siren sound is moving towards the ADV or moving away from the ADV,

determining, by the ADS using a second neural network model, a second detection result including a second probability that at least one image frame of the sequence of captured image frames is from an emergency vehicle, and a distance between the ADV and the emergency vehicle, wherein the distance between the ADV and the emergency vehicle is determined based on a size of a bounding box surrounding the at least one image frame and one or more extrinsic parameters of an image capturing device used to capture the at least one image frame, wherein the size of bounding box and the one or more extrinsic parameters of the image capturing device is used as part of labeling data of the at least one image frame, and wherein the one or more extrinsic parameters include a relative rotation and translation between cameras in a multi-camera arrangement,

determining that an emergency vehicle is present in the surrounding environment in response to at least one of the first probability of the first neural network model exceeds a first predefined threshold or the second probability of the second neural network model exceeds a second predefined threshold,

determining a position of the emergency vehicle and a moving direction of the emergency vehicle by fusing the first detection result of the first neural network model and the second detection result of the second neural network model,

determining whether the emergency vehicle is moving towards the ADV based on the position of the emergency vehicle and the moving direction of the emergency vehicle, and

controlling the ADV by steering out of a current driving lane, braking to decelerate, or steering to a side of a road, in response to determining that the emergency vehicle is moving towards the ADV.

14. The system of claim 13 , wherein the operations further comprise:

determining, using the first neural network model, an angle between the ADV and the source of the siren sound, and the moving direction of the source.

15. The system of claim 13 ,

wherein the controlling the ADV by steering out of the current driving lane, braking to decelerate, or steering to the side of the road comprises:

controlling the ADV, by steering out of the current driving lane or braking to decelerate.

16. The system of claim 13 , wherein the first neural network model is trained with audio data representing emergency vehicle siren collected from a plurality of emergency vehicles, and wherein the second neural network model is trained with visual data collected simultaneously to the collecting of the audio data.

17. The system of claim 13 , wherein each of the first neural network model and the second neural network model is a convolutional neural network.

18. The system of claim 13 , wherein the one or more audio capturing devices include one or more microphones, and wherein the one or more image capturing devices include one or more cameras.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2021
From: XU, KECHENG; SUN, HONGYI; LUO, QI; WANG, WEI; LIN, ZEJUN; REYNOLDS, WESLEY; LIU, FENG; HU, JIANGTAO; MIAO, JINGHAO
To: BAIDU USA LLC
Reel/Frame 054927/0541 →
Continuity (1)
Related Publication 20220219736A1 · Jul 14, 2022
References Cited (39)
US 11282385B2 · Lewis et al. · 2022 [cited by applicant]
US 11501532B2 · Gan et al. · 2022 [cited by applicant]
US 20050044053A1 · Moreno et al. · 2005 [cited by applicant]
US 20170249839A1 · Becker · 2017 [cited by examiner]
US 20180137756A1 · Moosaei · 2018 [cited by examiner]
US 20180189572A1 · Hori et al. · 2018 [cited by applicant]
US 20180233047A1 · Mandeville-Clarke · 2018 [cited by examiner]
US 20180364732A1 · Yaldo et al. · 2018 [cited by applicant]
US 20180374347A1 · Silver et al. · 2018 [cited by applicant]
US 20190027032A1 · Arunachalam · 2019 [cited by applicant]
US 20190049989A1 · Akotkar et al. · 2019 [cited by applicant]
US 20190163982A1 · Block · 2019 [cited by applicant]
US 20200089253A1 · Sudo · 2020 [cited by examiner]
US 20200265273A1 · Wei et al. · 2020 [cited by applicant]
US 20210150230A1 · Smolyanskiy · 2021 [cited by examiner]
US 20210200803A1 · Zhang et al. · 2021 [cited by applicant]
US 20210247201A1 · Hori et al. · 2021 [cited by applicant]
US 20210248183A1 · Pratt et al. · 2021 [cited by applicant]
US 20210358513A1 · Narisetty et al. · 2021 [cited by applicant]
US 20220067479A1 · Lee et al. · 2022 [cited by applicant]
US 20220093101A1 · Metallinou et al. · 2022 [cited by applicant]
US 20220101629A1 · Liu et al. · 2022 [cited by applicant]
US 20220121868A1 · Chen et al. · 2022 [cited by applicant]
US 20220141503A1 · Cui et al. · 2022 [cited by applicant]
US 20220147602A1 · Streit · 2022 [cited by applicant]
US 20220147607A1 · Streit · 2022 [cited by applicant]
US 20220150068A1 · Streit · 2022 [cited by applicant]
US 20220223037A1 · Xu et al. · 2022 [cited by applicant]
US 20220292809A1 · Choudhary et al. · 2022 [cited by applicant]
US 20220351348A1 · Chae et al. · 2022 [cited by applicant]
US 20220351439A1 · Chae et al. · 2022 [cited by applicant]
US 20220355814A1 · Sharifi · 2022 [cited by examiner]
US 20220358703A1 · Chae et al. · 2022 [cited by applicant]
Aparajit Garg et al, “Emergency Vehicle Detection by Autonomous Vehicle,” International Journal of Engineering Research and Technology (IJERT), May 1, 2019, 6 Pages. [cited by applicant]
Abhishek Raman et al, “A Hybrid Framework for Expediting Emergency Vehicle Movement on Indian Roads,” ICIMIA, Mar. 5, 2020, 6 Pages. [cited by applicant]
Hongyi Sun et al, “Emergency Vehicles Audio Detection and Localization in Autonomous Driving,” Arxiv.org, Cornell University library, Oct. 2, 2021, 6 Pages. [cited by applicant]
Van-Thuan Tran and Wei-Ho Tsai, “Audio-Vision Emergency Vehicle Detection,” IEEE sensors Journal, Nov. 15, 2021, 14 Pages. [cited by applicant]
Tran, Van-Thuan and Tsai, Wei-Ho, “Acousitc Based Emergency Vehicle Detection Using Convolutional Neural Networks,” IEEE Access, Apr. 18, 2020, 12 pages. [cited by applicant]
Selbes, Berkay and Sert, Mustafa, “Multimodal Vehicle Type Classification Using Convolutional Neural Network and Statistical Representations of MFCC,” IEEE Aug. 29, 2017, 6 pages. [cited by applicant]