IP Library Granted Patent US 11,567,510
Granted Patent B2
US 11,567,510 · App. 16/752,594 · Granted Jan 31, 2023

Using classified sounds and localized sound sources to operate an autonomous vehicle

Inventors: Metarsit Leenayongwut (Singapore, SG); Weng Fei Low (Singapore, SG)
Assignee: Motional AD LLC
G05D1/0255G05D1/0088G05D1/0214G06N20/00G10L25/18G10L25/51H04R1/406H04R3/005B60R11/0247H04R2499/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,567,510
App. No.
16/752,594
Granted
Jan 31, 2023
Kind
B2
Abstract

An ambient sound environment is captured by a microphone array of an autonomous vehicle traveling in the ambient sound environment. A perception module of the autonomous vehicle classifies sounds and localizes sound sources in the ambient sound environment. Classification is performed using spectrum analysis and/or machine learning. In an embodiment, sound sources within a field of view (FOV) of an image sensor of the autonomous vehicle are localized in a visual scene generated by the perception module. In an embodiment, one or more sound sources outside the FOV of the image sensors are localized in a static digital map. Localization is performed using parametric or non-parametric techniques and/or machine learning. The output of the perception module is input into a planning module of the autonomous vehicle to plan a route or trajectory for the autonomous vehicle in the ambient sound environment.

Claims (71)

1. An autonomous vehicle (AV), comprising:

a plurality of microphones;

a processing-circuit that performs operations including:

capturing, using the plurality of microphones, an ambient sound environment in which the AV is operating;

classifying, based on the captured ambient sound environment, a sound in the ambient sound environment, wherein the classifying includes:

determining a frequency spectrum of the sound;

matching the frequency spectrum to a reference frequency spectrum;

determining, based on the matching, whether the frequency spectrum of the sound has one or more tones; and

in accordance with the frequency spectrum of the sound having one or more tones, classifying the sound as a horn;

determining, based on the frequency spectrum of the sound, whether the sound sweeps between two or more frequencies; and

in accordance with determining that the sound sweeps between two or more frequencies, classifying the sound as a siren;

determining, based on the sound, a location of a source of the sound in the ambient sound environment; and

causing, using the processing-circuit, the AV to perform an action based on the classified sound and the determined location of the sound source in the ambient sound environment.

2. The AV of claim 1 , wherein the frequency spectrum is determined using a short-term Fourier Transform (STFT).

3. The AV of claim 1 , wherein the operations further comprise:

processing, using a machine learning-circuit, the sound; and

classifying, based on output of the machine learning circuit, the sound.

4. The AV of claim 3 , wherein the sound source is classified as a platoon of vehicles and the operations further comprise:

planning a route or trajectory for the AV to travel in the ambient sound environment to avoid the platoon of vehicles; and

operating, using a controller-circuit, the AV to travel the route or trajectory.

5. The AV of claim 3 , wherein the sound is classified as a construction zone sound and the operations further comprise:

planning a route or trajectory for the AV to travel in the ambient sound environment to avoid the construction zone; and

operating, using a controller-circuit, the AV to travel the route or trajectory.

6. The AV of claim 3 , wherein the sound is classified as a vehicle operation sound and the operations further comprise:

determining, using the processing circuit, a type of the sound source based on the vehicle operation sound; and

generating, using a perception-circuit, a bounding box in a vision scene for the sound source based on the type of the sound.

7. The AV of claim 3 , wherein the sound is classified as a pedestrian sound and the operations further comprise:

generating, using a perception-circuit, a bounding box in a vision scene for the pedestrian.

8. The AV of claim 3 , wherein the sound is classified as a vehicle operation sound and the operations further comprise:

determining, using the processing-circuit, a state of the sound source based on the vehicle operation sound; and

operating, using a controller-circuit, the AV based on the state of the sound source.

9. The AV of claim 1 , wherein the location is determined by a estimating a direction of arrival (DOA) relative to the plurality of microphones and a distance between the sound source and the plurality of microphones.

10. The AV of claim 9 , wherein the DOA is estimated by beamforming at least two of the plurality of microphones.

11. The AV of claim 9 , wherein the DOA is estimated using a narrowband multiple signal classification (MUSIC) algorithm.

12. The AV of claim 1 , wherein the sound source is a siren or horn of an emergency vehicle and the action is maneuvering or stopping the AV to allow the emergency vehicle to pass the AV.

13. The AV of claim 1 , wherein the location of the sound source is in a vision scene output by a perception module of the AV.

14. A method comprising:

capturing, using a plurality of microphones, an ambient sound environment in which an autonomous vehicle (AV) is operating;

classifying, using a processing circuit, a sound in the ambient sound environment, wherein the classifying includes:

determining a frequency spectrum of the sound;

matching the frequency spectrum to a reference frequency spectrum;

determining, based on the matching, whether the frequency spectrum of the sound has one or more tones;

in accordance with the frequency spectrum of the sound having one or more tones, classifying the sound as a horn;

determining, based on the frequency spectrum of the sound, whether the sound sweeps between two or more frequencies;

in accordance with determining that the sound sweeps between two or more frequencies, classifying the sound as a siren;

determining, using the processing-circuit, a location of a sound source in the ambient sound environment based on the sound; and

causing, using the processing-circuit, the AV to perform an action based on the classified sound and the determined location of the sound source in the ambient sound environment.

15. The method of claim 14 , wherein the frequency spectrum is determined using a short-term Fourier Transform (STFT).

16. The method of claim 14 wherein the operations further comprise:

processing, using a machine learning-circuit, the sound; and

classifying, based on output of the machine learning circuit, the sound.

17. The method of claim 16 , wherein the sound is classified as a platoon of vehicles and the operations further comprise:

planning a route or trajectory for the AV to travel in the ambient sound environment to avoid the platoon of vehicles; and

operating, using a controller-circuit, the AV to travel the route or trajectory.

18. The method of claim 16 , wherein the sound is classified as a construction zone sound and the operations further comprise:

planning a route or trajectory for the AV to travel in the ambient sound environment to avoid the construction zone; and

operating, using a controller-circuit, the AV to travel the route or trajectory.

19. The method of claim 16 , wherein the sound is classified as a vehicle operation sound and the operations further comprise:

determining, using the processing circuit, a type of the sound source based on the vehicle operation sound; and

generating, using a perception-circuit, a bounding box for the sound source based on the type of the sound.

20. The method of claim 16 , wherein the sound is classified as a pedestrian sound and the operations further comprise:

generating, using a perception-circuit, a bounding box for the pedestrian.

21. The method of claim 16 , wherein the sound is classified as a vehicle operation sound and the operations further comprise:

determining, using the processing-circuit, a state of the sound source based on the vehicle operation sound; and

operating, using a controller-circuit, the AV based on the state of the sound source.

22. The method of claim 14 , wherein the location is determined by a estimating a direction of arrival (DOA) relative to the plurality of microphones and a distance between the sound source and the plurality of microphones.

23. The method of claim 22 , wherein the DOA is estimated by beamforming at least two of the plurality of microphones.

24. The method of claim 22 , wherein the DOA is estimated using a narrowband multiple signal classification (MUSIC) algorithm.

25. The method of claim 14 , wherein the sound source is a siren or horn of an emergency vehicle and the action is maneuvering or stopping the AV to allow the emergency vehicle to pass the AV.

26. The method of claim 14 , wherein the location of the sound source is in a vision scene output by a perception module of the AV.

27. A non-transitory, computer-readable storage medium having instructions stored thereon, that when executed by one or more processors, cause the one or more processors to perform the method of claim 14 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2020
From: APTIV TECHNOLOGIES LIMITED
To: MOTIONAL AD LLC
Reel/Frame 053863/0746 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2020
From: LEENAYONGWUT, METARSIT; LOW, WENG FEI
To: APTIV TECHNOLOGIES LIMITED
Reel/Frame 053133/0771 →
Continuity (2)
Provisional Application 62796513 · Jan 24, 2019
Related Publication 20200241552A1 · Jul 30, 2020
Cited By (2)
US 12,344,249 US 12,406,577