IP Library Granted Patent US 12,441,239
Granted Patent B1
US 12,441,239 · App. 18/428,924 · Granted Oct 14, 2025

Autonomous vehicle sound determination by a model

Inventors: Seyed Parsa Beheshti (Belmont, CA); William Randolph Geddings, Jr. (San Francisco, CA); Anoop Jatavallabha Vijayakumar (Fremont, CA); Shaminda Subasingha (San Ramon, CA)
Assignee: Zoox, Inc.
B60Q5/006B60Q5/008G10K11/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,441,239
App. No.
18/428,924
Granted
Oct 14, 2025
Kind
B1
Abstract

Techniques for determining audio data for output by an autonomous vehicle are discussed herein. A system can determine audio data representing a frequency and a sound level for different object types in an environment. The audio data can be specific for a human or animal, and can be adjusted based on a noise level in the environment, weather, ambient light, or other criteria. A model may determine audio data for continuous output in a vicinity of a specific object or region in the environment. A sound level of the audio data may be adjusted over time base on a likelihood an intersection between a vehicle path and an object path.

Claims (71)

1. A method comprising:

receiving first data associated with a vehicle in an environment;

determining, based at least in part on the first data, a location associated with a first object that causes an occluded region in the environment relative to the vehicle;

determining an object type of the first object or a likelihood of a second object emerging from the occluded region;

determining ambient sound in the environment;

determining, based at least in part on the location, the ambient sound, and at least one of: the object type of the first object or the likelihood of the second object emerging from the occluded region, audio data including a frequency and a sound level for output by the vehicle; and

outputting, by a speaker of the vehicle, the audio data over a time period to notify the first object or the second object in the occluded region of the vehicle's presence.

2. The method of claim 1 , wherein:

the sound level is a first sound level, and

determining the audio data is further based at least in part on one or more of: weather in the environment, a sub-type of the first object or the second object, ambient light in the environment, a level of attention by the first object, a distance between the vehicle and one of the first object or the occluded region, a second sound level of the environment, or presence of a school zone.

3. The method of claim 2 , further comprising:

receiving, from a first machine learned model, a predicted intersection between the vehicle and the first object;

inputting, into a second machine learned model, the predicted intersection and the first data representing audio data in the environment; and

receiving, from the second machine learned model, a sound profile comprising the frequency and the sound level for output by the vehicle.

4. The method of claim 1 , wherein the first data represents one or more of: sensor data from a sensor coupled to the vehicle, log data, map data, or weather data.

5. The method of claim 1 , further comprising:

determining a first location of the first object or a second location associated with the occluded region relative to the vehicle,

wherein outputting the audio data comprises targeting the audio data towards the first object or the second object based at least in part on the first location or the second location.

6. The method of claim 5 , further comprising:

determining a portion of the occluded region from which the second object is likely to exit;

wherein the second location represents the portion of the occluded region from which the second object is likely to exit.

7. The method of claim 1 , further comprising:

determining, based at least in part on a change in position by the first object or the second object relative to the vehicle, has increased from a first time to a second time after the first time; and

modifying the sound level of the audio data from the first time to the second time based at least in part on the change in the position by the first object or the second object.

8. The method of claim 1 , further comprising:

determining whether the first object detects presence of the vehicle; and

determining the audio data based at least in part on whether the first object recognizing or acknowledging presence of the vehicle.

9. The method of claim 1 , further comprising:

determining a likelihood of an intersection between a first trajectory associated with the vehicle and a second trajectory associated with the first object; and

determining the audio data based at least in part on the likelihood.

10. The method of claim 1 , further comprising:

identifying an intensity or a frequency of sound associated with the environment; and

determining the audio data based at least in part on the intensity or the frequency of sound associated with the environment.

11. The method of claim 1 , wherein the frequency of the audio data is above a human threshold of hearing.

12. One or more non transitory computer readable media storing instructions executable by a processor, wherein the instructions, when executed, cause the processor to perform operations comprising:

receiving first data associated with a vehicle in an environment;

determining, based at least in part on the first data, a location associated with a first object that causes an occluded region in the environment relative to the vehicle;

determining an object type of the first object or a likelihood of a second object emerging from the occluded region;

determining ambient sound in the environment;

determining, based at least in part on the location and at least one of: the object type of the first object or the likelihood of the second object emerging from the occluded region, audio data including a frequency and a sound level for output by the vehicle; and

outputting, by a speaker of the vehicle, the audio data over a time period to notify the first object or the second object in the occluded region of the vehicle's presence.

13. The one or more non transitory computer readable media of claim 12 , wherein:

the sound level is a first sound level, and

determining the audio data is further based at least in part on one or more of: weather in the environment, a sub-type of the first object or the second object, ambient light in the environment, a level of attention by the first object, a distance between the vehicle and one of the first object or the occluded region, a second sound level of the environment, or presence of a school zone.

14. The one or more non transitory computer readable media of claim 12 , the operations further comprising:

emitting a sound beam using a plurality of speakers in a direction towards the first object or the occluded region.

15. The one or more non transitory computer readable media of claim 12 , the operations further comprising:

determining, based at least in part on a change in position by the first object or the second object relative to the vehicle, has increased from a first time to a second time after the first time; and

increasing the sound level of the audio data from the first time to the second time based at least in part on the change in the position by the first object or the second object.

16. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising:

receiving first data associated with a vehicle in an environment;

determining, based at least in part on the first data, a location associated with a first object that causes an occluded region in the environment relative to the vehicle;

determining an object type of the first object or a likelihood of a second object emerging from the occluded region;

determining ambient sound in the environment;

determining, based at least in part on the location, the ambient sound, and at least one of: the object type of the first object or the likelihood of the second object emerging from the occluded region, audio data including a frequency and a sound level for output by the vehicle; and

outputting, by a speaker of the vehicle, the audio data over a time period to notify the first object or the second object in the occluded region of the vehicle's presence.

17. The system of claim 16 , wherein:

the sound level is a first sound level, and

determining the audio data is further based at least in part on one or more of: weather in the environment, a sub-type of the first object or the second object, ambient light in the environment, a level of attention by the first object, a distance between the vehicle and one of the first object or the occluded region, a second sound level of the environment, or presence of a school zone.

18. The system of claim 16 , the operations further comprising:

receiving, from a first machine learned model, a predicted intersection between the vehicle and the first object;

inputting, into a second machine learned model, the predicted intersection and the first data representing audio data in the environment; and

receiving, from the second machine learned model, a sound profile comprising the frequency and the sound level for output by the vehicle.

19. The system of claim 16 , the operations further comprising:

determining a first location of the first object or a second location associated with the occluded region relative to the vehicle,

wherein outputting the audio data comprises targeting the audio data towards the first object or the second object based at least in part on the first location or the second location.

20. The system of claim 16 , the operations further comprising:

determining, based at least in part on a change in position by the first object or the second object relative to the vehicle, has increased from a first time to a second time after the first time; and

modifying the sound level of the audio data from the first time to the second time based at least in part on the change in the position by the first object or the second object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: BEHESHTI, SEYED PARSA; GEDDINGS, JR., WILLIAM RANDOLPH; JATAVALLABHA VIJAYAKUMAR, ANOOP; SUBASINGHA, SHAMINDA
To: ZOOX, INC.
Reel/Frame 066323/0431 →
References Cited (19)
US 8031085B1 · Anderson · 2011 [cited by examiner]
US 9878664B2 · Kentley-Klay et al. · 2018 [cited by applicant]
US 9981602B2 · Vincent · 2018 [cited by examiner]
US 10261514B2 · Zych · 2019 [cited by examiner]
US 10414336B1 · Harper · 2019 [cited by examiner]
US 10497264B2 · Rowell · 2019 [cited by examiner]
US 10547941B1 · Herman · 2020 [cited by examiner]
US 11016492B2 · Gier et al. · 2021 [cited by applicant]
US 11027648B2 · Harper · 2021 [cited by examiner]
US 11458891B1 · Kuehner · 2022 [cited by examiner]
US 11488472B2 · Parenti · 2022 [cited by examiner]
US 11500378B2 · Kentley-Klay · 2022 [cited by examiner]
US 20160362045A1 · Vegt · 2016 [cited by examiner]
US 20170222612A1 · Zollner · 2017 [cited by examiner]
US 20180290590A1 · Goldman-Shenhar · 2018 [cited by examiner]
US 20190329794A1 · Kim · 2019 [cited by examiner]
US 20200156538A1 · Harper · 2020 [cited by examiner]
US 20210245742A1 · Ha · 2021 [cited by examiner]
US 20220185267A1 · Beller et al. · 2022 [cited by applicant]