IP Library Granted Patent US 12679206
Granted Patent B2
US 12679206 · App. 18/655,766 · Granted Jul 14, 2026

Visualization of external audio commands

Inventors: Joseph Whinnery (Soquel, CA); Venkata Subrahmanyam Chandra Sekhar Chebiyyam (Mountain View, CA)
Assignee: Zoox, Inc.
B60K35/22G10L15/26G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12679206
App. No.
18/655,766
Granted
Jul 14, 2026
Kind
B2
Abstract

This disclosure relates to methods, systems, and techniques for visualizing external audio captured by a microphone of the vehicle. Using the techniques described herein, external audio signals may be interpreted and converted into clear visual indicium that can be understood by individuals associated with the vehicle, such as passengers inside the vehicle or individuals awaiting pickup.

Claims (78)

1 . A system comprising,

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising:

recording audio external to a vehicle traversing an environment using a microphone associated with the vehicle, the vehicle an autonomous vehicle;

determining a source of the audio in the environment using one or more sensors associated with the vehicle;

determining that the source is an authenticated source for external communication with the vehicle using at least one characteristic of the source;

analyzing the audio to determine a command from the source of the audio directed towards the vehicle or a passenger of the vehicle, the command comprising an instruction for a passenger of the vehicle to follow; and

causing control of the vehicle based at least in part on the command.

2 . The system of claim 1 , wherein the instructions further cause the system to perform actions comprising:

transmitting audio data derived from the audio to a computing device external to the vehicle, the computing device comprising one or more processors;

determining the command at least in part by applying, by the one or more processors of at the computing device, a speech-to-text algorithm to the audio data; and

transmitting the command to the vehicle.

3 . The system of claim 2 , wherein the instructions further cause the system to perform actions comprising:

displaying the command to a remote operator associated with the vehicle;

receiving input from the remote operator to update the command; and

transmitting the updated command to the vehicle.

4 . The system of claim 1 , wherein the instructions further cause the system to perform actions comprising:

determining the command as relating to an area in the environment which the passenger of the vehicle should avoid or proceed to upon exiting the vehicle;

generating a visual representation indicating the area in relation to a current position of the vehicle in the environment; and

displaying the visual representation on a display in the vehicle.

5 . The system of claim 1 , wherein the instructions further cause the system to perform actions comprising:

generating, based at least in part on the command, a textual description of the command; and

displaying the textual description on a display in the vehicle.

6 . A method comprising:

recording audio using a microphone associated with a vehicle traversing an environment and wherein the audio is external to the vehicle;

determining, based at least in part on the audio, that the audio is associated with the vehicle or an occupant of the vehicle;

generating audio data based at least in part on the audio;

determining, based at least in part on the audio data, a command directed towards the vehicle or a passenger of the vehicle, the command comprising an instruction for the vehicle or the passenger of the vehicle to follow; and

causing control of the vehicle based at least in part on the command.

7 . The method of claim 6 , further comprising:

determining, based at least in part on the audio data, an indicium for the passenger; and

at least one of:

providing the indicium on a display within the vehicle for a passenger of the vehicle; or

providing the indicium on a display of a handheld device associated with the vehicle.

8 . The method of claim 7 , further comprising:

generating, based at least in part on the command, a visual representation including the indicium associated with the command.

9 . The method of claim 8 , further comprising:

transmitting the audio data to a computing device external to the vehicle, the computing device comprising one or more processors;

determining the command at least in part by applying, by the one or more processors of at the computing device, a speech-to-text algorithm to the audio data; and

transmitting the command to the vehicle.

10 . The method of claim 9 , further comprising:

displaying the command to a remote operator associated with the vehicle;

receiving input from the remote operator to update the command; and

transmitting the updated command to the vehicle.

11 . The method of claim 10 , further comprising:

playing back at least a portion of the audio data to the remote operator.

12 . The method of claim 8 , further comprising:

determining the command as relating to an area in the environment which a passenger of the vehicle should avoid or proceed to upon exiting the vehicle; and

generating the visual representation indicating the area in relation to a current position of the vehicle in the environment.

13 . The method of claim 8 , further comprising:

generating, based at least in part on the command, a textual description of the command; and

displaying the textual description on a display associated with the vehicle for the individual.

14 . The method of claim 6 , further comprising:

determining a source of the audio in the environment using one or more sensors associated with the vehicle; and

determining that the source is an authenticated source for external communication with the vehicle using at least one characteristic of the source.

15 . One or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations comprising:

recording audio using a microphone associated with a vehicle traversing an environment and wherein the audio is external to the vehicle;

determining, based at least in part on the audio, that the audio is associated with the vehicle or an occupant of the vehicle;

generating audio data based at least in part on the audio;

determining, based at least in part on the audio data, a command directed towards the vehicle or a passenger of the vehicle, the command comprising an instruction for the vehicle or the passenger of the vehicle to follow; and

causing control of the vehicle based at least in part on the command.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein the operations further comprise:

determining, based at least in part on the audio data, an indicium for the passenger; and

at least one of:

providing the indicium on a display within the vehicle for a passenger of the vehicle; or

providing the indicium on a display of a handheld device associated with the vehicle.

17 . The one or more non-transitory computer-readable media of claim 16 , wherein the operations further comprise:

generating, based at least in part on the command, a visual representation including the indicium associated with the command.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein the operations further comprise:

transmitting the audio data to a computing device external to the vehicle, the computing device comprising one or more processors;

determining the command at least in part by applying, by the one or more processors of the computing device, a speech-to-text algorithm to the audio data; and

transmitting the command to the vehicle.

19 . The one or more non-transitory computer-readable media of claim 17 , wherein the operations further comprise:

determining the command as relating to an area in the environment which a passenger of the vehicle should avoid or proceed to upon exiting the vehicle; and

generating the visual representation indicating the area in relation to a current position of the vehicle in the environment.

20 . The one or more non-transitory computer-readable media of claim 15 , wherein the operations further comprise:

determining a source of the audio in the environment using one or more sensors associated with the vehicle; and

determining that the source is an authenticated source for external communication with the vehicle using at least one characteristic of the source.