IP Library Granted Patent US 12,592,237
Granted Patent B2
US 12,592,237 · App. 17/547,917 · Granted Mar 31, 2026

Driver interface with voice and image control

Inventors: Zili Li (San Jose, CA); Cristina Vasconcelos (Toronto, CA)
Assignee: SoundHound AI IP, LLC
G10L15/24G06F18/217G06V10/764G06V10/768G06V10/82G06V20/46G10L15/02G10L15/063G10L15/16G10L15/1822G10L15/187G10L15/22G10L15/30G10L2015/025G10L2015/223G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,237
App. No.
17/547,917
Granted
Mar 31, 2026
Kind
B2
Abstract

A driver interface for use within an automobile provides responses to voice commands issued for example by a driver of the automobile. The interface includes a camera and microphone for capturing image data such as gestures and audio data from the automobile driver. The image data and audio data are processed to extract image and linguistic features from the image and audio data, which image and linguistic features are processed to interpret and infer a meaning of the voice command.

Claims (39)

1 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

an image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual features, the visual features being an abstraction of the captured images, the abstraction being a compressed representation of the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, linguistic features;

a natural language processor enabled to infer a voice command from the linguistic features and the visual features; and

an output interface for rendering a response to the voice command; and

wherein the speech recognizer uses a machine-learned linguistic model based on a combination of audio feature tensors and visual feature tensors.

2 . The driver interface system of claim 1 wherein inferring the voice command from the linguistic features and the visual features is configured by a set of acoustic model configurations parameters.

3 . The driver interface system of claim 1 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the speech recognizer extracts linguistic features by making an application programming interface (API) call to the server through the data interface.

4 . The driver interface system of claim 3 wherein the data interface is a wireless network interface.

5 . The driver interface system of claim 1 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the natural language processor infers the voice command by making an application programming interface (API) call to the server through the data interface.

6 . The driver interface system of claim 5 wherein the data interface is a wireless network interface.

7 . The driver interface system of claim 1 wherein the visual features indicate a gesture.

8 . The driver interface system of claim 1 wherein the response is rendered by projection onto a display screen.

9 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

a neural network image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual features, the visual features being an abstraction of the captured images determined by the neural network image processor, the abstraction being a compressed representation of the captured images with reduced information content as compared to the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, linguistic features; and

a natural language processor enabled to infer a voice command from the linguistic features and the visual features, including visual features absracted from an environment surrounding the driver;

wherein the image processor, the speech recognizer, and the natural language processor are jointly configured by joint training, such that the extraction of the visual feature tensors and audio feature tensors is optimized in combination with the natural language processor to minimize errors in inferring the voice command.

10 . The driver interface system of claim 9 wherein inferring the voice command from the linguistic features and the visual features is configured by a set of acoustic model configurations parameters.

11 . The driver interface system of claim 9 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the speech recognizer extracts linguistic features by making an application programming interface (API) call to the server through the data interface.

12 . The driver interface system of claim 11 wherein the data interface is a wireless network interface.

13 . The driver interface system of claim 9 further comprising a data interface enabled to transmit visual features and captured speech to a server, wherein the natural language processor infers the voice command by making an application programming interface (API) call to the server through the data interface.

14 . The driver interface system of claim 13 wherein the data interface is a wireless network interface.

15 . The driver interface system of claim 9 wherein the speech recognizer uses a machine-learned linguistic model based on a combination of audio feature tensors and visual feature tensors.

16 . The driver interface system of claim 9 wherein the visual features indicate a gesture.

17 . The driver interface system of claim 9 further comprising an output interface for rendering a response to the voice command.

18 . The driver interface system of claim 17 wherein the response is rendered by projection onto a display screen.

19 . A driver interface system comprising:

a camera interface enabled to capture images of a driver and an environment surrounding the driver;

an image processor enabled to extract, from captured images of the driver and an environment surrounding the driver, visual feature tensors, the visual feature tensors being an abstraction of the captured images, the abstraction being a compressed representation of the captured images with reduced information content as compared to the captured images;

a microphone interface enabled to capture speech;

a speech recognizer enabled to extract, from the captured speech, audio feature tensors;

a natural language processor enabled to concatonate or merge the visual feature vectors and audio feature vectors to infer a voice command;

wherein the image processor, the speech recognizer, and the natural language processor are jointly configured by joint training, such that the extraction of the visual feature tensors and audio feature tensors is optimized in combination with the natural language processor to minimize errors in inferring the voice command; and

an output interface for rendering a response to the voice command.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: LI, ZILI; VASCONCELOS, CRISTINA
To: SOUNDHOUND, INC.
Reel/Frame 058362/0353 →