IP Library Granted Patent US 10,891,949
Granted Patent B2
US 10,891,949 · App. 16/125,944 · Granted Jan 12, 2021

Vehicle language processing

Inventors: Praveen Narayanan (San Jose, CA); Lisa Scaria (Milpitas, CA); Ryan Burke (Palo Alto, CA); Francois Charette (Tracy, CA); Punarjay Chakravarty (Mountain View, CA); Kaushik Balakrishnan (Mountain View, CA)
Assignee: FORD GLOBAL TECHNOLOGIES, LLC
G10L15/22B60W10/04B60W10/18B60W10/20B60W50/10G06N3/04G06N3/08G10L15/16G10L15/30G10L25/21B60W2050/007B60W2540/21G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,949
App. No.
16/125,944
Granted
Jan 12, 2021
Kind
B2
Abstract

A computing system can be programmed to receive a spoken language command in response to emitting a spoken language cue and process the spoken language command with a generalized adversarial neural network (GAN) to determine a vehicle command. The computing system can be further programmed to operate a vehicle based on the vehicle command.

Claims (35)

1. A method, comprising:

receiving a spoken language command in response to emitting a spoken language cue;

training a neural network (NN) at least in part by inputting mel-frequency data from the spoken language command combined with a loss function to a convolutional filter bank that outputs modified mel-frequency data to a bidirectional long short-term memory that in turn outputs audio spectrum data to determine a vehicle command using a generalized adversarial neural network (GAN) that determines a loss function by a combination of a binary classification of the output audio spectrum data as one of real or fake and a ground truth-based loss;

processing the spoken language command with the NN to determine the vehicle command; and

operating a vehicle based on the vehicle command.

2. The method of claim 1 , further comprising transforming the spoken language command into mel frequency samples before processing with the NN.

3. The method of claim 2 , wherein the mel frequency samples are each compressed into a single vector by convolving mel values along a y-axis before processing with the NN.

4. The method of claim 3 , wherein a mel frequency scale is a log power spectrum of spoken language command frequencies on a nonlinear scale of frequencies.

5. The method of claim 1 , further comprising training the GAN to determine real or fake spoken language using a plurality of recorded spoken language commands, ground truth that identifies the recorded spoken language commands as real or fake and a loss function based on ground truth.

6. The method of claim 1 , wherein operating the vehicle includes determining a path polynomial based on the vehicle command.

7. The method of claim 1 , wherein operating the vehicle includes determining a cognitive map based on vehicle sensor data.

8. The method of claim 1 , further comprising processing synthetic language data with a GAN to determine the spoken language cue.

9. The method of claim 1 , wherein the vehicle command is a request for goal-directed behavior of the vehicle.

10. A system, comprising a processor; and

a memory, programmed to:

receive a spoken language command in response to emitting a spoken language cue;

train a neural network (NN) at least in part by inputting mel-frequency data from the spoken language command combined with a loss function to a convolutional filter bank that outputs modified mel-frequency data to a bidirectional long short-term memory that in turn outputs audio spectrum data to determine a vehicle command using a generalized adversarial neural network (GAN) that determines a loss function by a combination of a binary classification of the output audio spectrum data as one of real or fake and a ground truth-based loss;

process the spoken language command with the NN to determine the vehicle command; and

operate a vehicle based on the vehicle command.

11. The system of claim 10 , further comprising transforming the spoken language command into mel frequency samples before processing with the NN.

12. The system of claim 11 , wherein the mel frequency samples are each compressed into a single vector by convolving mel values along a y-axis before processing with the NN.

13. The system of claim 12 , wherein a mel frequency scale is a log power spectrum of spoken language command frequencies on a nonlinear scale of frequencies.

14. The system of claim 10 , further comprising training the GAN to determine real or fake spoken language using a plurality of recorded spoken language commands, ground truth that identifies the recorded spoken language commands as real or fake and a loss function based on ground truth.

15. The system of claim 10 , wherein operating the vehicle includes determining a path polynomial based on the vehicle command.

16. The system of claim 10 , wherein operating the vehicle includes determining a cognitive map based on vehicle sensor data.

17. The system of claim 10 , further comprising processing synthetic language data with a GAN to determine the spoken language cue.

18. The system of claim 10 , wherein the vehicle command is a request for goal-directed behavior of the vehicle.

19. A system, comprising:

means for controlling second vehicle steering, braking and powertrain;

computer means for:

receiving a spoken language command in response to emitting a spoken language cue;

training a neural network (NN) at least in part by inputting mel-frequency data from the spoken language command combined with a loss function to a convolutional filter bank that outputs modified mel-frequency data to a bidirectional long short-term memory that in turn outputs audio spectrum data to determine a vehicle command using a generalized adversarial neural network (GAN) that determines a loss function by a combination of a binary classification of the output audio spectrum data as one of real or fake and a ground truth-based loss;

processing the spoken language command with the NN to determine the vehicle command; and

operating a vehicle based on the vehicle command and the means for controlling vehicle steering, braking and powertrain.

20. The system of claim 19 , further comprising transforming the spoken language command into a mel frequency samples before processing with the NN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2018
From: NARAYANAN, PRAVEEN; SCARIA, LISA; BURKE, RYAN; CHARETTE, FRANCOIS; CHAKRAVARTY, PUNARJAY; BALAKRISHNAN, KAUSHIK
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 046824/0900 →
Continuity (1)
Related Publication 20200082817A1 · Mar 12, 2020
Cited By (1)
US 12,711,947