IP Library › Granted Patent US 11,361,768
Granted Patent B2
US 11,361,768 · App. 16/935,112 · Granted Jun 14, 2022

Utterance classifier

Inventors: Nathan David Howard (Mountain View, CA); Gabor Simko (Santa Clara, CA); Maria Carolina Parada San Martin (Boulder, CO); Ramkarthik Kalyanasundaram (Cupertino, CA); Guru Prakash Arumugam (Sunnyvale, CA); Srinivas Vasudevan (Mountain View, CA)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/16G10L15/18G10L15/30G10L17/00G10L2015/223G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,361,768
App. No.
16/935,112
Filed
Jul 21, 2020
Granted
Jun 14, 2022
Kind
B2
Examiner
HAN, QI
Art Unit
2659
USPC
704/232
Abstract

A method includes receiving a spoken utterance that includes a plurality of words, and generating, using a neural network-based utterance classifier comprising a stack of multiple Long-Short Term Memory (LSTM) layers, a respective textual representation for each word of the of the plurality of words of the spoken utterance. The neural network-based utterance classifier trained on negative training examples of spoken utterances not directed toward an automated assistant server. The method further including determining, using the respective textual representation generated for each word of the plurality of words of the spoken utterance, that the spoken utterance is one of directed toward the automated assistant server or not directed toward the automated assistant server, and when the spoken utterance is directed toward the automated assistant server, generating instructions that cause the automated assistant server to generate a response to the spoken utterance.

Claims (36)

1. A method comprising:

receiving, at data processing hardware, a spoken utterance captured by an automated assistant device associated with a user, the spoken utterance comprising a plurality of words;

generating, by the data processing hardware, using a neural network-based utterance classifier comprising a stack of multiple Long-Short Term Memory (LSTM) layers, a respective textual representation for each word of the of the plurality of words of the spoken utterance, the neural network-based utterance classifier trained on negative training examples of spoken utterances not directed toward an automated assistant server;

determining, by the data processing hardware, using the respective textual representation generated for each word of the plurality of words of the spoken utterance, that the spoken utterance is one of:

directed toward the automated assistant server; or

not directed toward the automated assistant server; and

when the spoken utterance is directed toward the automated assistant server:

generating, by the data processing hardware, instructions that cause the automated assistant server to generate a response to the spoken utterance; and

providing, by the data processing hardware, for output from the automated assistant device, an indication that an audience for the spoken utterance is directed toward the automated assistant server.

2. The method of claim 1 , wherein the respective textual representation comprises a fixed-length vector.

3. The method of claim 2 , wherein the fixed-length vector comprises a 100-unit vector.

4. The method of claim 1 , wherein the automated assistant server generates the response to the spoken utterance by processing a transcription of the spoken utterance.

5. The method of claim 1 , wherein the spoken utterance is captured by a microphone of the automated assistant device.

6. The method of claim 1 , wherein the spoken utterance comprises an audio waveform.

7. The method of claim 1 , wherein the indication comprises an audible tone.

8. The method of claim 1 , wherein the indication comprises a flashing light.

9. The method of claim 1 , further comprising, when the spoken utterance is not directed toward the automated assistant server, discarding, by the data processing hardware, the spoken utterance captured without generating the instructions that cause the automated assistant server generate the response to the spoken utterance.

10. A system comprising:

data processing hardware; and

memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a spoken utterance captured by an automated assistant device associated with a user, the spoken utterance comprising a plurality of words;

generating, using a neural network-based utterance classifier comprising a stack of multiple Long-Short Term Memory (LSTM) layers, a respective textual representation for each word of the of the plurality of words of the spoken utterance, the neural network-based utterance classifier trained on negative training examples of spoken utterances not directed toward an automated assistant server;

determining, using the respective textual representation generated for each word of the plurality of words of the spoken utterance, that the spoken utterance is one of:

directed toward the automated assistant server; or

not directed toward the automated assistant server; and

when the spoken utterance is directed toward the automated assistant server:

generating instructions that cause the automated assistant server to generate a response to the spoken utterance; and

providing, for output from the automated assistant device, an indication that an audience for the spoken utterance is directed toward the automated assistant server.

11. The system of claim 10 , wherein the respective textual representation comprises a fixed-length vector.

12. The system of claim 11 , wherein the fixed-length vector comprises a 100-unit vector.

13. The system of claim 10 , wherein the automated assistant server generates the response to the spoken utterance by processing a transcription of the spoken utterance.

14. The system of claim 10 , wherein the spoken utterance is captured by a microphone of the automated assistant device.

15. The system of claim 10 , wherein the spoken utterance comprises an audio waveform.

16. The system of claim 10 , wherein the indication comprises an audible tone.

17. The system of claim 10 , wherein the indication comprises a flashing light.

18. The system of claim 10 , wherein the operations further comprise, when the spoken utterance is not directed toward the automated assistant server, discarding the spoken utterance captured without generating the instructions that cause the automated assistant server generate the response to the spoken utterance.

Continuity (3)
Continuation 16401349 · May 2, 2019
Continuation 15659016 · Jul 25, 2017
Related Publication 20200349946A1 · Nov 5, 2020