IP Library Granted Patent US 9,922,645
Granted Patent B2
US 9,922,645 · App. 15/460,342 · Granted Mar 20, 2018

Recognizing speech in the presence of additional audio

Inventors: Diego Melendo Casado (San Francisco, CA); Ignacio Lopez Moreno (New York, NY); Javier Gonzalez-Dominguez (Madrid, ES)
Assignee: Google LLC
G10L15/20G10L21/034G10L25/84G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,922,645
App. No.
15/460,342
Filed
Mar 16, 2017
Granted
Mar 20, 2018
Kind
B2
Art Unit
2673
USPC
704/233
Abstract

The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Claims (46)

1. A computer-implemented method comprising:

receiving, by a computing device that includes (i) a text-to-speech engine, (ii) an automated speech recognizer, and (iii) a barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, an audio signal corresponding to a user's utterance that is spoken while the computing device is outputting synthesized speech;

processing, using the barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, particular audio data that comprises data corresponding to the audio signal and data corresponding to the synthesized speech;

in response to receiving an indication from the barge-in model that the particular audio data comprises synthesized speech, suppressing a further output of the text-to-speech engine; and

outputting, by the automated speech recognizer, a transcription of the user's utterance without the synthesized speech.

2. The method of claim 1 , wherein processing the particular audio data comprises:

determining an obtained vector corresponding to the audio signal received by the computing device;

comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to the barge-in model; and

determining that the audio signal likely includes synthesized speech based on a result of comparing the obtained vector to the reference vector satisfying a predetermined threshold.

3. The method of claim 2 , wherein comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to a model that is trained to identify whether given audio data comprises additional audio other than a synthesized voice comprises calculating a cosine distance between the obtained vector and the reference vector, and

wherein the predetermined threshold corresponds to a value representing a threshold cosine distance between the obtained vector and the reference vector.

4. The method of claim 1 , wherein suppressing the further output of the text-to-speech engine comprises initiating a reduction in an audio output level of the text-to-speech engine.

5. The method of claim 1 , wherein suppressing the further output of the text-to-speech engine comprises at least temporarily ceasing an audio output of the text-to-speech engine.

6. The method of claim 1 , wherein the barge-in model is trained to identify whether given audio data comprises a user's voice.

7. The method of claim 1 , wherein the data corresponding to the audio signal comprises an i-vector.

8. A system comprising:

one or more computers; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a computing device that includes (i) a text-to-speech engine, (ii) an automated speech recognizer, and (iii) a barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, an audio signal corresponding to a user's utterance that is spoken while the computing device is outputting synthesized speech;

processing, using the barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, particular audio data that comprises data corresponding to the audio signal and data corresponding to the synthesized speech;

in response to receiving an indication from the barge-in model that the particular audio data comprises synthesized speech, suppressing a further output of the text-to-speech engine; and

outputting, by the automated speech recognizer, a transcription of the user's utterance without the synthesized speech.

9. The system of claim 8 , wherein processing the particular audio data comprises:

determining an obtained vector corresponding to the audio signal received by the computing device;

comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to the barge-in model; and

determining that the audio signal likely includes synthesized speech based on a result of comparing the obtained vector to the reference vector satisfying a predetermined threshold.

10. The system of claim 9 , wherein comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to a model that is trained to identify whether given audio data comprises additional audio other than a synthesized voice comprises calculating a cosine distance between the obtained vector and the reference vector, and

wherein the predetermined threshold corresponds to a value representing a threshold cosine distance between the obtained vector and the reference vector.

11. The system of claim 8 , wherein suppressing the further output of the text-to-speech engine comprises initiating a reduction in an audio output level of the text-to-speech engine.

12. The system of claim 8 , wherein suppressing the further output of the text-to-speech engine comprises at least temporarily ceasing an audio output of the text-to-speech engine.

13. The system of claim 8 , wherein the barge-in model is trained to identify whether given audio data comprises a user's voice.

14. The system of claim 8 , wherein the data corresponding to the audio signal comprises an i-vector.

15. A non-transitory computer readable storage device storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving, by a computing device that includes (i) a text-to-speech engine, (ii) an automated speech recognizer, and (iii) a barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, an audio signal corresponding to a user's utterance that is spoken while the computing device is outputting synthesized speech;

processing, using the barge-in detection model that is trained to output an indication of whether given audio data comprises synthesized speech, particular audio data that comprises data corresponding to the audio signal and data corresponding to the synthesized speech;

in response to receiving an indication from the barge-in model that the particular audio data comprises synthesized speech, suppressing a further output of the text-to-speech engine; and

outputting, by the automated speech recognizer, a transcription of the user's utterance without the synthesized speech.

16. The device of claim 15 , wherein processing the particular audio data comprises:

determining an obtained vector corresponding to the audio signal received by the computing device;

comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to the barge-in model; and

determining that the audio signal likely includes synthesized speech based on a result of comparing the obtained vector to the reference vector satisfying a predetermined threshold.

17. The device of claim 16 , wherein comparing the obtained vector corresponding to the audio signal to a reference vector corresponding to a model that is trained to identify whether given audio data comprises additional audio other than a synthesized voice comprises calculating a cosine distance between the obtained vector and the reference vector, and

wherein the predetermined threshold corresponds to a value representing a threshold cosine distance between the obtained vector and the reference vector.

18. The device of claim 15 , wherein suppressing the further output of the text-to-speech engine comprises initiating a reduction in an audio output level of the text-to-speech engine.

19. The device of claim 15 , wherein suppressing the further output of the text-to-speech engine comprises at least temporarily ceasing an audio output of the text-to-speech engine.

20. The device of claim 15 , wherein the barge-in model is trained to identify whether given audio data comprises a user's voice.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2017
From: MELENDO CASADO, DIEGO; LOPEZ MORENO, IGNACIO; GONZALEZ-DOMINGUEZ, JAVIER
To: GOOGLE INC.
Reel/Frame 041600/0396 →
Continuity (3)
Continuation 15093309 · Apr 7, 2016
Continuation 14181345 · Feb 14, 2014
Related Publication 20170186424A1 · Jun 29, 2017