IP Library Granted Patent US 9,601,116
Granted Patent B2
US 9,601,116 · App. 15/093,309 · Granted Mar 21, 2017

Recognizing speech in the presence of additional audio

Inventors: Diego Melendo Casado (San Francisco, CA); Ignacio Lopez Moreno (New York, NY); Javier Gonzalez-Dominguez (Madrid, ES)
Assignee: Google Inc.
G10L15/222G06F3/165G06F3/167G10L17/06H03G3/3005G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,601,116
App. No.
15/093,309
Granted
Mar 21, 2017
Kind
B2
Abstract

The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Claims (52)

1. A computer-implemented method comprising:

receiving, by a mobile device, an audio signal;

determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice;

in response to determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice, suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device;

after suppressing operation of the speech synthesis module, obtaining, by the mobile device, a transcription corresponding to the audio signal from an automated speech recognizer; and

providing, by the mobile device, the transcription for output.

2. The method of claim 1 , wherein suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device comprises initiating a reduction in an audio output level of the speech synthesis module.

3. The method of claim 2 , wherein initiating a reduction in an audio output level of the speech synthesis module comprises interrupting output of the speech synthesis module.

4. The method of claim 1 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal;

comparing the first vector to a second vector corresponding to the model that is trained to detect a presence of a synthesized voice; and

determining that the audio signal comprises additional audio other than the synthesized voice based on a result of the comparison satisfying a threshold.

5. The method of claim 1 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal; and

determining that the audio signal comprises additional audio other than the synthesized voice based on the first vector satisfying a threshold.

6. The method of claim 1 , wherein each of the model that is trained to detect a presence of a synthesized voice and the model that is trained to detect a presence of a user's voice is an i-vector based model.

7. The method of claim 1 , wherein each of the model that is trained to detect a presence of a synthesized voice and the model that is trained to detect a presence of a user's voice is a neural network based model.

8. A non-transitory computer readable storage device storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving, by a mobile device, an audio signal;

determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice;

in response to determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice, suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device;

after suppressing operation of the speech synthesis module, obtaining, by the mobile device, a transcription corresponding to the audio signal from an automated speech recognizer; and

providing, by the mobile device, the transcription for output.

9. The computer readable storage device of claim 8 , wherein suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device comprises initiating a reduction in an audio output level of the speech synthesis module.

10. The computer readable storage device of claim 9 , wherein initiating a reduction in an audio output level of the speech synthesis module comprises interrupting output of the speech synthesis module.

11. The computer readable storage device of claim 8 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal;

comparing the first vector to a second vector corresponding to the model that is trained to detect a presence of a synthesized voice; and

determining that the audio signal comprises additional audio other than the synthesized voice based on a result of the comparison satisfying a threshold.

12. The computer readable storage device of claim 8 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal; and

determining that the audio signal comprises additional audio other than the synthesized voice based on the first vector satisfying a threshold.

13. The computer readable storage device of claim 8 , wherein each of the model that is trained to detect a presence of a synthesized voice and the model that is trained to detect a presence of a user's voice is an i-vector based model.

14. The computer readable storage device of claim 8 , wherein each of the model that is trained to detect a presence of a synthesized voice and the model that is trained to detect a presence of a user's voice is a neural network based model.

15. A system comprising:

one or more computers; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a mobile device, an audio signal;

determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice;

in response to determining, by the mobile device and using a model that is trained to detect a presence of a synthesized voice and a model that is trained to detect a presence of a user's voice, that the audio signal likely includes both the synthesized voice and the user's voice, suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device;

after suppressing operation of the speech synthesis module, obtaining, by the mobile device, a transcription corresponding to the audio signal from an automated speech recognizer; and

providing, by the mobile device, the transcription for output.

16. The system of claim 15 , wherein suppressing, by the mobile device, operation of a speech synthesis module implemented by the mobile device comprises initiating a reduction in an audio output level of the speech synthesis module.

17. The system of claim 16 , wherein initiating a reduction in an audio output level of the speech synthesis module comprises interrupting output of the speech synthesis module.

18. The system of claim 15 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal;

comparing the first vector to a second vector corresponding to the model that is trained to detect a presence of a synthesized voice; and

determining that the audio signal comprises additional audio other than the synthesized voice based on a result of the comparison satisfying a threshold.

19. The system of claim 15 , further comprising:

obtaining a first vector corresponding to at least a portion of the audio signal; and

determining that the audio signal comprises additional audio other than the synthesized voice based on the first vector satisfying a threshold.

20. The system of claim 15 , wherein each of the model that is trained to detect a presence of a synthesized voice and the model that is trained to detect a presence of a user's voice is one of: an i-vector based model and a neural network based model.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044097/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2016
From: CASADO, DIEGO MELENDO; MORENO, IGNACIO LOPEZ; GONZALEZ-DOMINGUEZ, JAVIER
To: GOOGLE INC.
Reel/Frame 038248/0329 →
Continuity (2)
Continuation 14181345 · Feb 14, 2014
Related Publication 20160225373A1 · Aug 4, 2016