IP Library Granted Patent US 11,031,002
Granted Patent B2
US 11,031,002 · App. 16/548,947 · Granted Jun 8, 2021

Recognizing speech in the presence of additional audio

Inventors: Diego Melendo Casado (Mountain View, CA); Ignacio Lopez Moreno (New York, NY); Javier Gonzalez-Dominguez (Madrid, ES)
Assignee: Google LLC
G10L15/20G06F3/165G06F3/167G10L15/222G10L17/06G10L21/034G10L25/84H03G3/3005G10L15/26G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,031,002
App. No.
16/548,947
Granted
Jun 8, 2021
Kind
B2
Abstract

The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Claims (57)

1. A computer-implemented method comprising:

receiving, by a first computing device, first audio data that includes an utterance;

determining, by the first computing device, that a second computing device is outputting second audio data; and

based on determining that that the second computing device is outputting the second audio data and based on receiving the audio data that includes the utterance, providing, by the first computing device and for output to the second computing device, an instruction to suppress outputting the second audio data.

2. The method of claim 1 , comprising:

determining, by the first computing device, that the second audio data include speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the second audio data include speech.

3. The method of claim 1 , comprising:

determining, by the first computing device, that the first audio data includes speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the first audio data include speech.

4. The method of claim 3 , wherein determining that the first audio data includes speech comprises:

providing, as an input to a model that is configured to determine whether received audio data includes speech, the first audio data; and

receiving, from the model, data indicating that the first audio data includes speech.

5. The method of claim 4 , wherein the model configured to determine whether received audio data includes speech using audio fingerprinting or a neural network classifier.

6. The method of claim 1 , comprising:

after providing the instruction to suppress outputting the second audio data, obtaining, by the first computing device, a transcription of the utterance; and

providing, for output, the transcription.

7. The method of claim 1 , wherein providing the instruction to suppress outputting the second audio data comprises:

providing an instruction to mute the second computing device.

8. The method of claim 1 , wherein providing the instruction to suppress outputting the second audio data comprises:

providing an instruction to reduce a volume of the second audio data.

9. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a first computing device, first audio data that includes an utterance;

determining, by the first computing device, that a second computing device is outputting second audio data; and

based on determining that that the second computing device is outputting the second audio data and based on receiving the audio data that includes the utterance, providing, by the first computing device and for output to the second computing device, an instruction to suppress outputting the second audio data.

10. The system of claim 9 , wherein the operations comprise:

determining, by the first computing device, that the second audio data include speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the second audio data include speech.

11. The system of claim 9 , wherein the operations comprise:

determining, by the first computing device, that the first audio data includes speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the first audio data include speech.

12. The system of claim 11 , wherein determining that the first audio data includes speech comprises:

providing, as an input to a model that is configured to determine whether received audio data includes speech, the first audio data; and

receiving, from the model, data indicating that the first audio data includes speech.

13. The system of claim 12 , wherein the model configured to determine whether received audio data includes speech using audio fingerprinting or a neural network classifier.

14. The system of claim 9 , wherein the operations comprise:

after providing the instruction to suppress outputting the second audio data, obtaining, by the first computing device, a transcription of the utterance; and

providing, for output, the transcription.

15. The system of claim 9 , wherein providing the instruction to suppress outputting the second audio data comprises:

providing an instruction to mute the second computing device.

16. The system of claim 9 , wherein providing the instruction to suppress outputting the second audio data comprises:

providing an instruction to reduce a volume of the second audio data.

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a first computing device, first audio data that includes an utterance;

determining, by the first computing device, that a second computing device is outputting second audio data; and

based on determining that that the second computing device is outputting the second audio data and based on receiving the audio data that includes the utterance, providing, by the first computing device and for output to the second computing device, an instruction to suppress outputting the second audio data.

18. The medium of claim 17 , wherein the operations comprise:

determining, by the first computing device, that the second audio data include speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the second audio data include speech.

19. The medium of claim 17 , wherein the operations comprise:

determining, by the first computing device, that the first audio data includes speech,

wherein providing the instruction to suppress outputting the second audio data is based on determining that the first audio data include speech.

20. The medium of claim 17 , wherein the operations comprise:

after providing the instruction to suppress outputting the second audio data, obtaining, by the first computing device, a transcription of the utterance; and

providing, for output, the transcription.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2019
From: CASADO, DIEGO MELENDO; MORENO, IGNACIO LOPEZ; GONZALEZ-DOMINGUEZ, JAVIER
To: GOOGLE INC.
Reel/Frame 050143/0609 →
ENTITY CONVERSION Recorded Aug 23, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050154/0881 →
Continuity (5)
Continuation 15887034 · Feb 2, 2018
Continuation 15460342 · Mar 16, 2017
Continuation 15093309 · Apr 7, 2016
Continuation 14181345 · Feb 14, 2014
Related Publication 20200051553A1 · Feb 13, 2020
Cited By (1)
US 12,712,972