IP Library Granted Patent US 11,942,083
Granted Patent B2
US 11,942,083 · App. 17/303,139 · Granted Mar 26, 2024

Recognizing speech in the presence of additional audio

Inventors: Diego Melendo Casado (Mountain View, CA); Ignacio Lopez Moreno (New York, NY); Javier Gonzalez-Dominguez (Madrid, ES)
Assignee: Google LLC
G10L15/20G06F3/165G06F3/167G10L15/222G10L17/06G10L21/034G10L25/84H03G3/3005G10L15/26G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,942,083
App. No.
17/303,139
Granted
Mar 26, 2024
Kind
B2
Abstract

The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Claims (36)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

while audio is being played back from a computing device, receiving a first audio signal captured by a microphone of the computing device, the first audio signal comprising the played back audio and speech audio corresponding to a query, the played back audio different than the speech audio corresponding to the query;

processing, using a neural network-based model, the first audio signal to determine that the speech audio corresponding to the query was spoken by a user of the computing device; and

in response to determining that the speech audio corresponding to the query was spoken by the user, generating a second audio signal that comprises the speech audio corresponding to the query and suppresses the played back audio from the first audio signal captured by the microphone.

2. The computer-implemented method of claim 1 , wherein the neural network-based model is trained to recognize a presence of a voice of the user of the computing device.

3. The computer-implemented method of claim 1 , wherein the neural network-based model is trained to recognize output audio from the computing device.

4. The computer-implemented method of claim 1 , wherein the neural network-based model is trained to:

to recognize a presence of a voice of the user of the computing device; and

to recognize output audio from the computing device.

5. The computer-implemented method of claim 1 , wherein the operations further comprise processing the second audio signal to generate a transcription of the query spoken by the user.

6. The computer-implemented method of claim 5 , wherein the operations further comprise:

transforming the transcription of the query into a structured representation; and

processing, using a particular application, the structured representation.

7. The computer-implemented method of claim 1 , wherein the data processing hardware is implemented on the computing device.

8. The computer-implemented method of claim 1 , wherein the computing device comprises a mobile phone.

9. The computer-implemented method of claim 1 , wherein the computing device comprises a speaker device.

10. The computer-implemented method of claim 1 , wherein the operations further comprise providing, for audible output from the computing device, a text-to-speech (TTS) output conveying a response to the query in a synthesized voice.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

while audio is being played back from a computing device, receiving a first audio signal captured by a microphone of the computing device, the first audio signal comprising the played back audio and speech audio corresponding to a query, the played back audio different than the speech audio corresponding to the query;

processing, using a neural network-based model, the first audio signal to determine that the speech audio corresponding to the query was spoken by a user of the computing device; and

in response to determining that the speech audio corresponding to the query was spoken by the user, generating a second audio signal that comprises the speech audio corresponding to the query and suppresses the played back audio from the first audio signal captured by the microphone.

12. The system of claim 11 , wherein the neural network-based model is trained to recognize a presence of a voice of the user of the computing device.

13. The system of claim 11 , wherein the neural network-based model is trained to recognize output audio from the computing device.

14. The system of claim 11 , wherein the neural network-based model is trained to:

to recognize a presence of a voice of the user of the computing device; and

to recognize output audio from the computing device.

15. The system of claim 11 , wherein the operations further comprise processing the second audio signal to generate a transcription of the query spoken by the user.

16. The system of claim 15 , wherein the operations further comprise:

transforming the transcription of the query into a structured representation; and

processing, using a particular application, the structured representation.

17. The system of claim 11 , wherein the data processing hardware is implemented on the computing device.

18. The system of claim 11 , wherein the computing device comprises a mobile phone.

19. The system of claim 11 , wherein the computing device comprises a speaker device.

20. The system of claim 11 , wherein the operations further comprise providing, for audible output from the computing device, a text-to-speech (TTS) output conveying a response to the query in a synthesized voice.

Assignments (2)
CHANGE OF NAME Recorded Mar 22, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 066870/0400 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2021
From: MELENDO CASADO, DIEGO; LOPEZ MORENO, IGNACIO; GONZALEZ-DOMINGUEZ, JAVIER
To: GOOGLE INC.
Reel/Frame 056309/0523 →
Continuity (6)
Continuation 16548947 · Aug 23, 2019
Continuation 15887034 · Feb 2, 2018
Continuation 15460342 · Mar 16, 2017
Continuation 15093309 · Apr 7, 2016
Continuation 14181345 · Feb 14, 2014
Related Publication 20210272562A1 · Sep 2, 2021
Cited By (1)
US 12,695,528