IP Library Granted Patent US 12,254,876
Granted Patent B2
US 12,254,876 · App. 18/609,542 · Granted Mar 18, 2025

Recognizing speech in the presence of additional audio

Inventors: Diego Melendo Casado (Mountain View, CA); Ignacio Lopez Moreno (Brooklyn, NY); Javier Gonzalez-Dominguez (Madrid, ES)
Assignee: Google LLC
G10L15/20G06F3/165G06F3/167G10L15/222G10L17/06G10L21/034G10L25/84H03G3/3005G10L15/26G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,876
App. No.
18/609,542
Granted
Mar 18, 2025
Kind
B2
Abstract

The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Claims (40)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a first query spoken by a user and captured by a microphone of a computing device associated with the user;

providing, for audible playback from the computing device, a text-to-speech (TTS) output generated by a TTS system associated with the computing device, the TTS output comprising synthesized audio that conveys a response to the first query;

while the computing device is audibly playing back the TTS output:

detecting a barge-in event from the user to provide a second query;

in response to detecting the barge-in event, initiating a reduction in an audio output level of the computing device; and

receiving an audio signal captured by the microphone that conveys the second query spoken by the user; and

providing the audio signal characterizing the second query to a speech recognition engine.

2. The computer-implemented method of claim 1 , wherein a language model is configured to process an output of the speech recognition engine.

3. The computer-implemented method of claim 2 , wherein the output of the speech recognition engine comprises a vector of ordered pairs of speech units in a given language and corresponding weights.

4. The computer-implemented method of claim 3 , wherein the language model is configured to process the output of the speech recognition engine by determining a likelihood of word and phrase sequences from the speech units in the given language and the corresponding weights.

5. The computer-implemented method of claim 1 , wherein initiating the reduction in the audio output level of the computing device comprises providing a control signal to the computing device that causes the computing device to switch-off audible playback of the TTS output.

6. The computer-implemented method of claim 1 , wherein the data processing hardware is implemented on the computing device.

7. The computer-implemented method of claim 1 , wherein the speech recognition engine is configured to generate a transcription of the second query spoken by the user.

8. The computer-implemented method of claim 7 , wherein the operations further comprise:

transforming the transcription of the second query into a structured representation; and

processing, using a particular application, the structured representation.

9. The computer-implemented method of claim 1 , wherein the computing device comprises a mobile phone.

10. The computer-implemented method of claim 1 , wherein the computing device comprises a speaker device.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

receiving a first query spoken by a user and captured by a microphone of a computing device associated with the user;

providing, for audible playback from the computing device, a text-to-speech (TTS) output generated by a TTS system associated with the computing device, the TTS output comprising synthesized audio that conveys a response to the first query;

while the computing device is audibly playing back the TTS output:

detecting a barge-in event from the user to provide a second query;

in response to detecting the barge-in event, initiating a reduction in an audio output level of the computing device; and

receiving an audio signal captured by the microphone that conveys the second query spoken by the user; and

providing the audio signal characterizing the second query to a speech recognition engine.

12. The system of claim 11 , wherein a language model is configured to process an output of the speech recognition engine.

13. The system of claim 12 , wherein the output of the speech recognition engine comprises a vector of ordered pairs of speech units in a given language and corresponding weights.

14. The system of claim 13 , wherein the language model is configured to process the output of the speech recognition engine by determining a likelihood of word and phrase sequences from the speech units in the given language and the corresponding weights.

15. The system of claim 11 , wherein initiating the reduction in the audio output level of the computing device comprises providing a control signal to the computing device that causes the computing device to switch-off audible playback of the TTS output.

16. The system of claim 11 , wherein the data processing hardware is implemented on the computing device.

17. The system of claim 11 , wherein the speech recognition engine is configured to generate a transcription of the second query spoken by the user.

18. The system of claim 17 , wherein the operations further comprise:

transforming the transcription of the second query into a structured representation; and

processing, using a particular application, the structured representation.

19. The system of claim 11 , wherein the computing device comprises a mobile phone.

20. The system of claim 11 , wherein the computing device comprises a speaker device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2024
From: MELENDO CASADO, DIEGO; LOPEZ MORENO, IGNACIO; GONZALEZ-DOMINGUEZ, JAVIER
To: GOOGLE, INC.
Reel/Frame 066864/0265 →
CHANGE OF NAME Recorded Mar 22, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 066870/0400 →
Continuity (7)
Continuation 17303139 · May 21, 2021
Continuation 16548947 · Aug 23, 2019
Continuation 15887034 · Feb 2, 2018
Continuation 15460342 · Mar 16, 2017
Continuation 15093309 · Apr 7, 2016
Continuation 14181345 · Feb 14, 2014
Related Publication 20240221737A1 · Jul 4, 2024
References Cited (30)
US 6314402B1 · Monaco et al. · 2001 [cited by applicant]
US 7162421B1 · Zeppenfeld et al. · 2007 [cited by applicant]
US 7315613B2 · Kleindienst et al. · 2008 [cited by applicant]
US 7392188B2 · Junkawitsch et al. · 2008 [cited by applicant]
US 8073699B2 · Michelini et al. · 2011 [cited by applicant]
US 8112280B2 · Lu · 2012 [cited by applicant]
US 8527270B2 · Precoda et al. · 2013 [cited by applicant]
US 8566104B2 · Michelini et al. · 2013 [cited by applicant]
US 9240183B2 · Sharifi · 2016 [cited by examiner]
US 9318112B2 · Casado et al. · 2016 [cited by applicant]
US 9390725B2 · Graham · 2016 [cited by applicant]
US 9601116B2 · Melendo Casado et al. · 2017 [cited by applicant]
US 9922645B2 · Melendo Casado · 2018 [cited by examiner]
US 10869154B2 · Bruser et al. · 2020 [cited by applicant]
US 10943599B2 · Mitic et al. · 2021 [cited by applicant]
US 11605393B2 · Mitic et al. · 2023 [cited by applicant]
US 20020173333A1 · Buchholz et al. · 2002 [cited by applicant]
US 20080152095A1 · Kleindienst et al. · 2008 [cited by applicant]
US 20120029904A1 · Precoda et al. · 2012 [cited by applicant]
US 20120084084A1 · Zhu et al. · 2012 [cited by applicant]
US 20140128004A1 · Muralidhar et al. · 2014 [cited by applicant]
US 20150110263A1 · Johnston et al. · 2015 [cited by applicant]
US 20150112671A1 · Johnston et al. · 2015 [cited by applicant]
US 20150112684A1 · Scheffer et al. · 2015 [cited by applicant]
US 20160260440A1 · Joshi · 2016 [cited by applicant]
US 20200294487A1 · Donohoe et al. · 2020 [cited by applicant]
US 20210035554A1 · Iwase · 2021 [cited by examiner]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20230162752A1 · Mitic et al. · 2023 [cited by applicant]
WO WO2005015545A1 · 2005 [cited by examiner]