IP Library › Granted Patent US 12,711,952
Granted Patent B2
US 12,711,952 · App. 18/664,348 · Granted Aug 18, 2026

Background audio identification for speech disambiguation

Inventors: Jason Sanders (New York, NY); Gabriel Taubman (Brooklyn, NY); John J. Lee (Long Island City, NY)
Assignee: Google LLC
G10L15/08G06F16/685G10L15/1815G10L15/22G10L15/26G10L21/0272G10L25/48H04M3/4936G10L2015/225G10L21/0208H04M2201/40H04M2203/352
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,952
App. No.
18/664,348
Granted
Aug 18, 2026
Kind
B2
Abstract

Implementations relate to techniques for providing context-dependent search results. A computer-implemented method includes receiving an audio stream at a computing device during a time interval, the audio stream comprising user speech data and background audio, separating the audio stream into a first substream that includes the user speech data and a second substream that includes the background audio, identifying concepts related to the background audio, generating a set of terms related to the identified concepts, influencing a speech recognizer based on at least one of the terms related to the background audio, and obtaining a recognized version of the user speech data using the speech recognizer.

Claims (30)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a microphone of a computing device;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

processing, using a speech recognition model, the second audio data to determine one or more textual representations associated with the second audio data; and

selecting, based on the retrieved set of terms related to the identified entity, a particular textual representation from among the one or more textual representations as a transcription of the second audio data.

2 . The computer-implemented method of claim 1 , wherein the computing device captures the first audio data before capturing the second audio data.

3 . The computer-implemented method of claim 1 , wherein the second audio data corresponds to an utterance spoken by a user associated with the computing device.

4 . The computer-implemented method of claim 1 , wherein the data processing hardware resides on the computing device.

5 . The computer-implemented method of claim 1 , wherein the speech recognition language model executes on the computing device.

6 . The computer-implemented method of claim 1 , wherein the retrieved set of terms comprises a list of songs.

7 . The computer-implemented method of claim 1 , wherein the retrieved set of terms comprises a list of music performers.

8 . The computer-implemented method of claim 1 , wherein the computing device comprises a speaker.

9 . The computer-implemented method of claim 1 , wherein the particular textual representation comprises a lower relevance score than at least one other textual representation from the one or more textual representations.

10 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a microphone of a computing device;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

processing, using a speech recognition model, the second audio data to determine one or more textual representations associated with the second audio data; and

selecting, based on the retrieved set of terms related to the identified entity, a particular textual representation from among the one or more textual representations as a transcription of the second audio data.

11 . The system of claim 10 , wherein the computing device captures the first audio data before capturing the second audio data.

12 . The system of claim 10 , wherein the second audio data corresponds to an utterance spoken by a user associated with the computing device.

13 . The system of claim 10 , wherein the data processing hardware resides on the computing device.

14 . The system of claim 10 , wherein the speech recognition language model executes on the computing device.

15 . The system of claim 10 , wherein the retrieved set of terms comprises a list of songs.

16 . The system of claim 10 , wherein the retrieved set of terms comprises a list of music performers.

17 . The system of claim 10 , wherein the computing device comprises a speaker.

18 . The system of claim 10 , wherein the particular textual representation comprises a lower relevance score than at least one other textual representation from the one or more textual representations.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2024
From: SANDERS, JASON; TAUBMAN, GABRIEL; LEE, JOHN J.
To: GOOGLE, INC.
Reel/Frame 067414/0043 →
CHANGE OF NAME Recorded May 15, 2024
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 067414/0081 →
Continuity (11)
Continuation 18069663 · Dec 21, 2022
Continuation 17101946 · Nov 23, 2020
Continuation 16249211 · Jan 16, 2019
Continuation 15622341 · Jun 14, 2017
Continuation 14825648 · Aug 13, 2015
Continuation 13804986 · Mar 14, 2013
Provisional Application 61778570 · Mar 13, 2013
Provisional Application 61654387 · Jun 1, 2012
Provisional Application 61654518 · Jun 1, 2012
Provisional Application 61654407 · Jun 1, 2012
Related Publication 20240296835A1 · Sep 5, 2024
References Cited (87)
US 5685000A · Cox, Jr. · 1997 [cited by applicant]
US 6173266B1 · Marx et al. · 2001 [cited by applicant]
US 6829599B2 · Chidlovskii · 2004 [cited by applicant]
US 7024366B1 · Deyoe et al. · 2006 [cited by applicant]
US 7139717B1 · Abella et al. · 2006 [cited by applicant]
US 7197460B1 · Gupta et al. · 2007 [cited by applicant]
US 7249011B2 · Chou et al. · 2007 [cited by applicant]
US 7286987B2 · Roy · 2007 [cited by examiner]
US 7363282B2 · Karnawat et al. · 2008 [cited by applicant]
US 7418391B2 · Gayama et al. · 2008 [cited by applicant]
US 7451089B1 · Gupta et al. · 2008 [cited by applicant]
US 7587324B2 · Kaiser · 2009 [cited by applicant]
US 7702508B2 · Bennett · 2010 [cited by applicant]
US 7734471B2 · Paek et al. · 2010 [cited by applicant]
US 7869998B1 · Di Fabbrizio et al. · 2011 [cited by applicant]
US 7933775B2 · Quibria et al. · 2011 [cited by applicant]
US 8032481B2 · Pinckney et al. · 2011 [cited by applicant]
US 8090680B2 · Smeaton et al. · 2012 [cited by applicant]
US 8117197B1 · Cramer · 2012 [cited by applicant]
US 8126719B1 · Jochumson · 2012 [cited by applicant]
US 8160883B2 · Lecoeuche · 2012 [cited by applicant]
US 8185539B1 · Bhardwaj · 2012 [cited by applicant]
US 8214214B2 · Bennett · 2012 [cited by applicant]
US 8265939B2 · Kanevsky et al. · 2012 [cited by applicant]
US 8280888B1 · Bierner et al. · 2012 [cited by applicant]
US 8296144B2 · Weng et al. · 2012 [cited by applicant]
US 8473299B2 · Di Fabbrizio et al. · 2013 [cited by applicant]
US 8611876B2 · Miller · 2013 [cited by applicant]
US 8612223B2 · Minamino et al. · 2013 [cited by applicant]
US 8725512B2 · Claiborn et al. · 2014 [cited by applicant]
US 8825482B2 · Hernandez-Abrego et al. · 2014 [cited by applicant]
US 9123338B1 · Sanders · 2015 [cited by examiner]
US 9311915B2 · Weinstein · 2016 [cited by examiner]
US 9405363B2 · Hernandez-Abrego et al. · 2016 [cited by applicant]
US 9552816B2 · VanLund · 2017 [cited by examiner]
US 9812123B1 · Sanders · 2017 [cited by examiner]
US 10224024B1 · Sanders · 2019 [cited by examiner]
US 10872600B1 · Sanders · 2020 [cited by examiner]
US 11023520B1 · Sanders · 2021 [cited by examiner]
US 11557280B2 · Sanders · 2023 [cited by examiner]
US 11640426B1 · Sanders · 2023 [cited by examiner]
US 12002452B2 · Sanders · 2024 [cited by examiner]
US 12164562B1 · Sanders · 2024 [cited by examiner]
US 20010021909A1 · Shimomura et al. · 2001 [cited by applicant]
US 20020069058A1 · Jin et al. · 2002 [cited by applicant]
US 20020198707A1 · Zhou · 2002 [cited by applicant]
US 20040172252A1 · Aoki · 2004 [cited by examiner]
US 20040215449A1 · Roy · 2004 [cited by examiner]
US 20050027670A1 · Petropoulos · 2005 [cited by applicant]
US 20050091056A1 · Surace et al. · 2005 [cited by applicant]
US 20060122837A1 · Kim et al. · 2006 [cited by applicant]
US 20060149544A1 · Hakkani-Tur et al. · 2006 [cited by applicant]
US 20060190809A1 · Hejna · 2006 [cited by applicant]
US 20060248057A1 · Jacobs et al. · 2006 [cited by applicant]
US 20070003914A1 · Yang · 2007 [cited by applicant]
US 20070043571A1 · Michelini et al. · 2007 [cited by applicant]
US 20070061142A1 · Hernandez-Abrego et al. · 2007 [cited by applicant]
US 20070136246A1 · Stenchikova et al. · 2007 [cited by applicant]
US 20070192095A1 · Braho et al. · 2007 [cited by applicant]
US 20080133245A1 · Proulx et al. · 2008 [cited by applicant]
US 20080154828A1 · Antebi · 2008 [cited by applicant]
US 20080221901A1 · Cerra · 2008 [cited by examiner]
US 20080221902A1 · Cerra · 2008 [cited by examiner]
US 20090070113A1 · Gupta et al. · 2009 [cited by applicant]
US 20100104087A1 · Byrd et al. · 2010 [cited by applicant]
US 20100125456A1 · Weng et al. · 2010 [cited by applicant]
US 20110015928A1 · Odell et al. · 2011 [cited by applicant]
US 20110066634A1 · Phillips · 2011 [cited by examiner]
US 20110288855A1 · Roy · 2011 [cited by examiner]
US 20120041950A1 · Koll et al. · 2012 [cited by applicant]
US 20120059815A1 · Friedlander et al. · 2012 [cited by applicant]
US 20120063620A1 · Nomura et al. · 2012 [cited by applicant]
US 20120136667A1 · Emerick et al. · 2012 [cited by applicant]
US 20120265528A1 · Gruber · 2012 [cited by examiner]
US 20130063550A1 · Ritchey et al. · 2013 [cited by applicant]
US 20130086029A1 · Hebert · 2013 [cited by applicant]
US 20130304758A1 · Gruber et al. · 2013 [cited by applicant]
US 20140006019A1 · Paajanen · 2014 [cited by examiner]
US 20140229866A1 · Gottlieb · 2014 [cited by applicant]
US 20150039299A1 · Weinstein · 2015 [cited by examiner]
US 20150066479A1 · Pasupalak et al. · 2015 [cited by applicant]
US 20210082404A1 · Sanders · 2021 [cited by examiner]
US 20240296835A1 · Sanders · 2024 [cited by examiner]
Ingrid Lunden, “Another Siri-Like App, Voice Answer, Hits the App Store for Those of Us Without the iPhones 4S,” TechCmnch, Retrieved from <http://techcmnch.com/2012/04/18/another-siri-like-app-voice-answer-hits-the-app… [cited by applicant]
Qiaoling et al, “Predicting Web Searcher Satisfaction with Existing Community-based Answers,” Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. ACM, 2011. [cited by applicant]
Youtube video uploader—sparklingapps, “Voice Answer: a Siri like application for All iPhones and iPads”, Uploaded Jan. 20, 2012, Yotube Published Video link <https://youtu.be/zNufnccFIRc>. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/101,946, dated Aug. 18, 2022. [cited by applicant]