IP Library › Granted Patent US 12,632,496
Granted Patent B1
US 12,632,496 · App. 18/937,953 · Granted May 19, 2026

Background audio identification for query disambiguation

Inventors: Jason Sanders (New York, NY); John J. Lee (Long Island City, NY); Gabriel Taubman (Brooklyn, NY)
Assignee: GOOGLE LLC
G06F16/634
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,496
App. No.
18/937,953
Granted
May 19, 2026
Kind
B1
Abstract

Implementations relate to techniques for providing context-dependent search results. The techniques can include receiving a query and background audio. The techniques can also include identifying the background audio, establishing concepts related to the background audio and obtaining terms related to the concepts related to the background audio. The techniques can also include obtaining search results based on the query and on at least one of the terms. The techniques can also include providing the search results.

Claims (61)

1 . A computer-implemented method comprising:

receiving, at a computing device, a spoken query provided by a user, wherein the spoken query provided by the user satisfies a threshold volume level;

receiving, at the computing device, background audio produced by a source in an environment of the computing device;

determining that at least a portion of the background audio is collected in a fixed time interval before the user provided the spoken query and/or after the user provided the spoken query;

transmitting, from the computing device to a remote computing device, the spoken query provided by the user and at least the portion of the background audio that is collected in the fixed time interval before the user provided the spoken query and/or after the user provided the spoken query, wherein transmitting the spoken query from the computing device to the remote computing device causes the remote computing device to:

execute a search based on the spoken query to obtain one or more search results; and

filter, based on the portion of the background audio, the one or more search results, wherein the one or more search results are filtered based on one or more related terms of a set of related terms that describe entities being associated in an entity-relationship model with a known audio segment that is identified based on the portion of the background audio that is produced by the source in the environment of the computing device;

receiving, from the remote computing device, the one or more filtered search results; and

causing, in response to receiving the one or more filtered search results, one or more of the filtered search results to be rendered via the computing device for presentation to the user.

2 . The method of claim 1 , wherein the one or more filtered search results include only search results that include one or more of the related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

3 . The method of claim 1 , wherein the one or more search results are further filtered based on one or more of the search results satisfying a threshold score.

4 . The method of claim 1 , wherein the one or more filtered search results include only search results that include one or more related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

5 . The method of claim 1 , wherein receiving one or more of the filtered search results is in response to:

the spoken query being provided to a search engine to execute the search based on the spoken query to obtain the one or more search results;

scored results being received from the search engine; and

the search results being filtered based on a score for a search result containing at least one of the related terms being altered.

6 . The method of claim 1 , wherein the one or more search results are further filtered based on:

at least a portion of the background audio being recognized by matching at least the portion of the background audio to an acoustic fingerprint; and

the known audio segment being identified based on the background audio comprising the known audio segment associated with the acoustic fingerprint.

7 . The method of claim 1 , wherein the one or more search results are further filtered based on the set of related terms being generated based on a database being queried based on the known audio segment based on the background audio.

8 . A system comprising:

memory storing instructions; and

one or more processors operable to execute the instructions to:

receive, at a computing device, a spoken query provided by a user, wherein the spoken query provided by the user satisfies a threshold volume level;

receive, at the computing device, background audio produced by a source in an environment of the computing device;

determine that at least a portion of the background audio is collected in a fixed time interval before the user provided the spoken query and/or after the user provided the spoken query;

transmit, from the computing device to a remote computing device, the spoken query provided by the user and at least the portion of the background audio that is collected in the fixed time interval before the user provided the spoken query and/or after the user provided the spoken query, wherein transmitting the spoken query from the computing device to the remote computing device causes the remote computing device to:

execute a search based on the spoken query to obtain one or more search results; and

filter, based on the portion of the background audio, the one or more search results, wherein the one or more search results are filtered based on one or more related terms of a set of related terms that describe entities being associated in an entity-relationship model with a known audio segment that is identified based on the portion of the background audio that is produced by the source in the environment of the computing device;

receive, from the remote computing device, the one or more filtered search results; and

cause, in response to receiving the one or more filtered search results, one or more of the filtered search results to be rendered via the computing device for presentation to the user.

9 . The system of claim 8 , wherein the one or more filtered search results include only search results that include one or more of the related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

10 . The system of claim 8 , wherein the one or more search results are further filtered based on one or more of the search results satisfying a threshold score.

11 . The system of claim 8 , wherein the one or more filtered search results include only search results that include one or more related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

12 . The system of claim 8 , wherein receiving one or more of the filtered search results is in response to:

the spoken query being provided to a search engine to execute the search based on the spoken query to obtain the one or more search results;

scored results being received from the search engine; and

the search results being filtered based on a score for a search result containing at least one of the related terms being altered.

13 . The system of claim 8 , wherein the one or more search results are further filtered based on:

at least a portion of the background audio being recognized by matching at least the portion of the background audio to an acoustic fingerprint; and

the known audio segment being identified based on the background audio comprising the known audio segment associated with the acoustic fingerprint.

14 . The system of claim 8 , wherein the one or more search results are further filtered based on the set of related terms being generated based on a database being queried based on the known audio segment based on the background audio.

15 . A non-transitory computer readable storage medium configured to store instructions that, when executed by one or more processors, cause one or more of the processors to:

receive, at a computing device, a spoken query provided by a user, wherein the spoken query provided by the user satisfies a threshold volume level;

receive, at the computing device, background audio produced by a source in an environment of the computing device;

determine that at least a portion of the background audio is collected in a fixed time interval before the user provided the spoken query and/or after the user provided the spoken query;

transmit, from the computing device to a remote computing device, the spoken query provided by the user and at least the portion of the background audio that is collected in the fixed time interval before the user provided the spoken query and/or after the user provided the spoken query, wherein transmitting the spoken query from the computing device to the remote computing device causes the remote computing device to:

execute a search based on the spoken query to obtain one or more search results; and

filter, based on the portion of the background audio, the one or more search results, wherein the one or more search results are filtered based on one or more related terms of a set of related terms that describe entities being associated in an entity-relationship model with a known audio segment that is identified based on the portion of the background audio that is produced by the source in the environment of the computing device;

receive, from the remote computing device, the one or more filtered search results; and

cause, in response to receiving the one or more filtered search results, one or more of the filtered search results to be rendered via the computing device for presentation to the user.

16 . The non-transitory computer readable storage medium of claim 15 , wherein the one or more filtered search results include only search results that include one or more of the related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

17 . The non-transitory computer readable storage medium of claim 15 , wherein the one or more search results are further filtered based on one or more of the search results satisfying a threshold score.

18 . The non-transitory computer readable storage medium of claim 15 , wherein the one or more filtered search results include only search results that include one or more related terms of the set of related terms that describe entities that are associated in the entity-relationship model with the known audio segment that is identified based on the background audio.

19 . The non-transitory computer readable storage medium of claim 15 , wherein receiving one or more of the filtered search results is in response to:

the spoken query being provided to a search engine to execute the search based on the spoken query to obtain the one or more search results;

scored results being received from the search engine; and

the search results being filtered based on a score for a search result containing at least one of the related terms being altered.

20 . The non-transitory computer readable storage medium of claim 15 , wherein the one or more search results are further filtered based on:

at least a portion of the background audio being recognized by matching at least the portion of the background audio to an acoustic fingerprint; and

the known audio segment being identified based on the background audio comprising the known audio segment associated with the acoustic fingerprint.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2024
From: SANDERS, JASON; LEE, JOHN J.; TAUBMAN, GABRIEL
To: GOOGLE INC.
Reel/Frame 069578/0123 →
CHANGE OF NAME Recorded Dec 13, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 069689/0954 →
Continuity (7)
Continuation 18142006 · May 1, 2023
Continuation 17334378 · May 28, 2021
Continuation 16244366 · Jan 10, 2019
Continuation 13795153 · Mar 12, 2013
Provisional Application 61654518 · Jun 1, 2012
Provisional Application 61654407 · Jun 1, 2012
Provisional Application 61654387 · Jun 1, 2012
References Cited (62)
US 5721902A · Schultz · 1998 [cited by applicant]
US 6006622A · Bischof et al. · 1999 [cited by applicant]
US 6345252B1 · Beigi et al. · 2002 [cited by applicant]
US 6415258B1 · Reynar et al. · 2002 [cited by applicant]
US 6454520B1 · Pickelman et al. · 2002 [cited by applicant]
US 6584439B1 · Geilhufe · 2003 [cited by applicant]
US 6873604B1 · Surazski · 2005 [cited by examiner]
US 7027987B1 · Franz et al. · 2006 [cited by applicant]
US 7089188B2 · Logan et al. · 2006 [cited by applicant]
US 7783620B1 · Chevalier et al. · 2010 [cited by applicant]
US 8131716B2 · Chevalier et al. · 2012 [cited by applicant]
US 8150872B2 · Bernard · 2012 [cited by applicant]
US 8352467B1 · Guha · 2013 [cited by applicant]
US 9947333B1 · David · 2018 [cited by applicant]
US 11023520B1 · Sanders et al. · 2021 [cited by applicant]
US 11640426B1 · Sanders et al. · 2023 [cited by applicant]
US 20010034601A1 · Chujo · 2001 [cited by applicant]
US 20040225650A1 · Cooper et al. · 2004 [cited by applicant]
US 20050154713A1 · Glover et al. · 2005 [cited by applicant]
US 20060074883A1 · Teevan et al. · 2006 [cited by applicant]
US 20060200431A1 · Dwork et al. · 2006 [cited by applicant]
US 20060253427A1 · Wu et al. · 2006 [cited by applicant]
US 20070071206A1 · Gainsboro et al. · 2007 [cited by applicant]
US 20080005068A1 · Dumais et al. · 2008 [cited by applicant]
US 20080005076A1 · Payne et al. · 2008 [cited by applicant]
US 20080162471A1 · Bernard · 2008 [cited by applicant]
US 20080215597A1 · Nanba · 2008 [cited by applicant]
US 20080244675A1 · Sako et al. · 2008 [cited by applicant]
US 20080250011A1 · Haubold et al. · 2008 [cited by applicant]
US 20090006294A1 · Liu · 2009 [cited by applicant]
US 20090018898A1 · Genen · 2009 [cited by applicant]
US 20090157383A1 · Cho et al. · 2009 [cited by applicant]
US 20090228439A1 · Manolescu et al. · 2009 [cited by applicant]
US 20100076996A1 · Hu et al. · 2010 [cited by applicant]
US 20100145971A1 · Cheng et al. · 2010 [cited by applicant]
US 20100204986A1 · Kennewick et al. · 2010 [cited by applicant]
US 20110035382A1 · Bauer et al. · 2011 [cited by applicant]
US 20110093271A1 · Bernard · 2011 [cited by examiner]
US 20110153324A1 · Ballinger et al. · 2011 [cited by applicant]
US 20110257974A1 · Kristjansson et al. · 2011 [cited by applicant]
US 20110273213A1 · Rama · 2011 [cited by applicant]
US 20110295590A1 · Lloyd et al. · 2011 [cited by applicant]
US 20110307253A1 · Lloyd et al. · 2011 [cited by applicant]
US 20110313775A1 · Laligand · 2011 [cited by examiner]
US 20120034904A1 · Lebeau et al. · 2012 [cited by applicant]
US 20120117051A1 · Liu · 2012 [cited by applicant]
US 20120179465A1 · Cox et al. · 2012 [cited by applicant]
US 20120239175A1 · Mohajer et al. · 2012 [cited by applicant]
US 20120253802A1 · Heck et al. · 2012 [cited by applicant]
US 20120296458A1 · Koishida et al. · 2012 [cited by applicant]
US 20120296938A1 · Koishida et al. · 2012 [cited by applicant]
AU 2005200340 · 2005 [cited by applicant]
EP 2228737 · 2010 [cited by applicant]
WO 2007012120 · 2007 [cited by applicant]
WO 2007064640 · 2007 [cited by applicant]
WO 2012019020 · 2012 [cited by applicant]
Checkik et al., “Large-scale content-based audio retrieval from text queries”, Proceedings of the 1st ACM international conference on Multimedia information retrieval, pp. 106-112, 2008. [cited by applicant]
Egozi et al., “Concept-Based Feature Generation and Selection for Information Retrieval”, Proceedings of the 23rd AAAI Conference on Artificial Intelligence, pp. 1132-1137, 2008. [cited by applicant]
Haus et al., “An Audio front end for query-by-humming systems”, Proceedings of International Sympsium on Music Information Retrieval, 2001. [cited by applicant]
Mobile Image Search With Multimodal Context-Aware Queries, Yang et al, IEEE Computer Society Conference on Computer Vision and Pattern Recognition—Workshops, p. 25-32, Jun. 2010. [cited by applicant]
Multimodal Photo Annotation and Retrieval on a Mobile Phone, Anguera et al, 2008. [cited by applicant]
Natsev et al., Semantic concept-based query expansion and re-ranking for multimedia retrieval, Proceedings of the 15th International Conference on Multimedia, pp. 991-1000, 2007. [cited by applicant]