IP Library Granted Patent US 12711952
Granted Patent B2
US 12711952 · App. 18/664,348 · Granted Aug 18, 2026

Background audio identification for speech disambiguation

Inventors: Jason Sanders (New York, NY); Gabriel Taubman (Brooklyn, NY); John J. Lee (Long Island City, NY)
Assignee: Google LLC
G10L15/08G06F16/685G10L15/1815G10L15/22G10L15/26G10L21/0272G10L25/48H04M3/4936G10L2015/225G10L21/0208H04M2201/40H04M2203/352
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711952
App. No.
18/664,348
Granted
Aug 18, 2026
Kind
B2
Abstract

Implementations relate to techniques for providing context-dependent search results. A computer-implemented method includes receiving an audio stream at a computing device during a time interval, the audio stream comprising user speech data and background audio, separating the audio stream into a first substream that includes the user speech data and a second substream that includes the background audio, identifying concepts related to the background audio, generating a set of terms related to the identified concepts, influencing a speech recognizer based on at least one of the terms related to the background audio, and obtaining a recognized version of the user speech data using the speech recognizer.

Claims (30)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a microphone of a computing device;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

processing, using a speech recognition model, the second audio data to determine one or more textual representations associated with the second audio data; and

selecting, based on the retrieved set of terms related to the identified entity, a particular textual representation from among the one or more textual representations as a transcription of the second audio data.

2 . The computer-implemented method of claim 1 , wherein the computing device captures the first audio data before capturing the second audio data.

3 . The computer-implemented method of claim 1 , wherein the second audio data corresponds to an utterance spoken by a user associated with the computing device.

4 . The computer-implemented method of claim 1 , wherein the data processing hardware resides on the computing device.

5 . The computer-implemented method of claim 1 , wherein the speech recognition language model executes on the computing device.

6 . The computer-implemented method of claim 1 , wherein the retrieved set of terms comprises a list of songs.

7 . The computer-implemented method of claim 1 , wherein the retrieved set of terms comprises a list of music performers.

8 . The computer-implemented method of claim 1 , wherein the computing device comprises a speaker.

9 . The computer-implemented method of claim 1 , wherein the particular textual representation comprises a lower relevance score than at least one other textual representation from the one or more textual representations.

10 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a microphone of a computing device;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

processing, using a speech recognition model, the second audio data to determine one or more textual representations associated with the second audio data; and

selecting, based on the retrieved set of terms related to the identified entity, a particular textual representation from among the one or more textual representations as a transcription of the second audio data.

11 . The system of claim 10 , wherein the computing device captures the first audio data before capturing the second audio data.

12 . The system of claim 10 , wherein the second audio data corresponds to an utterance spoken by a user associated with the computing device.

13 . The system of claim 10 , wherein the data processing hardware resides on the computing device.

14 . The system of claim 10 , wherein the speech recognition language model executes on the computing device.

15 . The system of claim 10 , wherein the retrieved set of terms comprises a list of songs.

16 . The system of claim 10 , wherein the retrieved set of terms comprises a list of music performers.

17 . The system of claim 10 , wherein the computing device comprises a speaker.

18 . The system of claim 10 , wherein the particular textual representation comprises a lower relevance score than at least one other textual representation from the one or more textual representations.