IP Library › Granted Patent US 12,002,452
Granted Patent B2
US 12,002,452 · App. 18/069,663 · Granted Jun 4, 2024

Background audio identification for speech disambiguation

Inventors: Jason Sanders (New York, NY); Gabriel Taubman (Brooklyn, NY); John J. Lee (Long Island City, NY)
Assignee: Google LLC
G10L15/08G06F16/685G10L15/1815G10L15/22G10L15/26G10L21/0272G10L25/48H04M3/4936G10L2015/225G10L21/0208H04M2201/40H04M2203/352
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,002,452
App. No.
18/069,663
Granted
Jun 4, 2024
Kind
B2
Abstract

Implementations relate to techniques for providing context-dependent search results. A computer-implemented method includes receiving an audio stream at a computing device during a time interval, the audio stream comprising user speech data and background audio, separating the audio stream into a first substream that includes the user speech data and a second substream that includes the background audio, identifying concepts related to the background audio, generating a set of terms related to the identified concepts, influencing a speech recognizer based on at least one of the terms related to the background audio, and obtaining a recognized version of the user speech data using the speech recognizer.

Claims (32)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a computing device associated with a user;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

influencing, using the retrieved set of terms related to the identified entity, a speech recognition language model; and

generating, using the influenced speech recognition language model, a transcription of the second audio data.

2. The method of claim 1 , wherein the computing device captures the first audio data before capturing the second audio data.

3. The method of claim 1 , wherein influencing the speech recognition model using the set of terms related to the identified entity comprises adjusting a probability or relevance score associated with the speech recognition language model recognizing at least one term in the set of terms related to the identified entity.

4. The method of claim 1 , wherein the second audio data corresponds to an utterance spoken by the user of the computing device.

5. The method of claim 1 , wherein processing the first audio data to identify the entity.

6. The method of claim 1 , wherein the data processing hardware resides on the computing device.

7. The method of claim 1 , wherein the speech recognition language model executes on the computing device.

8. The method of claim 1 , wherein the retrieved list of terms comprises a list of songs.

9. The method of claim 1 , wherein the retrieved list of terms comprises a list of music performers.

10. The method of claim 1 , wherein the computing device comprises a speaker.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving first audio data and second audio data captured by a computing device associated with a user;

processing the first audio data to identify an entity associated with the first audio data;

retrieving a set of terms related to the identified entity;

influencing, using the retrieved set of terms related to the identified entity, a speech recognition language model; and

generating, using the influenced speech recognition language model, a transcription of the second audio data.

12. The system of claim 11 , wherein the computing device captures the first audio data before capturing the second audio data.

13. The system of claim 11 , wherein influencing the speech recognition model using the set of terms related to the identified entity comprises adjusting a probability or relevance score associated with the speech recognition language model recognizing at least one term in the set of terms related to the identified entity.

14. The system of claim 11 , wherein the second audio data corresponds to an utterance spoken by the user of the computing device.

15. The system of claim 11 , wherein processing the first audio data to identify the entity.

16. The system of claim 11 , wherein the data processing hardware resides on the computing device.

17. The system of claim 11 , wherein the speech recognition language model executes on the computing device.

18. The system of claim 11 , wherein the retrieved list of terms comprises a list of songs.

19. The system of claim 11 , wherein the retrieved list of terms comprises a list of music performers.

20. The system of claim 11 , wherein the computing device comprises a speaker.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: SANDERS, JASON; TAUBMAN, GABRIEL; LEE, JOHN J.
To: GOOGLE INC.
Reel/Frame 062174/0204 →
CHANGE OF NAME Recorded Dec 21, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 062204/0054 →
Continuity (11)
Continuation 17101946 · Nov 23, 2020
Continuation 16249211 · Jan 16, 2019
Continuation 15622341 · Jun 14, 2017
Continuation 14825648 · Aug 13, 2015
Continuation 13804986 · Mar 14, 2013
Continuation 14825648 · Aug 13, 2015
Provisional Application 61778570 · Mar 13, 2013
Provisional Application 61654407 · Jun 1, 2012
Provisional Application 61654518 · Jun 1, 2012
Provisional Application 61654387 · Jun 1, 2012
Related Publication 20230125170A1 · Apr 27, 2023
Cited By (1)
US 12,711,952