IP Library › Granted Patent US 11,557,280
Granted Patent B2
US 11,557,280 · App. 17/101,946 · Granted Jan 17, 2023

Background audio identification for speech disambiguation

Inventors: Jason Sanders (New York, NY); Gabriel Taubman (Brooklyn, NY); John J. Lee (Long Island City, NY)
Assignee: Google LLC
G10L15/08G06F16/685G10L15/1815G10L15/22G10L15/26G10L21/0272G10L25/48H04M3/4936G10L21/0208G10L2015/225H04M2201/40H04M2203/352
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,557,280
App. No.
17/101,946
Granted
Jan 17, 2023
Kind
B2
Abstract

Implementations relate to techniques for providing context-dependent search results. A computer-implemented method includes receiving an audio stream at a computing device during a time interval, the audio stream comprising user speech data and background audio, separating the audio stream into a first substream that includes the user speech data and a second substream that includes the background audio, identifying concepts related to the background audio, generating a set of terms related to the identified concepts, influencing a speech recognizer based on at least one of the terms related to the background audio, and obtaining a recognized version of the user speech data using the speech recognizer.

Claims (42)

1. A method comprising:

receiving, at data processing hardware, from a computing device associated with a user, first audio data and second audio data captured by the computing device;

processing, by the data processing hardware, the first audio data to identify a concept associated with the first audio data;

influencing, by the data processing hardware, a speech recognition language model based on the identified concept associated with the first audio data; and

generating, by the data processing hardware, using the influenced speech recognition language model, a textual representation of the second audio data.

2. The method of claim 1 , wherein the computing device captures the first audio data before capturing the second audio data.

3. The method of claim 1 , further comprising:

generating, by the data processing hardware, a set of terms related to the identified concept,

wherein influencing the speech recognition model based on the identified concept comprises adjusting a probability or relevance score associated with the speech recognition language model recognizing at least one term in the set of terms related to the identified concept.

4. The method of claim 3 , wherein generating the set of terms related to the identified concept comprises querying a conceptual expansion database for the set of terms using the identified concept associated with the first audio data.

5. The method of claim 3 , further comprising:

generating, by the data processing hardware, conceptual bias data using set of terms related to the identified concept associated with the first audio data; and

adjusting, by the data processing hardware, the probability or relevance score associated with the speech recognition language model recognizing the at least one term in the set of terms related to the identified concept based on the conceptual bias data.

6. The method of claim 5 , wherein generating the textual representation of the second audio data using the influenced speech recognition language model comprises, selecting, by the influenced speech recognition language model, the textual representation from a set of textual representations that have substantially similar frequencies of occurrence in a particular language by using the conceptual bias data to weigh a statistical selection of the textual representation from the set of textual representations.

7. The method of claim 1 , wherein the second audio data corresponds to an utterance spoken by the user of the computing device.

8. The method of claim 1 , wherein:

the data processing hardware resides on a server in communication with the computing device; and

the computing device is configured to transmit the first audio data and the second audio data over a communication channel to the data processing hardware residing on the server.

9. The method of claim 1 , further comprising transmitting, by the data processing hardware, the textual representation of the second audio data to the computing device, the textual representation when received by the computing device causing the computing device to perform a particular task based on the textual representation.

10. The method of claim 1 , further comprising transmitting, by the data processing hardware, the textual representation of the second audio data to the computing device, the textual representation when received by the computing device causing the computing device to display the textual representation on a display.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving, from a computing device associated with a user, first audio data and second audio data captured by the computing device;

processing the first audio data to identify a concept associated with the first audio data;

influencing a speech recognition language model based on the identified concept associated with the first audio data; and

generating, using the influenced speech recognition language model, a textual representation of the second audio data.

12. The system of claim 11 , wherein the computing device captures the first audio data before capturing the second audio data.

13. The system of claim 11 , wherein the operations further comprise:

generating a set of terms related to the identified concept,

wherein influencing the speech recognition model based on the identified concept comprises adjusting a probability or relevance score associated with the speech recognition language model recognizing at least one term in the set of terms related to the identified concept.

14. The system of claim 13 , wherein generating the set of terms related to the identified concept comprises querying a conceptual expansion database for the set of terms using the identified concept associated with the first audio data.

15. The system of claim 13 , wherein the operations further comprise:

generating conceptual bias data using set of terms related to the identified concept associated with the first audio data; and

adjusting the probability or relevance score associated with the speech recognition language model recognizing the at least one term in the set of terms related to the identified concept based on the conceptual bias data.

16. The system of claim 15 , wherein generating the textual representation of the second audio data using the influenced speech recognition language model comprises, selecting, by the influenced speech recognition language model, the textual representation from a set of textual representations that have substantially similar frequencies of occurrence in a particular language by using the conceptual bias data to weigh a statistical selection of the textual representation from the set of textual representations.

17. The system of claim 11 , wherein the second audio data corresponds to an utterance spoken by the user of the computing device.

18. The system of claim 11 , wherein:

the data processing hardware resides on a server in communication with the computing device; and

the computing device is configured to transmit the first audio data and the second audio data over a communication channel to the data processing hardware residing on the server.

19. The system of claim 11 , wherein the operations further comprise transmitting the textual representation of the second audio data to the computing device, the textual representation when received by the computing device causing the computing device to perform a particular task based on the textual representation.

20. The system of claim 11 , wherein the operations further comprise transmitting the textual representation of the second audio data to the computing device, the textual representation when received by the computing device causing the computing device to display the textual representation on a display.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2020
From: SANDERS, JASON; TAUBMAN, GABRIEL; LEE, JOHN J.
To: GOOGLE INC.
Reel/Frame 054661/0479 →
CHANGE OF NAME Recorded Dec 16, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 054771/0852 →
Continuity (10)
Continuation 16249211 · Jan 16, 2019
Continuation 15622341 · Jun 14, 2017
Continuation 14825648 · Aug 13, 2015
Continuation 14825648 · Aug 13, 2015
Continuation 13804986 · Mar 14, 2013
Provisional Application 61778570 · Mar 13, 2013
Provisional Application 61654387 · Jun 1, 2012
Provisional Application 61654518 · Jun 1, 2012
Provisional Application 61654407 · Jun 1, 2012
Related Publication 20210082404A1 · Mar 18, 2021
Cited By (1)
US 12,711,952