IP Library Granted Patent US 12,026,753
Granted Patent B2
US 12,026,753 · App. 17/308,624 · Granted Jul 2, 2024

Speech recognition for keywords

Inventors: Petar Aleksic (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G06Q30/0275G06Q30/0256G10L13/00G10L15/01G10L15/06G10L15/18G10L2015/088G10L15/187G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,026,753
App. No.
17/308,624
Granted
Jul 2, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition are disclosed. In one aspect, a method includes receiving a candidate adword from an advertiser. The method further includes generating a score for the candidate adword based on a likelihood of a speech recognizer generating, based on an utterance of the candidate adword, a transcription that includes a word that is associated with an expected pronunciation of the candidate adword. The method further includes classifying, based at least on the score, the candidate adword as an appropriate adword for use in a bidding process for advertisements that are selected based on a transcription of a speech query or as not an appropriate adword for use in the bidding process for advertisements that are selected based on the transcription of the speech query.

Claims (52)

1. A computer-implemented method executed by data processing hardware that causes the data processing hardware to perform operations comprising:

receiving, from an advertiser, a voice input of an adword;

transcribing the voice input of the adword into a plurality of potential phrases that differ from the adword using an automatic speech recognizer, each potential phrase of the plurality of potential phrases transcribed by the automatic speech recognizer indicating a misrecognition of the adword that is phonetically similar to the adword;

presenting, to the advertiser, a list of the plurality of potential phrases;

receiving, from the advertiser, a selection of one or more potential phrases from among the plurality of potential phrases to bid on for spoken queries, but not for typed queries;

receiving, from a user, a spoken query; and

distributing a content item based on the bid when the spoken query matches at least one of the selected one or more potential phrases.

2. The method of claim 1 , wherein:

the operations further comprise determining a frequency that each of the plurality of potential phrases occurs in a query log that includes transcriptions of previously spoken queries, and

presenting the list of the plurality of potential phrases comprises presenting the frequency that each of the plurality of potential phrases occurs in the query log.

3. The method of claim 1 , wherein:

the operations further comprise determining a most frequent location of users who spoke each of the plurality of potential phrases, and

presenting the list of the plurality of potential phrases comprises presenting the most frequent location of users who spoke each of the plurality of potential phrases.

4. The method of claim 1 , wherein presenting the list of the plurality of potential phrases comprises presenting data indicating whether an advertisement was presented in response to receiving each of the plurality of potential phrases as a spoken query.

5. The method of claim 1 , wherein transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer further comprises:

providing audio of the voice input as an input to an acoustic model that identifies candidate phonemes of the audio of the voice input;

providing data identifying the candidate phonemes that likely correspond to the audio of the voice input as an input to a language model that identifies candidate transcriptions of the candidate phonemes; and

selecting, from among the candidate transcriptions of the candidate phonemes, the potential phrases that differ from the adword.

6. The method of claim 5 , wherein selecting, from among the candidate transcriptions of the candidate phonemes, the potential phrases that differ from the adword comprises selecting the candidate transcriptions that are more likely to trigger displaying an advertisement.

7. The method of claim 1 , wherein:

the operations further comprise:

providing each of the potential phrases as an input to a speech to text module;

providing an output of the speech to text module as an input to the automatic speech recognizer; and

based on providing the output of the speech to text module as an input to the automatic speech recognizer, determining additional phrases that correspond to the voice input, and

wherein presenting the list of the plurality of potential phrases comprises presenting the list of the plurality of potential phrases and the additional phrases.

8. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving, from an advertiser, a voice input of an adword;

transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer, each potential phrase of the plurality of potential phrases transcribed by the automatic speech recognizer indicating a misrecognition of the adword that is phonetically similar to the adword;

presenting, to the advertiser, a list of the plurality of potential phrases;

receiving, from the advertiser, a selection of one or more potential phrases from among the plurality of potential phrases to bid on for spoken queries, but not for typed queries;

receiving, from a user, a spoken query; and

distributing a content item based on the bid when the spoken query matches at least one of the selected one or more potential phrases.

9. The system of claim 8 , wherein:

the operations further comprise determining a frequency that each of the plurality of potential phrases occurs in a query log that includes transcriptions of previously spoken queries, and

presenting the list of the plurality of potential phrases comprises presenting the frequency that each of the plurality of potential phrases occurs in the query log.

10. The system of claim 8 , wherein:

the operations further comprise determining a most frequent location of users who spoke each of the plurality of potential phrases, and

presenting the list of the plurality of potential phrases comprises presenting the most frequent location of users who spoke each of the plurality of potential phrases.

11. The system of claim 8 , wherein presenting the list of the plurality of potential phrases comprises presenting data indicating whether an advertisement was presented in response to receiving each of the plurality of potential phrases as a spoken query.

12. The system of claim 8 , wherein transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer further comprises:

providing audio of the voice input as an input to an acoustic model that identifies candidate phonemes of the audio of the voice input;

providing data identifying the candidate phonemes that likely correspond to the audio of the voice input as an input to a language model that identifies candidate transcriptions of the candidate phonemes; and

selecting, from among the candidate transcriptions of the candidate phonemes, the potential phrases that differ from the adword.

13. The system of claim 12 , wherein selecting, from among the candidate transcriptions of the candidate phonemes, the potential phrases that differ from the adword comprises selecting the candidate transcriptions that are more likely to trigger displaying an advertisement.

14. The system of claim 8 , wherein:

the operations further comprise:

providing each of the potential phrases as an input to a speech to text module;

providing an output of the speech to text module as an input to the automatic speech recognizer; and

based on providing the output of the speech to text module as an input to the automatic speech recognizer, determining additional phrases that correspond to the voice input, and

presenting the list of the plurality of potential phrases comprises presenting the list of the plurality of potential phrases and the additional phrases.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2021
From: ALEKSIC, PETAR; MORENO MENGIBAR, PEDRO J.
To: GOOGLE INC.
Reel/Frame 056146/0236 →
CHANGE OF NAME Recorded May 5, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 056151/0352 →
Continuity (3)
Continuation 16047723 · Jul 27, 2018
Continuation 14710928 · May 13, 2015
Related Publication 20210256567A1 · Aug 19, 2021