IP Library Granted Patent US 11,030,658
Granted Patent B2
US 11,030,658 · App. 16/047,723 · Granted Jun 8, 2021

Speech recognition for keywords

Inventors: Petar Aleksic (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G06Q30/0275G06Q30/0256G10L13/00G10L15/01G10L15/06G10L15/18G10L15/187G10L15/26G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,030,658
App. No.
16/047,723
Granted
Jun 8, 2021
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition are disclosed. In one aspect, a method includes receiving a candidate adword from an advertiser. The method further includes generating a score for the candidate adword based on a likelihood of a speech recognizer generating, based on an utterance of the candidate adword, a transcription that includes a word that is associated with an expected pronunciation of the candidate adword. The method further includes classifying, based at least on the score, the candidate adword as an appropriate adword for use in a bidding process for advertisements that are selected based on a transcription of a speech query or as not an appropriate adword for use in the bidding process for advertisements that are selected based on the transcription of the speech query.

Claims (68)

1. A computer-implemented method comprising:

receiving a voice input of an adword from an advertiser;

transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer;

for each potential phrase in the plurality of potential phrases that differ from the adword:

providing the potential phrase as an input to a text-to-speech module, the text-to-speech module converting the potential phrase to corresponding audio data as output;

providing the corresponding audio data output from the text-to-speech module as a corresponding input to the automatic speech recognizer; and

determining, based on providing the corresponding audio data output from the text-to-speech module as the corresponding input to the automatic speech recognizer, a corresponding additional phrase that corresponds to the voice input;

presenting, to the advertiser, a list of the plurality of potential phrases and the corresponding additional phrase determined for each potential phrase in the plurality of potential phrases;

receiving, from the advertiser, a selection of one or more potential phrases from among the plurality of potential phrases to bid on for spoken queries, but not for typed queries; and

distributing a content item based on the bid when a spoken query submitted by a user matches at least one of the selected one or more potential phrases.

2. The method of claim 1 , further comprising:

determining a frequency that each potential phrase in the plurality of potential phrases occurs in a query log that includes transcriptions of previously spoken queries,

wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting the frequency that each potential phrase in the plurality of potential phrases occurs in the query log.

3. The method of claim 1 , further comprising:

determining a most frequent location of users who spoke each potential phrase in the plurality of potential phrases,

wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting the most frequent location of users who spoke each potential phrase in the plurality of potential phrases.

4. The method of claim 1 , wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting data indicating whether an advertisement was presented in response to receiving each potential phrase in the plurality of potential phrases as a spoken query.

5. The method of claim 1 , wherein transcribing the voice input into the plurality of potential phrases that differ from the adword using the automatic speech recognizer comprises:

providing audio of the voice input as an input to an acoustic model that identifies candidate phonemes of the audio of the voice input;

providing data identifying the phonemes that likely correspond to the audio of the voice input as an input to a language model that identifies candidate transcriptions of the candidate phonemes; and

selecting, from among the candidate transcriptions of the candidate phonemes, the plurality of potential phrases that differ from the adword.

6. The method of claim 5 , wherein selecting, from among the candidate transcriptions of the candidate phonemes, the plurality of potential phrases that differ from the adword comprises selecting the candidate transcriptions that are more likely to trigger displaying an advertisement.

7. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving a voice input of an adword from an advertiser;

transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer;

for each potential phrase in the plurality of potential phrases that differ from the adword:

providing the potential phrase as an input to a text-to-speech module, the text-to-speech module converting the potential phrase to corresponding audio data as output;

providing the corresponding audio data output from the text-to-speech module as a corresponding input to the automatic speech recognizer; and

determining, based on providing the corresponding audio data output from the text-to-speech module as the corresponding input to the automatic speech recognizer, a corresponding additional phrase that corresponds to the voice input;

presenting, to the advertiser, a list of the plurality of potential phrases and the corresponding additional phrase determined for each potential phrase in the plurality of potential phrases;

receiving, from the advertiser, a selection of one or more potential phrases from among the plurality of potential phrases to bid on for spoken queries, but not for typed queries; and

distributing a content item based on the bid when a spoken query submitted by a user matches at least one of the selected one or more potential phrases.

8. The system of claim 7 , wherein the operations further comprise:

determining a frequency that each potential phrase in the plurality of potential phrases occurs in a query log that includes transcriptions of previously spoken queries,

wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting the frequency that each potential phrase in the plurality of potential phrases occurs in the query log.

9. The system of claim 7 , wherein the operations further comprise:

determining a most frequent location of users who spoke each potential phrase in the plurality of potential phrases,

wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting the most frequent location of users who spoke each potential phrase in the plurality of potential phrases.

10. The system of claim 7 , wherein presenting the list of the plurality of potential phrases and the additional phrases comprises presenting data indicating whether an advertisement was presented in response to receiving each potential phrase in the plurality of potential phrases as a spoken query.

11. The system of claim 7 , wherein transcribing the voice input into the plurality of potential phrases that differ from the adword using the automatic speech recognizer comprises:

providing audio of the voice input as an input to an acoustic model that identifies candidate phonemes of the audio of the voice input;

providing data identifying the phonemes that likely correspond to the audio of the voice input as an input to a language model that identifies candidate transcriptions of the candidate phonemes; and

selecting, from among the candidate transcriptions of the candidate phonemes, the plurality of potential phrases that differ from the adword.

12. The system of claim 11 , wherein selecting, from among the candidate transcriptions of the candidate phonemes, the plurality of potential phrases that differ from the adword comprises selecting the candidate transcriptions that are more likely to trigger displaying an advertisement.

13. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving a voice input of an adword from an advertiser;

transcribing the voice input into a plurality of potential phrases that differ from the adword using an automatic speech recognizer;

for each potential phrase in the plurality of potential phrases that differ from the adword:

providing the potential phrase as an input to a text-to-speech module, the text-to-speech module converting the potential phrase to corresponding audio data as output;

providing the corresponding audio data output from the text-to-speech module as a corresponding input to the automatic speech recognizer; and

determining, based on providing the corresponding audio data output from the text-to-speech module as the corresponding input to the automatic speech recognizer, a corresponding additional phrase that corresponds to the voice input;

presenting, to the advertiser, a list of the plurality of potential phrases and the corresponding additional phrase determined for each potential phrase in the plurality of potential phrases;

receiving, from the advertiser, a selection of one or more potential phrases from among the plurality of potential phrases to bid on for spoken queries, but not for typed queries; and

distributing a content item based on the bid when a spoken query submitted by a user matches at least one of the selected one or more potential phrases.

14. The medium of claim 13 , wherein the operations further comprise:

determining a frequency that each of the plurality of potential phrases occurs in a query log that includes transcriptions of previously spoken queries,

wherein presenting the list of the plurality of potential phrases comprises presenting the frequency that each of the plurality of potential phrases occurs in the query log.

15. The medium of claim 13 , wherein the operations further comprise:

determining a most frequent location of users who spoke each of the plurality of potential phrases,

wherein presenting the list of the plurality of potential phrases comprises presenting the most frequent location of users who spoke each of the plurality of potential phrases.

16. The medium of claim 13 , wherein presenting the list of the plurality of potential phrases comprises presenting data indicating whether an advertisement was presented in response to receiving each of the plurality of potential phrases as a spoken query.

17. The medium of claim 13 , wherein transcribing the voice input into the plurality of potential phrases that differ from the adword using the automatic speech recognizer comprises:

providing audio of the voice input as an input to an acoustic model that identifies candidate phonemes of the audio of the voice input;

providing data identifying the phonemes that likely correspond to the audio of the voice input as an input to a language model that identifies candidate transcriptions of the candidate phonemes; and

selecting, from among the candidate transcriptions of the candidate phonemes, the potential phrases that differ from the adword.

18. The medium of claim 17 , wherein selecting, from among the candidate transcriptions of the candidate phonemes, the plurality of potential phrases that differ from the adword comprises selecting the candidate transcriptions that are more likely to trigger displaying an advertisement.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2018
From: ALEKSIC, PETAR; MORENO MENGIBAR, PEDRO J.
To: GOOGLE INC.
Reel/Frame 046500/0281 →
CHANGE OF NAME Recorded Jul 30, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 046660/0258 →
Continuity (2)
Continuation 14710928 · May 13, 2015
Related Publication 20190026787A1 · Jan 24, 2019