IP Library Patent Application 12976972
Patent Application
App. No. 12/976,972

Speech to Text Conversion

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
12/976,972
Abstract

Methods, computer program products and systems are described for speech-to-text conversion. A voice input is received from a user of an electronic device and contextual metadata is received that describes a context of the electronic device at a time when the voice input is received. Multiple base language models are identified, where each base language model corresponds to a distinct textual corpus of content. Using the contextual metadata, an interpolated language model is generated based on contributions from the base language models. The contributions are weighted according to a weighting for each of the base language models. The interpolated language model is used to convert the received voice input to a textual output. The voice input is received at a computer server system that is remote to the electronic device. The textual output is transmitted to the electronic device.

Claims (62)

1 . A computer-implemented speech-to-text conversion method, comprising:

receiving a voice input from a user of an electronic device and contextual metadata that describes a context of the electronic device at a time when the voice input is received;

identifying a plurality of base language models, wherein each base language model corresponds to a distinct textual corpus of content;

using the contextual metadata to generate an interpolated language model based on contributions from the plurality of base language models, wherein the contributions are weighting according to a weighting for each of the base language models; and

using the interpolated language model to convert the received voice input to a textual output.

2 . The method of claim 1 , wherein the voice input is received at a computer server system that is remote to the electronic device, the method further comprising:

transmitting the textual output to the electronic device.

3 . The method of claim 1 , wherein the contextual metadata identifies a field in an electronic document on which the electronic device was focused when the voice input was received at the electronic device.

4 . The method of claim 1 , further comprising:

determining the weightings for each of the base language models based on the contextual metadata.

5 . The method of claim 1 , further comprising:

building one or more of the base language models based on text input data collected from a plurality of users and metadata that corresponds to the text input data.

6 . The method of claim 5 , wherein for a particular text input data the corresponding metadata identifies an input field that corresponds to the text input data.

7 . The method of claim 5 , wherein the text input data and the metadata are formed as individual pairs, the method further comprising:

forming a bipartite cluster graph of the individual pairs.

8 . The method of claim 7 , further comprising identifying clusters in the graph and using the clusters to generate the interpreted language model.

9 . The method of claim 8 , further comprising training the base language models by using sample voice utterances from a plurality of users of a plurality of electronic devices.

10 . A computer-implemented system for converting speech to text, the system comprising:

a plurality of base language models, each base language model corresponding to a particular semantic category;

an interpolated language model that is linked to the plurality of base language models; and

wherein each link between the interpolated language model and each of the base language models is associated with a weight.

11 . The system of claim 10 , wherein the weight for each link between the interpolated language mode and a base language model is based on an accuracy of the base language model in associating a voice input with a text output representing a conversion of the voice input into text.

12 . The system of claim 10 , wherein the weights represent likelihoods of usage in the interpolated language model matching usage in the particular base language model.

13 . The system of claim 12 , wherein the weightings are a function of the semantic category.

14 . The system of claim 13 , further comprising:

a network interface configured to:

receive a voice input; and

cause the interpolated language model to be applied to the voice input to generate a text output.

15 . The system of claim 14 , wherein the network interface is further configured to:

use metadata received with the voice input to match to the semantic category to determine weightings for the base language models from the interpolated language model.

16 . The system of claim 15 , wherein the system is configured to dynamically apply the weightings for the plurality of base language models in real-time substantially as the voice input is received by the network interface.

17 . A computer-readable storage device encoded with a computer program product, the computer program product including instructions for speech-to-text conversion that, when executed, cause data processing apparatus to perform operations comprising:

receiving a voice input from a user of an electronic device and contextual metadata that describes a context of the electronic device at a time when the voice input is received;

identifying a plurality of base language models, wherein each base language model corresponds to a distinct textual corpus of content;

using the contextual metadata to generate an interpolated language model based on contributions from the plurality of base language models, wherein the contributions are weighting according to a weighting for each of the base language models; and

using the interpolated language model to convert the received voice input to a textual output.

18 . The computer-readable storage device of claim 17 , wherein the voice input is received at a computer server system that is remote to the electronic device, the operations further comprising:

transmitting the textual output to the electronic device.

19 . The computer-readable storage device of claim 17 , wherein the contextual metadata identifies a field in an electronic document on which the electronic device was focused when the voice input was received at the electronic device.

20 . The computer-readable storage device of claim 17 , the operations further comprising:

determining weightings for each of the base language models based on the contextual metadata.

21 . The computer-readable storage device of claim 20 , wherein the interpolated language model is based on contributions from each of the base language models that are proportionate to the respective weightings for each of the base language models.

22 . The computer-readable storage device of claim 17 , the operations further comprising:

building one or more of the base language models based on text input data collected from a plurality of users and metadata that corresponds to the text input data.

23 . The computer-readable storage device of claim 22 , wherein for a particular text input data the corresponding metadata identifies an input field that corresponds to the text input data.

24 . The computer-readable storage device of claim 22 , wherein the text input data and the metadata are formed as individual pairs, the operations further comprising:

forming a bipartite cluster graph of the individual pairs.

25 . The computer-readable storage device of claim 24 , the operations further comprising:

identifying clusters in the graph and using the clusters to generate the interpreted language model.

26 . The computer-readable storage device of claim 25 , the operations further comprising:

training the base language models by using sample voice utterances from a plurality of users of a plurality of electronic devices.

27 . A computer-implemented method, comprising;

extracting pairs from a historical log of query search results that includes a plurality of search queries and corresponding search results, each pair including a query and a website that corresponds to a search result for the query;

generating a bipartite cluster graph based on the extracted pairs of queries and corresponding websites;

training a plurality of language models based on clusters identified in the bipartite cluster graph;

based on sample data obtained from input by one or more users into a web from, the sample data comprising one or more sample queries, identifying K clusters from the cluster graph that are most significant to the sample queries, K being an integer; and

generating an interpolated language model for the web form based on weighted contributions from the language models trained for each of the identified K clusters.

28 . The method of claim 27 further comprising:

receiving a voice input into the web form; and

based on the interpolated language model generating a text output that represents the voice input.

29 . The method of claim 28 , wherein the voice input is received from a user of an electronic device and the text output is generated by a computer server system that is remote to the electronic device, the method further comprising:

transmitting the text output to the electronic device.

Assignments (2)
CHANGE OF NAME Recorded Oct 6, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044142/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2011
From: BALLINGER, BRANDON M.; SCHALKWYK, JOHAN; COHEN, MICHAEL H.; LUC ALLAUZEN, CYRIL GEORGES; RILEY, MICHAEL D.
To: GOOGLE INC.
Reel/Frame 025999/0930 →