IP Library › Granted Patent US 12,165,638
Granted Patent B2
US 12,165,638 · App. 17/659,330 · Granted Dec 10, 2024

Personalizable probabilistic models

Inventors: Leonid Aleksandrovich Velikovich (New York, NY); Petar Stanisa Aleksic (Jersey City, NJ)
Assignee: Google LLC
G10L15/197G06F40/166G06F40/284G06F40/295G10L15/063G10L15/22G10L15/32G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,165,638
App. No.
17/659,330
Granted
Dec 10, 2024
Kind
B2
Abstract

A method includes receiving audio data corresponding to an utterance spoken by a user and processing, using a first recognition model, the audio data to generate a non-contextual candidate hypothesis as output from the first recognition model. The non-contextual candidate hypothesis has a corresponding likelihood score assigned by the first recognition model. The method also includes generating, using a second recognition model configured to receive personal context information, a contextual candidate hypothesis that includes a personal named entity. The method also includes scoring, based on the personal context information and the corresponding likelihood score assigned to the non-contextual candidate hypothesis, the contextual candidate hypothesis relative to the non-contextual candidate hypotheses. Based on the scoring of the contextual candidate hypothesis relative to the non-contextual candidate hypothesis, the method also includes generating a transcription of the utterance by selecting one of the contextual candidate hypothesis or the non-contextual candidate hypothesis.

Claims (56)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance spoken by a user;

processing the audio data, using a first recognition model trained by a training process as an aggregate devoid of any user-specific data, to generate a non-contextual candidate hypothesis as output from the first recognition model, the non-contextual candidate hypothesis corresponding to a candidate transcription of the utterance and having a corresponding likelihood score assigned by the first recognition model to the non-contextual candidate hypothesis, wherein the first recognition model is devoid of incorporating any user-specific data when generating the non-contextual candidate hypothesis as output from the first recognition model;

separately processing the audio data, using a second recognition model configured to receive personal contextual information, to generate a contextual candidate hypothesis as output from the second recognition model that includes a personal named entity, the contextual candidate hypothesis corresponding to another candidate transcription for the utterance;

scoring, based on the personal contextual information and the corresponding likelihood score assigned to the non-contextual candidate hypothesis, the contextual candidate hypothesis relative to the non-contextual candidate hypotheses; and

based on the scoring of the contextual candidate hypothesis relative to the non-contextual candidate hypothesis, generating a transcription of the utterance spoken by the user by selecting one of the contextual candidate hypothesis or the non-contextual candidate hypothesis,

wherein the training process trains the first recognition model by:

receiving a training corpus comprising a plurality of utterances, each utterance of the plurality of utterances paired with a corresponding ground-truth transcription;

for each corresponding ground-truth transcription in the training corpus of ground-truth transcriptions:

determining whether a personal named entity is identified in the corresponding ground-truth transcription; and

when the personal named entity is identified in the corresponding ground-truth transcription, replacing the personal named entity with an entity class token; and

training the first recognition model by adjusting coefficients of the first recognition model based on the plurality of utterances and the corresponding ground-truth transcriptions that include personal named entities replaced by the entity class tokens.

2. The computer-implemented method of claim 1 , wherein the first recognition model processes the audio data to generate the non-contextual candidate hypothesis without incorporating any of the personal contextual information associated with the user.

3. The computer-implemented method of claim 1 , wherein the first recognition model comprises an end-to-end speech recognition model configured to generate the corresponding likelihood score for the non-contextual candidate hypothesis.

4. The computer-implemented method of claim 1 , wherein:

the second recognition model comprises an end-to-end speech recognition model configured to generate the contextual candidate hypothesis; and

wherein generating the contextual candidate hypothesis that includes the personal named entity comprises processing, using the end-to-end speech recognition model, the audio data to generate the contextual candidate hypothesis based on the personal context information.

5. The computer-implemented method of claim 4 , wherein the end-to-end speech recognition model is configured to generate the contextual candidate hypothesis after the non-contextual candidate hypothesis is generated as output from the first recognition model.

6. The computer-implemented method of claim 4 , wherein the end-to-end speech recognition model is configured to generate the contextual candidate hypothesis in parallel with the processing of the audio data to generate the non-contextual candidate hypothesis as output from first recognition model.

7. The computer-implemented method of claim 1 , wherein the second recognition model comprises a language model.

8. The computer-implemented method of claim 1 , wherein the personal context information associated with the user comprises personal named entities specific to the user, the personal named entities comprising at least one of:

contact names in a personal contact list of the user;

names in in a media library associated with the user;

names of installed applications;

names of nearby locations; or

user-defined entity names.

9. The computer-implemented method of claim 1 , wherein the personal contextual information associated with the user further indicates a user history related to the personal named entity specific to the user.

10. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance spoken by a user;

processing the audio data, using a first recognition model trained by a training process as an aggregate devoid of any user-specific data, to generate a non-contextual candidate hypothesis as output from the first recognition model, the non-contextual candidate hypothesis corresponding to a candidate transcription of the utterance and having a corresponding likelihood score assigned by the first recognition model to the non-contextual candidate hypothesis, wherein the first recognition model is devoid of incorporating any user-specific data when generating the non-contextual candidate hypothesis as output from the first recognition model;

separately processing the audio data, using a second recognition model configured to receive personal contextual information, to generate a contextual candidate hypothesis as output from the second recognition model that includes a personal named entity, the contextual candidate hypothesis corresponding to another candidate transcription for the utterance;

scoring, based on the personal context information and the corresponding likelihood score assigned to the non-contextual candidate hypothesis, the contextual candidate hypothesis relative to the non-contextual candidate hypotheses; and

based on the scoring of the contextual candidate hypothesis relative to the non-contextual candidate hypothesis, generating a transcription of the utterance spoken by the user by selecting one of the contextual candidate hypothesis or the non-contextual candidate hypothesis,

wherein the training process trains the first recognition model by:

receiving a training corpus comprising a plurality of utterances, each utterance of the plurality of utterances paired with a corresponding ground-truth transcription;

for each corresponding ground-truth transcription in the training corpus of ground-truth transcriptions:

determining whether a personal named entity is identified in the corresponding ground-truth transcription; and

when the personal named entity is identified in the corresponding ground-truth transcription, replacing the personal named entity with an entity class token; and

training the first recognition model by adjusting coefficients of the first recognition model based on the plurality of utterances and the corresponding ground-truth transcriptions that include personal named entities replaced by the entity class tokens.

11. The system of claim 10 , wherein the first recognition model processes the audio data to generate the non-contextual candidate hypothesis without incorporating any of the personal contextual information associated with the user.

12. The system of claim 10 , wherein the first recognition model comprises an end-to-end speech recognition model configured to generate the corresponding likelihood score for the non-contextual candidate hypothesis.

13. The system of claim 10 , wherein:

the second recognition model comprises an end-to-end speech recognition model configured to generate the contextual candidate hypothesis; and

wherein generating the contextual candidate hypothesis that includes the personal named entity comprises processing, using the end-to-end speech recognition model, the audio data to generate the contextual candidate hypothesis based on the personal context information.

14. The system of claim 13 , wherein the end-to-end speech recognition model is configured to generate the contextual candidate hypothesis after the non-contextual candidate hypothesis is generated as output from the first recognition model.

15. The system of claim 13 , wherein the end-to-end speech recognition model is configured to generate the contextual candidate hypothesis in parallel with the processing of the audio data to generate the non-contextual candidate hypothesis as output from first recognition model.

16. The system of claim 10 , wherein the second recognition model comprises a language model.

17. The system of claim 10 , wherein the personal context information associated with the user comprises personal named entities specific to the user, the personal named entities comprising at least one of:

contact names in a personal contact list of the user;

names in in a media library associated with the user;

names of installed applications;

names of nearby locations; or

user-defined entity names.

18. The system of claim 10 , wherein the personal contextual information associated with the user further indicates a user history related to the personal named entity specific to the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: VELIKOVICH, LEONID ALEKSANDROVICH; ALEKSIC, PETAER STANISA
To: GOOGLE LLC
Reel/Frame 059621/0472 →
Continuity (1)
Related Publication 20230335125A1 · Oct 19, 2023