IP Library Granted Patent US 9,263,033
Granted Patent B2
US 9,263,033 · App. 14/314,295 · Granted Feb 16, 2016

Utterance selection for automated speech recognizer training

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,263,033
App. No.
14/314,295
Granted
Feb 16, 2016
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a set of training utterances. The methods, systems, and apparatus include actions of obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances and selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances. Additional actions include selecting a particular candidate utterance from the subset of the initial set of utterances and determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution. Further actions include adding the particular candidate utterance to the set of training utterances.

Claims (70)

1. A computer-implemented method comprising:

obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances;

selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances;

selecting a particular candidate utterance from the subset of the initial set of utterances;

determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution; and

adding the particular candidate utterance to the set of training utterances.

2. The method of claim 1 , wherein the characteristics comprise two or more of: sub-words included in utterances, gender of speaker, accent of speaker, age of speaker, application that the utterance originates, or confidence score.

3. The method of claim 1 , wherein obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances comprises:

obtaining the initial set of candidate utterances;

calculating a distribution of sub-words in the initial set of candidate utterances based on transcriptions associated with the initial set of candidate utterances;

calculating a distribution of another characteristic in the initial set of candidate utterances; and

generating the target multi-dimensional distribution from the calculated distributions.

4. The method of claim 1 , wherein selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances comprises:

filtering the initial set of candidate utterances based on the speech recognition confidence scores to obtain the subset of the initial set of candidate utterances.

5. The method of claim 1 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

obtaining a set of utterances consisting of candidate utterances that were previously added from the initial set of candidate utterances.

6. The method of claim 1 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

determining a first multi-dimensional distribution of characteristics in the set of training utterances without the candidate utterance;

determining a second multi-dimensional distribution of characteristics in the set of training utterances with the candidate utterance;

determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution; and

in response to determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution, determining that adding the particular candidate utterance to the set of training utterances reduces the divergence of the multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution.

7. The method of claim 1 , comprising:

training an automated speech recognizer using the set of training utterances including the particular candidate utterance.

8. A system comprising:

one or more computers; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances;

selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances;

selecting a particular candidate utterance from the subset of the initial set of utterances;

determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution; and

adding the particular candidate utterance to the set of training utterances.

9. The system of claim 8 , wherein the characteristics comprise two or more of:

sub-words included in utterances, gender of speaker, accent of speaker, age of speaker, application that the utterance originates, or confidence score.

10. The system of claim 8 , wherein obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances comprises:

obtaining the initial set of candidate utterances;

calculating a distribution of sub-words in the initial set of candidate utterances based on transcriptions associated with the initial set of candidate utterances;

calculating a distribution of another characteristic in the initial set of candidate utterances; and

generating the target multi-dimensional distribution from the calculated distributions.

11. The system of claim 8 , wherein selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances comprises:

filtering the initial set of candidate utterances based on the speech recognition confidence scores to obtain the subset of the initial set of candidate utterances.

12. The system of claim 8 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

obtaining a set of utterances consisting of candidate utterances that were previously added from the initial set of candidate utterances.

13. The system of claim 8 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

determining a first multi-dimensional distribution of characteristics in the set of training utterances without the candidate utterance;

determining a second multi-dimensional distribution of characteristics in the set of training utterances with the candidate utterance;

determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution; and

in response to determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution, determining that adding the particular candidate utterance to the set of training utterances reduces the divergence of the multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution.

14. The system of claim 8 , the operations comprising:

training an automated speech recognizer using the set of training utterances including the particular candidate utterance.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances;

selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances;

selecting a particular candidate utterance from the subset of the initial set of utterances;

determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution; and

adding the particular candidate utterance to the set of training utterances.

16. The medium of claim 15 , wherein the characteristics comprise two or more of: sub-words included in utterances, gender of speaker, accent of speaker, age of speaker, application that the utterance originates, or confidence score.

17. The medium of claim 15 , wherein obtaining a target multi-dimensional distribution of characteristics in an initial set of candidate utterances comprises:

obtaining the initial set of candidate utterances;

calculating a distribution of sub-words in the initial set of candidate utterances based on transcriptions associated with the initial set of candidate utterances;

calculating a distribution of another characteristic in the initial set of candidate utterances; and

generating the target multi-dimensional distribution from the calculated distributions.

18. The medium of claim 15 , wherein selecting a subset of the initial set of candidate utterances based on speech recognition confidence scores associated with the candidate utterances comprises:

filtering the initial set of candidate utterances based on the speech recognition confidence scores to obtain the subset of the initial set of candidate utterances.

19. The medium of claim 15 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

obtaining a set of utterances consisting of candidate utterances that were previously added from the initial set of candidate utterances.

20. The medium of claim 15 , wherein determining that adding the particular candidate utterance to a set of training utterances reduces a divergence of a multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution comprises:

determining a first multi-dimensional distribution of characteristics in the set of training utterances without the candidate utterance;

determining a second multi-dimensional distribution of characteristics in the set of training utterances with the candidate utterance;

determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution; and

in response to determining that the first multi-dimensional distribution is more divergent from the target multi-dimensional distribution than the second distribution, determining that adding the particular candidate utterance to the set of training utterances reduces the divergence of the multi-dimensional distribution of the characteristics in the set of training utterances from the target multi-dimensional distribution.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044566/0657 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2014
From: SIOHAN, OLIVIER; MENGIBAR, PEDRO J.
To: GOOGLE INC.
Reel/Frame 033594/0667 →