System and method for unsupervised and active learning for automatic speech recognition
View Patent ↗A system and method is provided for combining active and unsupervised learning for automatic speech recognition. This process enables a reduction in the amount of human supervision required for training acoustic and language models and an increase in the performance given the transcribed and un-transcribed data.
1. A method comprising:
identifying, in a database of utterances, transcribed utterances and un-transcribed utterances;
selecting, via a processor, transcription candidate utterances from the un-transcribed utterances using confidence scores of the un-transcribed utterances;
transcribing the transcription candidate utterances, to yield additional transcribed utterances; and
adding the additional transcribed utterances to the database of utterances.
2. The method of claim 1 , the method further comprising:
determining the confidence scores using an acoustic model and a language model.
3. The method of claim 1 , wherein word posterior probability estimates are used for confidence scores associated with the database of utterances.
4. The method of claim 1 , wherein the transcribing the transcription candidate utterances is conducted by a human being.
5. The method of claim 1 , wherein the transcribing the transcription candidate utterances is conducted by the processor.
6. The method of claim 1 , further comprising:
upon adding the additional transcribed utterances to the database of utterances, removing the additional transcribed utterances from the un-transcribed utterances.
7. A system comprising:
a processor; and
a computer-readable storage medium having instructions stored which, when executed on the processor, perform operations comprising:
identifying, in a database of utterances, transcribed utterances and un-transcribed utterances;
selecting transcription candidate utterances from the un-transcribed utterances;
transcribing the transcription candidate utterances, to yield additional transcribed utterances using confidence scores of the un-transcribed utterances; and
adding the additional transcribed utterances to the database of utterances.
8. The system of claim 7 , wherein the non-transitory computer-readable storage medium stores additional instructions which, when executed on the processor, perform a method comprising:
determining the confidence scores using an acoustic model and a language model.
9. The system of claim 7 , wherein word posterior probability estimates are used for confidence scores associated with the database of utterances.
10. The system of claim 7 , wherein the transcribing the transcription candidate utterances is conducted by a human being.
11. The system of claim 7 , wherein the transcribing the transcription candidate utterances is conducted by the processor.
12. The system of claim 7 , wherein the computer-readable storage medium has additional instructions stored which result in the operations further comprising:
upon adding the additional transcribed utterances to the database of utterances, removing the additional transcribed utterances from the un-transcribed utterances.
13. A computer-readable storage device having instructions stored which, when executed on a computing device, cause the computing device to perform operations comprising:
identifying, in a database of utterances, transcribed utterances and un-transcribed utterances;
selecting transcription candidate utterances from the un-transcribed utterances using confidence scores of the un-transcribed utterances;
transcribing the transcription candidate utterances, to yield additional transcribed utterances; and
adding the additional transcribed utterances to the database of utterances.
14. The computer-readable storage device of claim 13 , wherein the computer-readable storage device has additional instructions stored which result in the operations further comprising:
determining the confidence scores using an acoustic model and a language model.
15. The computer-readable storage device of claim 13 , wherein word posterior probability estimates are used for confidence scores associated with the database of utterances.
16. The computer-readable storage device of claim 13 , wherein the transcribing the transcription candidate utterances is conducted by a human.
17. The computer-readable storage device of claim 13 , wherein the computer-readable storage device has additional instructions stored which result in the operations further comprising:
upon adding the additional transcribed utterances to the database of utterances, removing the additional transcribed utterances from the un-transcribed utterances.