Self-learning end-to-end automatic speech recognition
The present invention provides a system for self-learning end to end automatic speech recognition. The system comprises a memory. The memory stores a language model, a noise database and an automatic speech recognition, ASR, model. The system further comprises a processor. The processor may be hardware based or it may be cloud based. The processor is configured to run: an application programming interface (API); a text extractor; a language model builder; a pronunciation model; a text to speech engine; and a mixer. The text to speech engine may also be referred to as a text to speech converter. The system provides automatic updating of the ASR model to improve accuracy of the ASR functionality of the system without the need for complete retraining.
1 . A system for self-learning end to end automatic speech recognition (ASR), the system comprising:
a memory; wherein the memory stores a language model, a noise database and an automatic speech recognition model;
a processor;
the processor configured to run:
an application programming interface (API);
a text extractor;
a language model builder;
a pronunciation model;
a text to speech engine; and
a mixer; and
the processor configured to:
receive client text documents via the API; and
using the text extractor:
determine out of vocabulary (OOV), words from the client text documents;
extract sentences containing OOV words from the client text documents;
use the extracted sentences to create an n-gram based language model using the language model builder;
add the n-gram based language model to the language model stored in memory;
process the extracted sentences using the pronunciation model to produce a phonetic sequence of the OOV words;
generate a plurality of audio samples for each sentence using the text to speech engine;
in response to generating the plurality of audio samples for each sentence using the text to speech engine:
augment, using the mixer, the plurality of audio samples for each sentence generated using the text to speech engine with noise from the noise database; and
fine tune the ASR model stored in memory using noise augmented audio samples.
2 . The system of claim 1 wherein the memory stores speakers and styles that the text to speech engine can emulate.
3 . The system of claim 2 wherein the processor is configured to:
receive client audio via the API;
process the audio using the ASR model to produce a speaker separated transcript;
analyze confidence of the speaker separated transcript for each speaker utterance; and
if the confidence of the speaker separated transcript of a speaker's utterance is below a threshold confidence limit:
extract an x-vector voice print of the speaker from the client audio;
match the x-vector voice print to a nearest matching speaker in the speakers and styles in the memory;
select the nearest matching speaker from the speakers and styles stored in the memory;
process the speaker's utterance using a processing model to produce a phonetic sequence of the speaker's utterance;
process the phonetic sequence of the speaker's utterance with the nearest matching speaker using the text to speech engine to generate a plurality of nearest matching speaker audio samples;
augment, using the mixer, the plurality of nearest matching speaker audio samples with noise from the noise database; and
fine tune the ASR model stored in memory using noise augmented nearest matching speaker audio samples.
4 . The system of claim 3 wherein the processor is further configured to concatenate the x-vector voice print with the phonetic sequence of the speaker's utterance;
use a resulting concatenation to generate a problematic speaker audio sample using the text to speech engine;
augment, using the mixer, the problematic speaker audio sample with noise from the noise database; and
fine tune the ASR model stored in memory using a noise augmented problematic speaker audio sample.
5 . The system of claim 1 wherein the memory stores ASR training text and the processor is further configured to:
process the ASR training text using the pronunciation model to produce a phonetic sequence of the ASR training text;
use the text to speech engine to generate one or more ASR training text audio samples using the phonetic sequence of the ASR training text and a first speaker and style from the speakers and styles in the memory;
augment, using the mixer, the one or more ASR training text audio samples with noise from the noise database; and
fine tune the ASR model stored in memory using the one or more noise augmented ASR training text audio samples.
6 . The system of claim 5 wherein the processor is further configured to:
process the ASR training text using the pronunciation model to produce a second phonetic sequence of the ASR training text;
use the text to speech engine to generate one or more second ASR training text audio samples using the second phonetic sequence of the ASR training text and a second speaker and style from the speakers and styles in the memory;
augment, using the mixer, the one or more second ASR training text audio samples with noise from the noise database to generate one or more second noise augmented ASR training text audio samples; and
fine tune the ASR model stored in memory using one or more second noise augmented ASR training text audio samples.
7 . The system of claim 1 wherein the processor uses interpolation using an n-gram language modelling toolkit that uses Neyser Ney Smoothing to add the n-gram based language model to the language model stored in memory.
8 . The system of claim 1 wherein the pronunciation model is a Grapheme-to-phoneme model (G2P).
9 . The system of claim 1 wherein the language model stored in the memory comprises a lexicon; and the OOV words are words that are words in the client text documents that are not present in the lexicon.
10 . A computer implemented method for self-learning end to end automatic speech recognition (ASR), the method comprising:
receiving client text documents via an application programming interface (API);
determining out of vocabulary (OOV) words from the client text documents using a text extractor;
extracting sentences containing OOV words from the client text documents using the text extractor;
using the extracted sentences and a language model builder to create an n-gram based language model;
adding the n-gram based language model to a language model stored in memory;
processing the extracted sentences using a pronunciation model to produce a phonetic sequence of the OOV words; and
generating a plurality of audio samples for each sentence using a text to speech engine; and
in response to generating the plurality of audio samples for each sentence using the text to speech engine:
augmenting, using a mixer, the plurality of audio samples for each sentence generated using the text to speech engine with noise from a noise database; and
fine tuning an ASR model stored in memory using noise augmented audio samples.
11 . The method of claim 10 wherein the memory stores speakers and styles that the text to speech engine can emulate.
12 . The method of claim 11 wherein the method further comprises:
receiving client audio via the API;
processing the audio using the ASR model to produce a speaker separated transcript;
analyze confidence of the speaker separated transcript for each speaker utterance;
if the confidence of the speaker separated transcript of a speaker's utterance is below a threshold confidence limit:
extracting an x-vector voice print of the speaker from the client audio;
matching the x-vector voice print to a nearest matching speaker in the speakers and styles in the memory;
selecting the nearest matching speaker from the speakers and styles stored in the memory;
processing the speaker's utterance using a processing model to produce a phonetic sequence of the speaker's utterance;
processing the phonetic sequence of the speaker's utterance with the nearest matching speaker using the text to speech engine to generate a plurality of nearest matching speaker audio samples;
augmenting, using the mixer, the plurality of nearest matching speaker audio samples with noise from the noise database; and
fine tuning the ASR model stored in memory using noise augmented nearest matching speaker audio samples.
13 . The method of claim 12 wherein the method further comprises:
concatenating the x-vector voice print with the phonetic sequence of the speaker's utterance;
using a resulting concatenation to generate a problematic speaker audio sample using the text to speech engine;
augmenting, using the mixer, the problematic speaker audio sample with noise from the noise database; and
fine tuning the ASR model stored in memory using a noise augmented problematic speaker audio sample.
14 . The method of claim 10 wherein the memory stores ASR training text and the method further comprises:
processing the ASR training text using the pronunciation model to produce a phonetic sequence of the ASR training text;
using the text to speech engine to generate one or more ASR training text audio samples using the phonetic sequence of the ASR training text and a first speaker and style from the speakers and styles in the memory;
augmenting, using the mixer, the one or more ASR training text audio samples with noise from the noise database; and
fine tuning the ASR model stored in memory using the one or more noise augmented ASR training text audio samples.
15 . The method of claim 14 further comprising:
processing the ASR training text using the pronunciation model to produce a second phonetic sequence of the ASR training text;
using the text to speech engine to generate one or more second ASR training text audio samples using the second phonetic sequence of the ASR training text and a second speaker and style from the speakers and styles in the memory;
augmenting, using the mixer, the one or more second ASR training text audio samples with noise from the noise database to generate one or more second noise augmented ASR training text audio samples; and
fine tuning the ASR model stored in memory using one or more second noise augmented ASR training text audio samples.