IP Library › Granted Patent US 12,725,617
Granted Patent B2
US 12,725,617 · App. 18/552,405 · Granted Sep 1, 2026

Self-learning end-to-end automatic speech recognition

Inventors: Cornelius Patrick Glackin (London, GB); Nigel Henry Cannings (London, GB)
Assignee: Verint Systems UK Limited
G10L17/04G10L13/02G10L17/02G10L17/14
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,617
App. No.
18/552,405
Granted
Sep 1, 2026
Kind
B2
Abstract

The present invention provides a system for self-learning end to end automatic speech recognition. The system comprises a memory. The memory stores a language model, a noise database and an automatic speech recognition, ASR, model. The system further comprises a processor. The processor may be hardware based or it may be cloud based. The processor is configured to run: an application programming interface (API); a text extractor; a language model builder; a pronunciation model; a text to speech engine; and a mixer. The text to speech engine may also be referred to as a text to speech converter. The system provides automatic updating of the ASR model to improve accuracy of the ASR functionality of the system without the need for complete retraining.

Claims (91)

1 . A system for self-learning end to end automatic speech recognition (ASR), the system comprising:

a memory; wherein the memory stores a language model, a noise database and an automatic speech recognition model;

a processor;

the processor configured to run:

an application programming interface (API);

a text extractor;

a language model builder;

a pronunciation model;

a text to speech engine; and

a mixer; and

the processor configured to:

receive client text documents via the API; and

using the text extractor:

determine out of vocabulary (OOV), words from the client text documents;

extract sentences containing OOV words from the client text documents;

use the extracted sentences to create an n-gram based language model using the language model builder;

add the n-gram based language model to the language model stored in memory;

process the extracted sentences using the pronunciation model to produce a phonetic sequence of the OOV words;

generate a plurality of audio samples for each sentence using the text to speech engine;

in response to generating the plurality of audio samples for each sentence using the text to speech engine:

augment, using the mixer, the plurality of audio samples for each sentence generated using the text to speech engine with noise from the noise database; and

fine tune the ASR model stored in memory using noise augmented audio samples.

2 . The system of claim 1 wherein the memory stores speakers and styles that the text to speech engine can emulate.

3 . The system of claim 2 wherein the processor is configured to:

receive client audio via the API;

process the audio using the ASR model to produce a speaker separated transcript;

analyze confidence of the speaker separated transcript for each speaker utterance; and

if the confidence of the speaker separated transcript of a speaker's utterance is below a threshold confidence limit:

extract an x-vector voice print of the speaker from the client audio;

match the x-vector voice print to a nearest matching speaker in the speakers and styles in the memory;

select the nearest matching speaker from the speakers and styles stored in the memory;

process the speaker's utterance using a processing model to produce a phonetic sequence of the speaker's utterance;

process the phonetic sequence of the speaker's utterance with the nearest matching speaker using the text to speech engine to generate a plurality of nearest matching speaker audio samples;

augment, using the mixer, the plurality of nearest matching speaker audio samples with noise from the noise database; and

fine tune the ASR model stored in memory using noise augmented nearest matching speaker audio samples.

4 . The system of claim 3 wherein the processor is further configured to concatenate the x-vector voice print with the phonetic sequence of the speaker's utterance;

use a resulting concatenation to generate a problematic speaker audio sample using the text to speech engine;

augment, using the mixer, the problematic speaker audio sample with noise from the noise database; and

fine tune the ASR model stored in memory using a noise augmented problematic speaker audio sample.

5 . The system of claim 1 wherein the memory stores ASR training text and the processor is further configured to:

process the ASR training text using the pronunciation model to produce a phonetic sequence of the ASR training text;

use the text to speech engine to generate one or more ASR training text audio samples using the phonetic sequence of the ASR training text and a first speaker and style from the speakers and styles in the memory;

augment, using the mixer, the one or more ASR training text audio samples with noise from the noise database; and

fine tune the ASR model stored in memory using the one or more noise augmented ASR training text audio samples.

6 . The system of claim 5 wherein the processor is further configured to:

process the ASR training text using the pronunciation model to produce a second phonetic sequence of the ASR training text;

use the text to speech engine to generate one or more second ASR training text audio samples using the second phonetic sequence of the ASR training text and a second speaker and style from the speakers and styles in the memory;

augment, using the mixer, the one or more second ASR training text audio samples with noise from the noise database to generate one or more second noise augmented ASR training text audio samples; and

fine tune the ASR model stored in memory using one or more second noise augmented ASR training text audio samples.

7 . The system of claim 1 wherein the processor uses interpolation using an n-gram language modelling toolkit that uses Neyser Ney Smoothing to add the n-gram based language model to the language model stored in memory.

8 . The system of claim 1 wherein the pronunciation model is a Grapheme-to-phoneme model (G2P).

9 . The system of claim 1 wherein the language model stored in the memory comprises a lexicon; and the OOV words are words that are words in the client text documents that are not present in the lexicon.

10 . A computer implemented method for self-learning end to end automatic speech recognition (ASR), the method comprising:

receiving client text documents via an application programming interface (API);

determining out of vocabulary (OOV) words from the client text documents using a text extractor;

extracting sentences containing OOV words from the client text documents using the text extractor;

using the extracted sentences and a language model builder to create an n-gram based language model;

adding the n-gram based language model to a language model stored in memory;

processing the extracted sentences using a pronunciation model to produce a phonetic sequence of the OOV words; and

generating a plurality of audio samples for each sentence using a text to speech engine; and

in response to generating the plurality of audio samples for each sentence using the text to speech engine:

augmenting, using a mixer, the plurality of audio samples for each sentence generated using the text to speech engine with noise from a noise database; and

fine tuning an ASR model stored in memory using noise augmented audio samples.

11 . The method of claim 10 wherein the memory stores speakers and styles that the text to speech engine can emulate.

12 . The method of claim 11 wherein the method further comprises:

receiving client audio via the API;

processing the audio using the ASR model to produce a speaker separated transcript;

analyze confidence of the speaker separated transcript for each speaker utterance;

if the confidence of the speaker separated transcript of a speaker's utterance is below a threshold confidence limit:

extracting an x-vector voice print of the speaker from the client audio;

matching the x-vector voice print to a nearest matching speaker in the speakers and styles in the memory;

selecting the nearest matching speaker from the speakers and styles stored in the memory;

processing the speaker's utterance using a processing model to produce a phonetic sequence of the speaker's utterance;

processing the phonetic sequence of the speaker's utterance with the nearest matching speaker using the text to speech engine to generate a plurality of nearest matching speaker audio samples;

augmenting, using the mixer, the plurality of nearest matching speaker audio samples with noise from the noise database; and

fine tuning the ASR model stored in memory using noise augmented nearest matching speaker audio samples.

13 . The method of claim 12 wherein the method further comprises:

concatenating the x-vector voice print with the phonetic sequence of the speaker's utterance;

using a resulting concatenation to generate a problematic speaker audio sample using the text to speech engine;

augmenting, using the mixer, the problematic speaker audio sample with noise from the noise database; and

fine tuning the ASR model stored in memory using a noise augmented problematic speaker audio sample.

14 . The method of claim 10 wherein the memory stores ASR training text and the method further comprises:

processing the ASR training text using the pronunciation model to produce a phonetic sequence of the ASR training text;

using the text to speech engine to generate one or more ASR training text audio samples using the phonetic sequence of the ASR training text and a first speaker and style from the speakers and styles in the memory;

augmenting, using the mixer, the one or more ASR training text audio samples with noise from the noise database; and

fine tuning the ASR model stored in memory using the one or more noise augmented ASR training text audio samples.

15 . The method of claim 14 further comprising:

processing the ASR training text using the pronunciation model to produce a second phonetic sequence of the ASR training text;

using the text to speech engine to generate one or more second ASR training text audio samples using the second phonetic sequence of the ASR training text and a second speaker and style from the speakers and styles in the memory;

augmenting, using the mixer, the one or more second ASR training text audio samples with noise from the noise database to generate one or more second noise augmented ASR training text audio samples; and

fine tuning the ASR model stored in memory using one or more second noise augmented ASR training text audio samples.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2025
From: INTELLIGENT VOICE LIMITED
To: VERINT SYSTEMS UK LIMITED
Reel/Frame 070814/0757 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: CANNINGS, NIGEL HENRY, MR.; GLACKIN, CORNELIUS PATRICK, DR.
To: INTELLIGENT VOICE LIMITED
Reel/Frame 066030/0267 →
Priority Claims (1)
GB 2116251 · Nov 11, 2021 · national
Continuity (1)
Related Publication 20240119942A1 · Apr 11, 2024
References Cited (30)
US 6044343A · Cong · 2000 [cited by examiner]
US 6308151B1 · Smith · 2001 [cited by examiner]
US 8457959B2 · Kaiser · 2013 [cited by examiner]
US 8583432B1 · Biadsy · 2013 [cited by examiner]
US 9990926B1 · Pearce · 2018 [cited by examiner]
US 10614827B1 · Korjani · 2020 [cited by examiner]
US 11024194B1 · Beigman Klebanov · 2021 [cited by examiner]
US 20040230420A1 · Kadambe · 2004 [cited by examiner]
US 20080153537A1 · Khawand · 2008 [cited by examiner]
US 20110238407A1 · Kent · 2011 [cited by examiner]
US 20120109649A1 · Talwar · 2012 [cited by examiner]
US 20150058019A1 · Chen · 2015 [cited by examiner]
US 20150161370A1 · North · 2015 [cited by examiner]
US 20150293908A1 · Mathur · 2015 [cited by examiner]
US 20160027452A1 · Kalinli-Akbacak · 2016 [cited by examiner]
US 20160140951A1 · Agiomyrgiannakis · 2016 [cited by examiner]
US 20160379626A1 · Deisher et al. · 2016 [cited by applicant]
US 20180144736A1 · Huang · 2018 [cited by examiner]
US 20180233129A1 · Bakish · 2018 [cited by examiner]
US 20180295240A1 · Dickins · 2018 [cited by examiner]
US 20190012594A1 · Fukuda · 2019 [cited by examiner]
US 20200111484A1 · Aleksic et al. · 2020 [cited by applicant]
US 20200243094A1 · Thomson · 2020 [cited by examiner]
US 20200312337A1 · Stafylakis · 2020 [cited by examiner]
US 20200357388A1 · Zhao et al. · 2020 [cited by applicant]
US 20200372897A1 · Battenberg · 2020 [cited by examiner]
US 20210043220A1 · Baek · 2021 [cited by examiner]
US 20210248801A1 · Li · 2021 [cited by examiner]
US 20220392478A1 · Hijazi · 2022 [cited by examiner]
Search Report and Written Opinion of the International Searching Authority in PCT application PCT/EP2022/080969 mailed on Feb. 7, 2023 (8 pages). [cited by applicant]