AUTOMATIC SPEECH RECOGNITION SYSTEM CONTEXTUALLY BIASED FOR MEDICAL SPEECH
Methods and systems of generating text representation of spoken medical speech are presented herein. Some methods may include the steps of providing a pre-trained automatic speech recognition (ASR) system stored in memory and executed on a processor; receiving, by the pre-trained ASR system, spoken medical speech; and generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, where the contextual language model may include medical terminology that is not included in a vocabulary used to train the pre-trained ASR system.
1 . A method of generating text of medical speech, the method comprising:
providing a pre-trained automatic speech recognition (ASR) system stored in memory and executed on a processor;
receiving, by the pre-trained ASR system, spoken medical speech; and
generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, wherein the contextual language model comprises medical terminology that is not included in a vocabulary used to train the pre-trained ASR system.
2 . The method of claim 1 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.
3 . The method of claim 1 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.
4 . The method of claim 1 , wherein the biased ASR system is a shallow fusion model.
5 . The method of claim 1 , wherein the medical terminology comprises a plurality of medical terms.
6 . The method of claim 1 , wherein the contextual language model is a contextual n-gram language model.
7 . The method of claim 6 , wherein the step of generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model to generate a textual representation of a medical term that is not included in the vocabulary used to train the pre-trained ASR system.
8 . The method of claim 1 , wherein the language model biases the ASR system during beam searching.
9 . The method of claim 1 , wherein the language model biases the ASR system before beam searching.
10 . A method of generating a medical report, comprising:
the method of claim 1 ; and,
writing a report based on the text of the medical speech.
11 . A system for generating text of spoken medical speech comprising:
an input interface configured to receive spoken medical speech;
a memory configured to store a plurality of processor-executable instruction, the memory including:
a pre-trained ASR system; and,
a contextual language model, wherein the contextual language model receives a plurality of medical terms; and,
a processor configured to execute the plurality of processor-executable instructions to perform operations including:
biasing the pre-trained ASR system using the contextual language model; and,
generating text of the spoken medical speech using the biased pre-trained ASR system, wherein at least one of the plurality of medical terms is not included in a vocabulary used to train the pre-trained ASR system.
12 . The system of claim 11 , wherein the biased pretrained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.
13 . The system of claim 11 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.
14 . The system of claim 12 , wherein generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model.
15 . The system of claim 11 , wherein the contextual language model biases the pre-trained ASR system during beam search decoding.
16 . A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for generating text of spoken medical speech, the instructions being executed by a processor to perform operations comprising:
providing a pre-trained automatic speech recognition (ASR) model;
biasing the pre-trained ASR model using a contextual language model, wherein the contextual language model comprises medical terminology; and,
generating text of the spoken medical speech by biasing the pre-trained ASR system using a contextual language model, wherein the contextual language model comprises medical terminology that is not included in a vocabulary used to train the pre-trained ASR system
17 . The non-transitory processor-readable storage medium of claim 16 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been jointly trained using the vocabulary.
18 . The non-transitory processor-readable storage medium of claim 16 , wherein the pre-trained ASR system comprises an acoustic model, a pronunciation model, and a language model that have been separately trained, and wherein the language model is trained using the vocabulary.
19 . The non-transitory processor-readable storage medium of claim 16 , wherein the contextual language model is a contextual n-gram language model.
20 . The non-transitory processor-readable storage medium of claim 19 , wherein generating text of the spoken medical speech comprises determining an n-gram score based on an overall model score generated by the pre-trained ASR system and a bias score generated by the contextual language model.