METHOD AND APPARATUS FOR IMPROVING EFFICIENCY OF AUTOMATIC SPEECH RECOGNITION
A method and an apparatus for improving efficiency of automatic speech recognition (ASR) is provided. The apparatus includes a call analytics server comprising a processor and a memory, which perform the method. The method comprises removing non-speech portions from a call audio to produce a pre-processed audio, sending the pre-processed audio from the CAS to an ASR engine, and receiving a call text from the ASR engine. The call text is the speech-to-text conversion of the pre-processed audio, and the call text comprises text corresponding to the speech in the pre-processed audio.
1 . A method for improving efficiency of automatic speech recognition (ASR), the method comprising:
removing, at a call analytics server (CAS), non-speech portions from a call audio to produce a pre-processed audio;
sending the pre-processed audio from the CAS to an ASR engine; and
receiving, at the CAS, a call text from the ASR engine, wherein the call text is the speech-to-text conversion of the pre-processed audio, and the call text comprises text corresponding to the speech in the pre-processed audio.
2 . The method of claim 1 , further comprising receiving, at the CAS, the call audio from a call audio source.
3 . The method of claim 1 , wherein the removing comprises removing portions comprising at least one of beeps, rings, silence, noise, or music.
4 . The method of claim 1 , further comprising performing offset correction on the call text.
5 . The method of claim 4 , wherein the performing offset correction comprises:
adding, at the CAS, to a timestamp of a text in the call text, time corresponding to a duration of the non-speech portion occurring prior to the speech corresponding to the text.
6 . An apparatus for improving efficiency of automatic speech recognition (ASR), the apparatus comprising:
a processor; and
a memory communicably coupled to the processor, wherein the memory comprises computer-executable instructions, which when executed using the processor, perform a method comprising:
removing, at a call analytics server (CAS), non-speech portions from a call audio to produce a pre-processed audio,
sending the pre-processed audio from the CAS to an ASR engine, and
receiving, at the CAS, a call text from the ASR engine, wherein the call text is the speech-to-text conversion of the pre-processed audio, and the call text comprises text corresponding to the speech in the pre-processed audio.
7 . The apparatus of claim 6 , wherein the method further comprises receiving, at the CAS, the call audio from a call audio source.
8 . The apparatus of claim 6 , wherein the removing comprises removing portions comprising at least one of beeps, rings, silence, noise, or music.
9 . The apparatus of claim 6 , further comprising performing offset correction on the call text.
10 . The method of claim 9 , wherein the performing offset correction comprises:
adding, at the CAS, to a timestamp of a text in the call text, time corresponding to a duration of the non-speech portion occurring prior to the speech corresponding to the text.