SEMIAUTOMATED RELAY METHOD AND APPARATUS
A method to transcribe communications includes the steps of obtaining a voice signal originating at a first device configured for verbal communication with a second device and providing the voice signal to an automated speech recognition system configured to transcribe the signal. Before determining a final transcription, the method obtains a first hypothesis transcription including one or more first words determined to be a transcription of at least a first portion of the voice signal and a second hypothesis transcription including a plurality of second words determined to be a transcription of at least a second portion of the voice signal that includes the first portion of the voice signal, determines one or more consistent words included in both hypothesis transcriptions, and in response, transmits the consistent words to the second device for presentation before the final transcription of the voice signal is provided to the second device.
1 . A method to transcribe communications, the method comprising:
obtaining a voice signal originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication;
providing the voice signal to an automated speech recognition system configured to transcribe the voice signal;
before a final transcription of the voice signal is determined by the automated speech recognition system, the method including:
obtaining a first hypothesis transcription generated by the automated speech recognition system, the first hypothesis transcription including one or more first words determined by the automated speech recognition system to be a transcription of at least a first portion of the voice signal;
obtaining a second hypothesis transcription generated by the automated speech recognition system, the second hypothesis transcription including a plurality of second words determined by the automated speech recognition system to be a transcription of at least a second portion of the voice signal that includes the first portion of the voice signal;
determining one or more consistent words that are included in both the one or more first words of the first hypothesis transcription and the plurality of second words of the second hypothesis transcription; and
in response to determining the one or more consistent words, transmitting the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before the final transcription of the voice signal is provided to the second device.
2 . The method of claim 1 , wherein the first hypothesis transcription is not provided to the second device.
3 . The method of claim 1 , wherein the voice signal is a portion of total voice signal originating at the first device during the communication session.
4 . The method of claim 1 , further comprising:
obtaining a third hypothesis transcription generated by the automated speech recognition system, the third hypothesis transcription including a plurality of third words determined by the automated speech recognition system to be a transcription of at least a third portion of the voice signal that includes the first portion of the voice signal.
5 . The method of claim 1 , further comprising:
obtaining the final transcription of the voice signal, the final transcription including a plurality of third words that the automated speech recognition system outputs together as the finalized transcription of the voice signal;
determining when a portion of the final transcription that corresponds to the one or more consistent words includes a corrected word that is different from any of the one or more consistent words; and
in response to determining that the consistent words include the corrected word, transmitting an indication of the corrected word to the second device such that the second device changes the presentation of the one or more consistent words to include the corrected word.
6 . At least one storage device configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 1 .
7 . A system comprising:
at least one processor; and
at least one memory device communicatively coupled to the at least one processor and configured to store one or more instructions that when executed by the at least one processor cause the system to perform operations comprising:
obtaining a voice signal originating at a first device during a communication session between the first device and a second device;
providing the voice signal to an automated speech recognition system configured to transcribe the voice signal;
obtaining a plurality of hypothesis transcriptions generated by the automated speech recognition system, each of the plurality of hypothesis transcriptions including one or more words determined by the automated speech recognition system to be a transcription of a portion of the voice signal;
determining one or more consistent words that are included in two or more of the plurality of hypothesis transcriptions; and
in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before a final transcription of the voice signal is provided to the second device.
8 . The system of claim 7 , wherein the plurality of hypothesis transcriptions are not provided to the second device.
9 . The system of claim 7 , wherein the automated speech recognition system is included in the system.
10 . The system of claim 7 , wherein the automated speech recognition system includes a plurality of differently tuned automated speech recognition engines and wherein each of the hypothesis transcriptions is generated by a different one of the ASR engines.
11 . The system of claim 7 , wherein at least one of the hypothesis transcriptions for at least one word is based on other words corresponding to the voice signal that are temporally proximate the at least one word.
12 . The system of claim 7 , wherein the operations further comprise:
obtaining a subsequent transcription of the voice signal;
determining when a portion of the subsequent transcription that corresponds to the one or more consistent words includes a corrected word that is different from any of the one or more consistent words; and
in response to determining the corrected word, provide an indication of the corrected word to the second device such that the second device changes the presentation of the one or more consistent words to include the corrected word, the presentation of the corrected word configured to occur before the final transcription of the voice signal is provided to the second device.
13 . The system of claim 12 , wherein the operations further comprise:
obtain the final transcription of the voice signal, the final transcription including a plurality of words that the automated speech recognition system outputs together as the final transcription of the voice signal;
determine when a portion of the final transcription that corresponds to the one or more consistent words includes a final word that is different from any of the one or more consistent words as corrected via the corrected word; and
in response to determining the final word, provide an indication of the final word to the second device such that the second device changes the presentation of the words to include the final word.
14 . A method to transcribe communications, the method comprising:
obtaining voice signal originating at a first device during a communication session between the first device and a second device;
providing the voice signal to an automated speech recognition system configured to transcribe the voice signal;
obtaining a plurality of hypothesis transcriptions generated by the automated speech recognition system, each of the plurality of hypothesis transcriptions including one or more words determined by the automated speech recognition system to be a transcription of a portion of the voice signal;
determining one or more consistent words that are included in two or more of the plurality of hypothesis transcriptions; and
in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before a final transcription of the voice signal is provided to the second device.
15 . The method of claim 14 , wherein the plurality of hypothesis transcriptions are not provided to the second device.
16 . The method of claim 14 , wherein the plurality of hypothesis transcriptions are obtained sequentially over time and a first portion of the voice signal associated with a first one of the plurality of hypothesis transcriptions includes all of the voice signal associated with all of the plurality of hypothesis transcriptions obtained previous to obtaining the first one of the plurality of hypothesis transcriptions.
17 . The method of claim 14 , further comprising:
determining a corrected word in a subsequent transcription of the voice signal that is different from any of the consistent words; and
in response to determining the corrected word, providing an indication of the corrected word to the second device, the second device using the corrected word to replace one or more of the consistent words in the presentation of the consistent words.
18 . The method of claim 17 wherein the automated speech recognition system generates each of the hypothesis transcriptions as well as the subsequent transcription.
19 . The method of claim 17 wherein the automated speech recognition system generates each of the plurality of hypothesis transcriptions and wherein the subsequent transcription is received via a call assistant interface device.
20 . The method of claim 14 wherein the automated speech recognition system includes a plurality of differently tuned ASR engines and wherein the two or more of the plurality of hypothesis transcriptions are generated via two or more of the differently tuned ASR engines, respectively.
21 . The method of claim 1 wherein the first and second portions of the voice signal are identical.