Transcription of audio
A method may include obtaining first features of first audio data that includes speech and obtaining second features of second audio data that is a revoicing of the first audio data. The method may further include providing the first features and the second features to an automatic speech recognition system and obtaining a single transcription generated by the automatic speech recognition system using the first features and the second features.
1. A method comprising:
obtaining first features of first audio data that includes speech;
obtaining second features of second audio data that is a revoicing of the first audio data;
providing the first features and the second features to an automatic speech recognition system; and
obtaining a single transcription generated by the automatic speech recognition system using the first features and the second features.
2. The method of claim 1 , wherein the first audio data is from a communication session between a first device and a second device.
3. The method of claim 2 , further comprising directing the transcription to the first device during the communication session.
4. The method of claim 1 , further comprising aligning the first features and the second features in time.
5. The method of claim 4 , wherein aligning the first features and the second features in time comprises providing the first features and the second features to a convolutional neural network.
6. The method of claim 5 , wherein the convolutional neural network includes a multiplier on each input path to each node of a convolutional layer of the convolutional neural network and the method further comprises adjusting a value of each multiplier based on a time difference between the first audio data and the second audio data.
7. The method of claim 4 , further comprising generating, using the automatic speech recognition system, phoneme probabilities for words in the first audio data using the aligned first features and the aligned second features.
8. The method of claim 1 , further comprising:
generating, using a first decoder of the automatic speech recognition system, a plurality of first words;
generating, using a second decoder of the automatic speech recognition system, a plurality of second words;
comparing the plurality of first words and the plurality of second words; and
generating the single transcription based on the comparison of the plurality of first words and the plurality of second words.
9. The method of claim 8 , wherein the plurality of first words is organized in a word graph, a word lattice, or a plurality of text strings.
10. A non-transitory computer-readable medium configured to store instructions that when executed by a computer system perform the method of claim 1 .
11. A system comprising:
at least one computer-readable media configured to store instructions;
at least one processor coupled to the one computer-readable media, the processor configured to execute the instructions to cause the system to perform operations, the operations comprising:
obtain first features of first audio data that includes speech;
obtain second features of second audio data that is a revoicing of the first audio data;
provide the first features and the second features to an automatic speech recognition system; and
obtain a single transcription generated by the automatic speech recognition system using the first features and the second features.
12. The system of claim 11 , wherein the first audio data is from a communication session between a first device and a second device.
13. The system of claim 12 , wherein the operations further comprise direct the transcription to the first device during the communication session.
14. The system of claim 11 , wherein the operations further comprise align the first features and the second features in time.
15. The system of claim 14 , wherein time aligning the first audio data and the second audio data includes time shifting the second audio data, the first audio data, or both the first audio data and the second audio data.
16. The system of claim 14 , wherein the operations further comprise generate, using the automatic speech recognition system, phoneme probabilities for words in the first audio data using the aligned first features and the aligned second features.
17. The system of claim 14 , wherein aligning the first features and the second features in time comprises providing the first features and the second features to a convolutional neural network.
18. The system of claim 17 , wherein the convolutional neural network includes a multiplier on each input path to each node of a convolutional layer of the convolutional neural network and the operations further comprise adjust a value of each multiplier based on a time difference between the first audio data and the second audio data.
19. The system of claim 11 , wherein the operations further comprise:
generate, using a first decoder of the automatic speech recognition system, a plurality of first words;
generate, using a second decoder of the automatic speech recognition system, a plurality of second words;
compare the plurality of first words and the plurality of second words; and
generate the single transcription based on the comparison of the plurality of first words and the plurality of second words.
20. The system of claim 19 , wherein the plurality of first words is organized in a word graph, a word lattice, or a plurality of text strings.