IP Library Granted Patent US 11,158,322
Granted Patent B2
US 11,158,322 · App. 16/595,211 · Granted Oct 26, 2021

Human resolution of repeated phrases in a hybrid transcription system

Inventors: Eric Ariel Shellef (Ramat Gan, IL); Yaakov Kobi Ben Tsvi (Ramat Hasharon, IL); Iris Getz (Ramat Hasharon, IL); Tom Livne (Herzliya, IL); Eli Asor (Tel Aviv, IL); Elisha Yehuda Rosensweig (Ra'anana, IL)
Assignee: Verbit Software Ltd.
G10L15/26G06F40/20G10L15/01G10L15/02G10L15/04G10L15/063G10L15/08G10L15/183G10L15/187G10L15/1815G10L15/19G10L15/20G10L15/22G10L15/30G10L25/60H04R1/406H04R3/005H04R5/027G06F3/0484G10L2015/0631G10L2015/0635G10L2015/0638G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,158,322
App. No.
16/595,211
Granted
Oct 26, 2021
Kind
B2
Abstract

When transcribing audio recordings, such as legal depositions, phrases may be repeated throughout the recordings, but these repeated phrases get transcribed incorrectly by an automatic speech recognition (ASR) system. In order to assist a transcriber to correctly resolve such phrases, some embodiments described herein involve a computer that receives an audio recording that includes speech, generates a transcription of the audio recording utilizing an ASR system, and clusters segments of the audio recording into clusters of similar utterances. The computer provides a transcriber with certain segments of the audio recording, which include similar utterances belonging to a certain cluster, along with transcriptions of the certain segments. The computer receives from the transcriber: an indication of which of the certain segments include repetitions of a phrase, and a correct transcription of the phrase. The computer then updates the transcription of the audio recording based on the indication and the correct transcription.

Claims (48)

1. A system configured to assist in transcription of a repeated phrase, comprising:

a frontend server configured to transmit an audio recording comprising speech of first and second people; and

a backend server configured to:

generate a transcription of the audio recording utilizing an automatic speech recognition (ASR) system;

cluster segments of the audio recording into clusters of similar utterances;

select first and second segments of the audio recording that comprise similar utterances spoken by the first and second people, respectively;

generate, for each transcriber from among transcribers, feature values based on the first and second segments and utilize a machine learning-based model to calculate, based on the feature values, expected accuracies of transcriptions of the first and second segments were they transcribed by the transcriber; wherein the machine learning-based model is generated based on training data comprising additional feature values generated based on additional segments of additional audio recordings, and values of accuracies of transcriptions, by the transcriber, of the additional segments;

select a certain transcriber, from among the transcribers, whose expected accuracies reach a predetermined threshold;

provide the certain transcriber with the first and second segments of the audio recording and with transcriptions of the first and second segments;

receive from the certain transcriber: an indication indicating whether the first and second segments comprise repetitions of a phrase, and a correct transcription of said phrase; and

update the transcription of the audio recording based on the indication and the correct transcription.

2. The system of claim 1 , wherein the backend server is further configured to utilize the indication to update a phonetic model utilized by the ASR system to reflect one or more pronunciations of the phrase.

3. The system of claim 1 , wherein the backend server is further configured to update a language model utilized by the ASR system to include the correct transcription of the phrase.

4. The system of claim 1 , wherein the backend server is further configured to cluster the segments utilizing dynamic time warping (DTW) of acoustic feature representations of the segments.

5. The system of claim 1 , wherein the backend server is further configured cluster the segments based on similarity of paths corresponding to the segments in a lattice constructed by the ASR system.

6. The system of claim 1 , wherein the backend server is further configured to represent each segment of audio and a product of ASR of the segment using a vector of feature values that comprises: one or more feature values indicative of acoustic properties of the segment, and at least some feature values indicative of phonetic transcription properties calculated by the ASR system; the backend server is further configured to utilize a distance function that operates on pairs of vectors of feature values.

7. The system of claim 1 , wherein the backend server is further configured to update the transcription of the audio recording responsive to the indication indicating that a number of the segments that comprise an utterance of the phrase is greater than a threshold that is at least two.

8. The system of claim 1 , wherein the expected accuracies of the transcriptions of the first and second segments were they transcribed by the certain transcriber are indicative of expected word error rates (WER) in said transcriptions of the first and second segments.

9. A method for assisting in transcription of a repeated phrase, comprising:

receiving an audio recording comprising speech of first and second people;

generating a transcription of the audio recording utilizing an automatic speech recognition (ASR) system;

clustering segments of the audio recording into clusters of similar utterances;

selecting first and second segments of the audio recording that comprise similar utterances spoken by the first and second people, respectively;

generating, for each transcriber from among transcribers, feature values based on the first and second segments and utilizing a machine learning-based model to calculate, based on the feature values, expected accuracies of transcriptions of the first and second segments were they transcribed by the transcriber; wherein the machine learning-based model is generated based on training data comprising additional feature values generated based on additional segments of additional audio recordings, and values of accuracies of transcriptions, by the transcriber, of the additional segments;

selecting a certain transcriber, from among the transcribers, whose expected accuracies reach a predetermined threshold;

providing the certain transcriber with the first and second segments of the audio recording and with transcriptions of the first and second segments;

receiving from the certain transcriber: an indication indicating whether the first and second segments comprise repetitions of a phrase, and a correct transcription of said phrase; and

updating the transcription of the audio recording based on the indication and the correct transcription.

10. The method of claim 9 , further comprising utilizing the indication to update a phonetic model utilized by the ASR system to reflect one or more pronunciations of the phrase.

11. The method of claim 10 , wherein the expected accuracies of the transcriptions of the first and second segments were they transcribed by the certain transcriber are indicative of expected word error rates (WER) in said transcriptions of the first and second segments.

12. The method of claim 9 , further comprising updating a language model utilized by the ASR system to include the correct transcription of the phrase.

13. The method of claim 9 , further comprising clustering the segments utilizing dynamic time warping (DTW) of acoustic feature representations of the segments.

14. The method of claim 9 , further comprising clustering the segments based on similarity of paths corresponding to the segments in a lattice constructed by the ASR system.

15. The method of claim 9 , further comprising updating the transcription of the audio recording based on the indication indicating that a number of the segments that comprise an utterance of the phrase is greater than a threshold that is at least two.

16. A non-transitory computer-readable medium having instructions stored thereon that, in response to execution by a system including a processor and memory, causes the system to perform operations comprising:

receiving an audio recording comprising speech of first and second people;

generating a transcription of the audio recording utilizing an automatic speech recognition (ASR) system;

clustering segments of the audio recording into clusters of similar utterances;

selecting first and second segments of the audio recording that comprise similar utterances spoken by the first and second people, respectively;

generating, for each transcriber from among transcribers, feature values based on the first and second segments and utilizing a machine learning-based model to calculate, based on the feature values, expected accuracies of transcriptions of the first and second segments were they transcribed by the transcriber; wherein the machine learning-based model is generated based on training data comprising additional feature values generated based on additional segments of additional audio recordings, and values of accuracies of transcriptions, by the transcriber, of the additional segments;

selecting a certain transcriber, from among the transcribers, whose expected accuracies reach a predetermined threshold;

providing the certain transcriber with the first and second segments of the audio recording and with transcriptions of the first and second segments;

receiving from the certain transcriber: an indication indicating whether the first and second segments comprise repetitions of a phrase, and a correct transcription of said phrase; and

updating the transcription of the audio recording based on the indication and the correct transcription.

17. The non-transitory computer-readable medium of claim 16 , further comprising instructions defining a step of utilizing the indication to update a phonetic model utilized by the ASR system to reflect one or more pronunciations of the phrase.

18. The non-transitory computer-readable medium of claim 16 , further comprising instructions defining a step of updating a language model utilized by the ASR system to include the correct transcription of the phrase.

19. The non-transitory computer-readable medium of claim 16 , wherein the expected accuracies of the transcriptions of the first and second segments were they transcribed by the certain transcriber are indicative of expected word error rates (WER) in said transcriptions of the first and second segments.

20. The non-transitory computer-readable medium of claim 16 , further comprising instructions defining a step of clustering the segments utilizing dynamic time warping (DTW) of acoustic feature representations of the segments.

Assignments (3)
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Feb 3, 2021
From: VERBIT SOFTWARE LTD
To: SILICON VALLEY BANK
Reel/Frame 055209/0474 →
SECURITY INTEREST Recorded Jun 4, 2020
From: VERBIT SOFTWARE LTD.
To: SILICON VALLEY BANK
Reel/Frame 052836/0856 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2020
From: SHELLEF, ERIC ARIEL; BEN TSVI, YAAKOV KOBI; GETZ, IRIS; LIVNE, TOM; HIMMELREICH, ROMAN; ASOR, ELI; ROSENSWEIG, ELISHA YEHUDA; SHTILERMAN, ELAD
To: VERBIT SOFTWARE LTD.
Reel/Frame 052228/0567 →
Continuity (2)
Provisional Application 62896617 · Sep 6, 2019
Related Publication 20210074272A1 · Mar 11, 2021