IP Library Patent Application 16595264
Patent Application
App. No. 16/595,264

User interface to assist in hybrid transcription of audio that includes a repeated phrase

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/595,264
Abstract

When transcribing an audio recording, certain phrases may be difficult to resolve, especially if they involve names and/or infrequently used terms. However, often such phrases may be repeated multiple times throughout the audio recording. Embodiments described herein interact with a transcriber to resolve such cases of repeated phrases. In one embodiment, a computer plays segments of an audio recording to the transcriber, and at least some of the segments include an utterance of a phrase. The computer also presents, to the transcriber, transcriptions of the segments, and at least some of the transcriptions do not include a correct transcription of the phrase. The computer receives from the transcriber an indication of which of the segments include an utterance of the phrase and the correct transcription of the phrase, and then updates a transcription of the audio recording accordingly.

Claims (38)

1 . A system configured to interact with a transcriber to resolve a repeated phrase, comprising:

a computer configured to:

play segments of an audio recording to the transcriber; wherein the audio recording comprises first and second channels, recorded by first and second microphones, respectively; wherein the first and second channels comprise recordings of speech of first and second people, respectively; and wherein the segments belong to a cluster comprising similar utterances, and at least some of the segments comprise an utterance of a phrase;

present, to the transcriber, first and second transcriptions of respective first and second segments of the audio recording; wherein the first and second segments comprise utterances of the first and second people, respectively;

receive from the transcriber an indication indicating that at least one of the first and second segments comprises an utterance of the phrase, and a correct transcription of the phrase; and

update a transcription of the audio recording based on the indication and the correct transcription.

2 . The system of claim 1 , wherein the first and second transcriptions are generated by an automatic speech recognition (ASR) system.

3 . The system of claim 2 , wherein the correct transcription of the phrase includes a term that is not represented in a language model utilized by the ASR system to generate the first and second transcriptions, and wherein the computer is further configured to update said language model to include a representation of the term.

4 . The system of claim 1 , wherein the first transcription and/or the second transcription comprises the correct transcription of the phrase, and the correct transcription is selected by the transcriber.

5 . The system of claim 1 , wherein the computer is further configured to present for each segment, from among the segments, a value indicative of at least one of: a similarity of the segment to a consensus of the segments, a similarity of the segment to other segments from among the segments, and a similarity of a transcription of the segment to transcriptions of the other segments.

6 . (canceled)

7 . The system of claim 1 , wherein the computer is further configured to present the first and second transcriptions in an order that is based on confidence in the first and second transcriptions.

8 . (canceled)

9 . The system of claim 1 , wherein the computer is further configured to update the transcription of the audio recording based on an indication indicating that a number of the segments that comprise an utterance of the phrase is greater than a threshold that is at least two.

10 . A method for interacting with a transcriber to resolve a repeated phrase, comprising:

playing segments of an audio recording to the transcriber; wherein the audio recording comprises first and second channels, recorded by first and second microphones, respectively; wherein the first and second channels comprise recordings of speech of first and second people, respectively; and wherein the segments belong to a cluster comprising similar utterances, and at least some of the segments comprise an utterance of a phrase;

presenting, to the transcriber, first and second transcriptions of respective first and second segments of the audio recording; wherein the first and second segments comprise utterances of the first and second people, respectively;

receiving from the transcriber an indication indicating that at least one of the first and second segments comprises an utterance of the phrase, and a correct transcription of the phrase; and

updating a transcription of the audio recording based on the indication and the correct transcription.

11 . The method of claim 10 , further comprising generating the transcription of the audio recording utilizing an automatic speech recognition (ASR) system.

12 . The method of claim 10 , further comprising generating the first and second transcriptions utilizing an automatic speech recognition (ASR) system.

13 . The method of claim 12 , wherein the correct transcription of the phrase includes a term that is not represented in a language model utilized by the ASR system to generate the first and second transcriptions, and further comprising updating the language model to include a representation of the term.

14 . The method of claim 10 , wherein the first transcription and/or the second transcription comprises the correct transcription of the phrase, and further comprising receiving a selection by the transcriber of the correct transcription of the phrase, from among the first transcription and the second transcription.

15 . The method of claim 10 , further comprising presenting for each segment, from among the segments, a value indicative of at least one of: a similarity of the segment to a consensus of the segments, a similarity of the segment to other segments from among the segments, and a similarity of a transcription of the segment to transcriptions of the other segments.

16 . (canceled)

17 . The method of claim 10 , further comprising presenting the first and second transcriptions in an order that is based on confidence in the first and second transcriptions.

18 . The method of claim 10 , further comprising updating the transcription of the audio recording based on an indication indicating that a number of the segments that comprise an utterance of the phrase is greater than a threshold that is at least two.

19 . A non-transitory computer-readable medium having instructions stored thereon that, in response to execution by a system including a processor and memory, causes the system to perform operations comprising:

playing segments of an audio recording to a transcriber; wherein the audio recording comprises first and second channels, recorded by first and second microphones, respectively; wherein the first and second channels comprise recordings of speech of first and second people, respectively; and wherein the segments belong to a cluster comprising similar utterances;

presenting, to the transcriber, first and second transcriptions of respective first and second segments of the audio recording: wherein the first and second segments comprise utterances of the first and second people, respectively;

receiving from the transcriber an indication indicating that at least one of the first and second segments comprises an utterance of a phrase, and a correct transcription of the phrase; and

updating a transcription of the audio recording based on the indication and the correct transcription.

20 . The non-transitory computer-readable medium of claim 19 , wherein the first transcription and/or the second transcription comprises the correct transcription of the phrase, and further comprising instructions defining a step of receiving a selection by the transcriber of the correct transcription of the phrase, from among the first transcription and the second transcription.

21 . (canceled)

22 . (canceled)

23 . The non-transitory computer-readable medium of claim 19 , further comprising instructions defining a step of presenting for each segment, from among the segments, a value indicative of at least one of: a similarity of the segment to a consensus of the segments, a similarity of the segment to other segments from among the segments, and a similarity of a transcription of the segment to transcriptions of the other segments.

24 . The non-transitory computer-readable medium of claim 19 , further comprising instructions defining a step of presenting the first and second transcriptions in an order that is based on confidence in the first and second transcriptions.

25 . The non-transitory computer-readable medium of claim 19 , further comprising instructions defining a step of updating the transcription of the audio recording based on the an indication indicating that a number of the segments that comprise an utterance of the phrase is greater than a threshold that is at least two.

Assignments (2)
SECURITY INTEREST Recorded Jun 4, 2020
From: VERBIT SOFTWARE LTD.
To: SILICON VALLEY BANK
Reel/Frame 052836/0856 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2020
From: SHELLEF, ERIC ARIEL; BEN TSVI, YAAKOV KOBI; GETZ, IRIS; LIVNE, TOM; HIMMELREICH, ROMAN; ASOR, ELI; ROSENSWEIG, ELISHA YEHUDA; SHTILERMAN, ELAD
To: VERBIT SOFTWARE LTD.
Reel/Frame 052228/0567 →