IP Library Granted Patent US 11,562,743
Granted Patent B2
US 11,562,743 · App. 16/775,295 · Granted Jan 24, 2023

Analysis of an automatically generated transcription

Inventor: Maayan Shir (Tel-Aviv, IL)
Assignee: salesforce.com, inc.
G10L15/26G10L15/065G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,743
App. No.
16/775,295
Granted
Jan 24, 2023
Kind
B2
Abstract

There is provided a computer implemented method of aligning an automatically generated transcription of an audio recording to a manually generated transcription of the audio recording comprising: identifying non-aligned text fragments, each located between respective two non-continuous aligned text-fragments of the automatically generated transcription, each aligned text-fragment matching words of the manually generated transcription, for each respective non-aligned text fragment: mapping a target keyword of the manually generated transcription to phonemes, mapping the respective non-aligned text fragment to a corresponding audio-fragment of the audio recording, mapping the audio-fragment to phonemes, identifying at least some of the phonemes of the audio-fragment that correspond to the phonemes of the target keyword, and mapping the identified at least some of the phonemes of the audio-fragment to a corresponding word of the automatically generated transcript, wherein the corresponding word is an incorrect automated transcription of the target word appearing in the manually generated transcription.

Claims (30)

1. A computer implemented method of aligning an automatically generated transcription of an audio recording to a manually generated transcription of the audio recording comprising:

identifying a plurality of non-aligned text fragments, each located between respective two non-continuous aligned text-fragments of the automatically generated transcription, each aligned text-fragment matching a plurality of words of the manually generated transcription;

for each respective non-aligned text fragment:

selecting a target keyword of the non-aligned text fragment of the manually generated transcription,

wherein the target keyword is not in a lexicon used for creation of the automatically generated transcription;

breaking down the target keyword of the manually generated transcription to a plurality of phonemes;

mapping the respective non-aligned text fragment that includes the target keyword and at least one non-target keyword to a corresponding audio-fragment of the audio recording;

breaking down the audio-fragment to a plurality of phonemes;

identifying at least some of the plurality of phonemes of the audio-fragment that map to the plurality of phonemes of the target keyword; and

mapping the identified at least some of the plurality of phonemes of the audio-fragment to a corresponding word of the automatically generated transcript, wherein the corresponding word is an incorrect automated transcription of the target word appearing in the manually generated transcription.

2. The method of claim 1 , wherein the at least some of the plurality of phonemes of the audio-fragment are identified as corresponding to the plurality of phonemes of the target keyword according to a closest matched computed based on shortest phoneme weighted distance.

3. The method of claim 1 , wherein words of the automatically generated transcription are associated with a timestamp indicating a mapping to the audio recording, and wherein the respective non-aligned text fragment is mapped to the corresponding audio-fragment according to the timestamp.

4. The method of claim 1 , wherein the matching comprises selecting the at least some of the plurality of phonemes of the audio-fragment, according a lowest value of a phoneme distance to the plurality of phonemes of the keyword of the manually generated transcript.

5. The method of claim 4 , wherein the phoneme distance is selected from the group consisting of: a binary phoneme distance that assigns a binary value indicative of whether each respective phenome is matched or is not matched, and a weighted phoneme distance that assigns a non-binary value indicative of an amount of similarity between corresponding phonemes.

6. The method of claim 1 , further comprising feeding the target keyword of the manually generated transcription and the corresponding word of the automatically generated transcription for automatically updating a model that computes the automatically generated transcript.

7. The method of claim 1 , wherein the target keyword of the manually generated transcription and the corresponding word of the automatically generated transcription are used for adjusting the model for correctly automatically transcribing an audio-fragment corresponding to the audio-fragment of the audio recording to the target keyword of the manually generated transcription instead of to the corresponding word of the automatically generated transcript.

8. The method of claim 1 , further comprising computing a value for precision and/or recall of transcription of the target keyword of the manually generated transcription in the automatically generated transcript.

9. The method of claim 1 , wherein the automatically generated transcription is created by an acoustic model that extracts phonemes from the audio recording and assigns a probability value to each phoneme denoting likelihood of accurate extraction, and a language model that receives the extracted phonemes and outputs the automatically generated transcription by mapping phonemes to words and determines a word sequence probability.

10. The method of claim 1 , wherein each of the plurality of aligned text-fragments includes a sequence of at least 4 matching words.

11. A computer implemented method of post-processing an automatically generated transcription of an audio recording to correct transcription errors, comprising:

receiving an audio recording;

computing the automatically generated transcription of the audio recording by an acoustic model that extracts phonemes from the audio recording, and a language model that receives the extracted phonemes and outputs the automatically generated transcription by mapping phonemes to words selected from a lexicon;

receiving a plurality of target words, wherein the plurality of target words are excluded from the lexicon;

computing a respective weighted phoneme distance that assigns a non-binary value indicative of an amount of similarity between corresponding phonemes, from an automatically transcribed word of the automatically generated transcription to each of the plurality of target words, and

when the respective phoneme distance is according to a requirement, switching the respective automatically transcribed word to a certain target word of the plurality of target words corresponding to a lowest value of the respective phoneme distance.

12. The method of claim 11 , wherein the requirement denotes that the automatically transcribed word is similar to but not identical to the plurality of target words.

13. The method of claim 12 , wherein the requirement is a range having an upper threshold value of the phoneme distance denoting identical words and a lower threshold value of the phoneme distance denoting similar but difference words.

14. The method of claim 11 , wherein the respective automatically transcribed word and an indication of a switch to the certain target word are used to update the language model for improved accuracy in mapping phonemes to the certain target word.

15. The method of claim 11 , wherein the automatically transcribed word is selected for inclusion in the automatically generated transcription when the automatically transcribed word is assigned a confidence value by the language model above a threshold.

16. The method of claim 11 , further comprising confirming the switching when a phoneme distance computed between phonemes extracted from a portion of the audio recording corresponding to the automatically transcribed word and the certain target word denotes statistical equivalence.

Assignments (2)
CHANGE OF NAME Recorded Sep 20, 2023
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 064975/0956 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2020
From: SHIR, MAAYAN
To: SALESFORCE.COM, INC.
Reel/Frame 052532/0492 →
Continuity (1)
Related Publication 20210233535A1 · Jul 29, 2021