IP Library Granted Patent US 11,580,959
Granted Patent B2
US 11,580,959 · App. 17/034,082 · Granted Feb 14, 2023

Improving speech recognition transcriptions

Inventors: Andrew R. Freed (Cary, NC); Marco Noel (Quebec, CA); Aishwarya Hariharan (Morrisville, NC); Martha Holloman (Raleigh, NC); Mohammad Gorji-Sefidmazgi (Matthews, NC); Daniel Zyska (Charlotte, NC)
Assignee: International Business Machines Corporation
G10L15/10G10L15/065G10L15/187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,959
App. No.
17/034,082
Granted
Feb 14, 2023
Kind
B2
Abstract

An approach to correcting transcriptions of speech recognition models may be provided. A list of similar sounding phonemes from associated with the phonemes of high frequency terms may be generated for a particular node associated with a virtual assistant. An utterance may be transcribed and receive a confidence score regarding the correctness of the transcription based on audio metrics and other factors. The phonemes of the utterance can be compared to the phonemes of the high frequency terms from the list and a score for the matching phonemes and similar sounding phonemes can be determined. If it is determined the sounds similar score for a term from the high frequency term list is above a threshold, the transcription can be replaced with the term, providing a corrected transcription.

Claims (44)

1. A computer-implemented method for training a model for improving speech recognition, the computer-implemented method comprising:

receiving, by one or more processors, a history of utterances and corresponding audio metrics for the utterances;

identifying, by the one or more processors, one or more high frequency terms based on the history of utterances;

converting the identified one or more high frequency terms into one or more phonemes;

generating, by the one or more processors, a sounds similar list for the identified one or more high frequency terms, based at least in part, on the one or more phonemes;

transcribing, by the one or more processors, an utterance from a virtual assistant into a transcription including one or more words, wherein transcribing comprises instructions to transform the utterance from a virtual assistant into an audio spectrogram and identifying one or more phonemes of the utterance from the virtual assistant based, at least in part, on the audio spectrogram;

calculating, by the one or more processors, a transcription score for each of the more or more words included in the transcription; and

responsive to the transcription score for a word from the transcription being below a threshold, comparing, by the one or more processors, one or more phonemes of the word having the transcription score below the threshold to the one or more phonemes of the high frequency terms on the sounds similar list to determine a sounds similar score and replacing the word included in the transcription having the transcription score below the threshold with a high frequency term on the sounds similar list if the sounds similar score is above a threshold.

2. The computer-implemented method of claim 1 , wherein the audio metrics identify the frequency of an utterance and the frequency of one or more terms corresponding to the utterance.

3. The computer-implemented method of claim 1 ,

wherein transcribing the utterance from the virtual assistant into the transcription including the one or more words is performed by a speech recognition model based on a deep neural network.

4. The computer-implemented method of claim 1 , further comprising:

assigning, by the one or more processors, a sounds similar value to the term phonemes which corresponds to the utterance phonemes.

5. The computer-implemented method of claim 1 , wherein the history of utterances is from a virtual assistant.

6. A computer system for improving speech recognition transcriptions, the system comprising:

one or more computer processors;

one or more computer readable storage media device;

computer program instructions stored on the computer readable storage device, comprising instructions to:

receive a history of utterances and corresponding audio metrics for the utterances;

identify one or more high frequency terms based on the history of utterances;

convert the identified one or more high frequency terms into one or more phonemes;

generate a sounds similar list for the identified one or more high frequency terms, based at least in part, on the one or more phonemes;

transcribe an utterance from a virtual assistant into a transcription including one or more words, wherein transcribe comprises instructions to transform the utterance from a virtual assistant into an audio spectrogram and identify one or more phonemes of the utterance from the virtual assistant based, at least in part, on the audio spectrogram;

calculate a transcription score for each of the one or more words included in the transcription; and

responsive to the transcription score for a word from the transcription being below a threshold, compare one or more phonemes of the word having the transcription score below the threshold to the one or more phonemes of the high frequency terms on the sounds similar list to determine a sounds similar score and replacing the word included in the transcription having the transcription score below the threshold with a high frequency term on the sounds similar list if the sounds similar score is above a threshold.

7. The computer system of claim 6 , wherein the audio metrics identify the frequency of an utterance and the frequency of one or more terms corresponding to the utterance.

8. The computer system of claim 7 ,

wherein transcribing the utterance from the virtual assistant into the transcription including the one or more words is performed by a speech recognition model based on a deep neural network.

9. The computer system of claim 6 , further comprising instructions to:

assign a sounds similar value to the term phonemes which corresponds to the utterance phonemes.

10. The computer system of claim 6 , wherein the history of utterances is from a virtual assistant.

11. A computer program product for improving speech recognition transcriptions, the computer program product comprising a computer readable storage media device and program instructions sorted on the computer readable storage media device, the program instructions including instructions to:

receive a history of utterances and corresponding audio metrics for the utterances;

identify one or more high frequency terms based on the history of utterances;

convert the identified one or more high frequency terms into one or more phonemes;

generate a sounds similar list for the identified one or more high frequency terms, based at least in part, on the one or more phonemes;

transcribe an utterance from a virtual assistant into a transcription including one or more words, wherein transcribe comprises instructions to transform the utterance from a virtual assistant into an audio spectrogram and identify one or more phonemes of the utterance from the virtual assistant based, at least in part, on the audio spectrogram;

calculate a transcription score for each of the one or more words included in the transcription; and

responsive to the transcription score for a word from the transcription being below a threshold, compare one or more phonemes of the word having the transcription score below the threshold to the one or more phonemes of the high frequency terms on the sounds similar list to determine a sounds similar score and replacing the word included in the transcription having the transcription score below the threshold with a high frequency term on the sounds similar list if the sounds similar score is above a threshold.

12. The computer program product of claim 11 , wherein the audio metrics identify the frequency of an utterance and the frequency of one or more terms corresponding to the utterance.

13. The computer program product of claim 12 , wherein transcribing the utterance from the virtual assistant into the transcription including the one or more words is performed by a speech recognition model based on a deep neural network.

14. The computer program product of claim 11 , further comprising instructions to:

assign a sounds similar value to the term phonemes which corresponds to the utterance phonemes.

15. The computer program product of claim 11 , wherein the history of utterances is from a virtual assistant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2020
From: FREED, ANDREW R.; NOEL, MARCO; HARIHARAN, AISHWARYA; HOLLOMAN, MARTHA; GORJI-SEFIDMAZGI, MOHAMMAD; ZYSKA, DANIEL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 053898/0036 →
Continuity (1)
Related Publication 20220101830A1 · Mar 31, 2022
Cited By (5)
US 12,223,948 US 12,387,720 US 12,394,411 US 12,562,151 US 12,592,222