IP Library › Granted Patent US 12,322,375
Granted Patent B1
US 12,322,375 · App. 18/496,797 · Granted Jun 3, 2025

Transcription analysis platform

Inventors: Michael J. Szentes (San Antonio, TX); Carlos Chavez (San Antonio, TX); Robert E. Lewis (Boerne, TX); Nicholas S. Walker (El Paso, TX)
Assignee: United Services Automobile Association (USAA)
G10L15/01G06F16/685G06F40/226G06F40/268G10L15/1822G10L15/26G10L17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,375
App. No.
18/496,797
Granted
Jun 3, 2025
Kind
B1
Abstract

Various embodiments of the present disclosure evaluate transcription accuracy. In some implementations, the system normalizes a first transcription of an audio file and a baseline transcription of the audio file. The baseline transcription can be used as an accurate transcription of the audio file. The system can further determine an error rate of the first transcription by aligning each portion of the first transcription with the portion of the baseline transcription, and assigning a label to each portion based on a comparison of the portion of the first transcription with the portion of the baseline transcription.

Claims (43)

1. A method comprising:

receiving, at a graphical user interface (GUI), a transcription of an audio file and a baseline transcription of the audio file;

displaying, by the GUI, one or more portions of the transcription aligned with one or more corresponding portions of the baseline transcription;

displaying, by the GUI, a differentiator label for each portion of the transcription based on a comparison of a portion of the transcription with a corresponding portion of the baseline transcription;

displaying, by the GUI, an error rate of the transcription based on determinations of whether each portion, as aligned between the transcription and the baseline transcription, has a differentiator label corresponding to correct, inserted, substituted in, or deleted respective differentiator labels assigned to each portion of the transcription; and

displaying, by the GUI, a heat map that is based on A) the error rate and B) determined word lengths of the transcription and/or the baseline transcription.

2. The method of claim 1 , further comprising:

displaying an inserted section into the transcription at a location where a deleted word was deleted from the transcription.

3. The method of claim 1 , further comprising:

displaying a normalization of the transcription and the baseline transcription by automatically changing words and/or numbers to a standardized spelling or appearance.

4. The method of claim 1 , wherein an x-axis on the heat map represents the error rate and a y-axis represents a length difference of the transcription compared to the baseline transcription.

5. The method of claim 1 , wherein an x-axis on the heat map represents a word accuracy rate and a y-axis represents a length of a call according to the baseline transcription.

6. The method of claim 1 , wherein the error rate is determined by dividing a number of correct portions by a sum of at least a number of portions with a label corresponding to inserted, a number of portions with a label corresponding to deleted, and a number of portions with a label corresponding to substituted.

7. The method of claim 1 , wherein the error rate is an error rate of incorrect phrases between the transcription and the baseline transcription.

8. A system comprising:

one or more processors; and

one or more memories storing instructions that, when executed by the one or more processors, cause the system to perform a process comprising:

receiving, at a graphical user interface (GUI), a transcription of an audio file and a baseline transcription of the audio file;

displaying, by the GUI, one or more portions of the transcription aligned with one or more corresponding portions of the baseline transcription;

displaying, by the GUI, a differentiator label for each portion of the transcription based on a comparison of a portion of the transcription with a corresponding portion of the baseline transcription;

displaying, by the GUI, an error rate of the transcription based on determinations of whether each portion, as aligned between the transcription and the baseline transcription, has a differentiator label corresponding to correct, inserted, substituted in, or deleted respective differentiator labels assigned to each portion of the transcription; and

displaying, by the GUI, a heat map that is based on A) the error rate and B) determined word lengths of the transcription and/or the baseline transcription.

9. The system according to claim 8 , wherein the process further comprises:

displaying an inserted section into the transcription at a location where a deleted word was deleted from the transcription.

10. The system according to claim 8 , wherein the process further comprises:

displaying a normalization of the transcription and the baseline transcription by automatically changing words and/or numbers to a standardized spelling or appearance.

11. The system according to claim 8 , wherein an x-axis on the heat map represents the error rate and a y-axis represents a length difference of the transcription compared to the baseline transcription.

12. The system according to claim 8 , wherein an x-axis on the heat map represents a word accuracy rate and a y-axis represents a length of a call according to the baseline transcription.

13. The system according to claim 8 , wherein the error rate is determined by dividing a number of correct portions by a sum of at least a number of portions with a label corresponding to inserted, a number of portions with a label corresponding to deleted, and a number of portions with a label corresponding to substituted.

14. The system according to claim 8 , wherein the error rate is an error rate of incorrect phrases between the transcription and the baseline transcription.

15. A non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receiving, at a graphical user interface (GUI), a transcription of an audio file and a baseline transcription of the audio file;

displaying, by the GUI, one or more portions of the transcription aligned with one or more corresponding portions of the baseline transcription;

displaying, by the GUI, a differentiator label for each portion of the transcription based on a comparison of a portion of the transcription with a corresponding portion of the baseline transcription;

displaying, by the GUI, an error rate of the transcription based on determinations of whether each portion, as aligned between the transcription and the baseline transcription, has a differentiator label corresponding to correct, inserted, substituted in, or deleted respective differentiator labels assigned to each portion of the transcription; and

displaying, by the GUI, a heat map that is based on A) the error rate and B) determined word lengths of the transcription and/or the baseline transcription.

16. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

displaying an inserted section into the transcription at a location where a deleted word was deleted from the transcription.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

displaying a normalization of the transcription and the baseline transcription by automatically changing words and/or numbers to a standardized spelling or appearance.

18. The non-transitory computer-readable medium of claim 15 , wherein an x-axis on the heat map represents the error rate and a y-axis represents a length difference of the transcription compared to the baseline transcription.

19. The non-transitory computer-readable medium of claim 15 , wherein an x-axis on the heat map represents a word accuracy rate and a y-axis represents a length of a call according to the baseline transcription.

20. The non-transitory computer-readable medium of claim 15 , wherein the error rate is determined by dividing a number of correct portions by a sum of at least a number of portions with a label corresponding to inserted, a number of portions with a label corresponding to deleted, and a number of portions with a label corresponding to substituted.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: SZENTES, MICHAEL J.; CHAVEZ, CARLOS; LEWIS, ROBERT E.; WALKER, NICHOLAS S.
To: UIPCO, LLC
Reel/Frame 065483/0429 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: UIPCO, LLC
To: UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)
Reel/Frame 065483/0916 →
Continuity (3)
Continuation 17084513 · Oct 29, 2020
Continuation 15614119 · Jun 5, 2017
Provisional Application 62349290 · Jun 13, 2016
References Cited (67)
US 5649060A · Ellozy et al. · 1997 [cited by applicant]
US 6064957A · Brandow et al. · 2000 [cited by applicant]
US 6370504B1 · Zick et al. · 2002 [cited by applicant]
US 6473778B1 · Gibbon · 2002 [cited by applicant]
US 6535849B1 · Pakhomov et al. · 2003 [cited by applicant]
US 8571851B1 · Tickner et al. · 2013 [cited by applicant]
US 8892447B1 · Srinivasan et al. · 2014 [cited by applicant]
US 9224386B1 · Weber · 2015 [cited by applicant]
US 9478218B2 · Shu · 2016 [cited by applicant]
US 9508338B1 · Kaszczuk et al. · 2016 [cited by applicant]
US 9679564B2 · Daborn et al. · 2017 [cited by applicant]
US 10854190B1 · Szentes et al. · 2020 [cited by applicant]
US 20020065653A1 · Kriechbaum et al. · 2002 [cited by applicant]
US 20030004724A1 · Kahn et al. · 2003 [cited by applicant]
US 20040015350A1 · Gandhi et al. · 2004 [cited by applicant]
US 20040199385A1 · Deligne et al. · 2004 [cited by applicant]
US 20050010407A1 · Jaroker · 2005 [cited by applicant]
US 20050137867A1 · Miller · 2005 [cited by examiner]
US 20050177369A1 · Stoimenov et al. · 2005 [cited by applicant]
US 20050228667A1 · Duan et al. · 2005 [cited by applicant]
US 20060136205A1 · Song · 2006 [cited by applicant]
US 20060190249A1 · Kahn · 2006 [cited by examiner]
US 20060190263A1 · Finke et al. · 2006 [cited by applicant]
US 20070192095A1 · Braho et al. · 2007 [cited by applicant]
US 20080077583A1 · Castro et al. · 2008 [cited by applicant]
US 20090141875A1 · Demmitt · 2009 [cited by examiner]
US 20090271192A1 · Marquette et al. · 2009 [cited by applicant]
US 20100205628A1 · Davis et al. · 2010 [cited by applicant]
US 20100228548A1 · Liu et al. · 2010 [cited by applicant]
US 20110208522A1 · Pereg · 2011 [cited by examiner]
US 20110239119A1 · Phillips et al. · 2011 [cited by applicant]
US 20120016671A1 · Jaggi et al. · 2012 [cited by applicant]
US 20120179694A1 · Sciacca et al. · 2012 [cited by applicant]
US 20120278337A1 · Acharya · 2012 [cited by applicant]
US 20130013305A1 · Thompson et al. · 2013 [cited by applicant]
US 20130080150A1 · Levit et al. · 2013 [cited by applicant]
US 20130124984A1 · Kuspa · 2013 [cited by applicant]
US 20130311181A1 · Bachtiger et al. · 2013 [cited by applicant]
US 20140088962A1 · Corfield · 2014 [cited by examiner]
US 20140153709A1 · Byrd et al. · 2014 [cited by applicant]
US 20140163981A1 · Cook · 2014 [cited by examiner]
US 20150039306A1 · Sidi et al. · 2015 [cited by applicant]
US 20150058006A1 · Proux · 2015 [cited by applicant]
US 20150269136A1 · Alphonso et al. · 2015 [cited by applicant]
US 20160078861A1 · Mathias et al. · 2016 [cited by applicant]
US 20160091967A1 · Prokofieva et al. · 2016 [cited by applicant]
US 20160133251A1 · Kadirkamanathan et al. · 2016 [cited by applicant]
US 20160246929A1 · Zenati et al. · 2016 [cited by applicant]
US 20160342706A1 · Galle · 2016 [cited by applicant]
US 20170323643A1 · Arslan et al. · 2017 [cited by applicant]
US 20180061404A1 · Devaraj et al. · 2018 [cited by applicant]
US 20180122367A1 · Ingmarsson · 2018 [cited by applicant]
US 20180294014A1 · Ekambaram et al. · 2018 [cited by applicant]
US 20220300306A1 · Leung et al. · 2022 [cited by applicant]
CA 2977076A1 · 2011 [cited by examiner]
FR 3032574A1 · 2016 [cited by examiner]
Suhm, Bernhard, Brad Myers, and Alex Waibel. “Multimodal error correction for speech user interfaces.” ACM transactions on computer-human interaction (TOCHI) 8.1 (2001): 60-98. (Year: 2001). [cited by examiner]
Kowal, Sabine, and Daniel C. O'Connell. “Transcription as a crucial step of data analysis.” The SAGE handbook of qualitative data analysis 7.5 (2014): 64-79. (Year: 2014). [cited by examiner]
Panayotov, Vassil, et al. “Librispeech: an asr corpus based on public domain audio books.” 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015. (Year: 2015). [cited by examiner]
Wagner et al., “The String-to-String Correction Problem,” Journal of the Association for Computing Machinery, vol. 21, Issue 1, Jan. 1974, pp. 168-173. [cited by applicant]
Barras, C. et al., “Improving Speaker Diarization,” Retrieved from <https://hal.archives-ouvertes.fr/hal-01451540/document> on Apr. 3, 2020 (Year 2004). [cited by applicant]
Barras, Claude et al., “Multistage speaker diarization of broadcast news,” IEEE Transactions on Audio, Speech, and Language Processing 14.5: pp. 1505-1512, (Year 2006). [cited by applicant]
Huijbregts, Marijn et al., “Speaker diarization error analysis using oracle components,” IEEE Transactions on Audio, Speech, and Language Processing 20.2: pp. 393-403, (Year 2011). [cited by applicant]
Elsahar, Hady and Samhaa R. El-Beltagy, “Building large Arabic multi-domain resources for sentiment analysis,” 16th International Conference on Computational Linguistics and Intelligent Text Processing, CICLing, pp. 23-… [cited by applicant]
U.S. Appl. No. 17/084,513, filed Oct. 29, 2020, Transcription Analysis Platform. [cited by applicant]
U.S. Appl. No. 15/614,119 U.S. Pat. No. 10,854,190, filed Jun. 5, 2017 Dec. 1, 2020, Transcription Analysis Platform. [cited by applicant]
U.S. Appl. No. 62/349,290, filed Jun. 13, 2016, Transcription Analysis Platform. [cited by applicant]