IP Library › Granted Patent US 12,367,859
Granted Patent B2
US 12,367,859 · App. 17/809,202 · Granted Jul 22, 2025

Artificial intelligence factsheet generation for speech recognition

Inventors: Shreya Khare (Bangalore, IN); Ashish R. Mittal (Bengaluru, IN); Saneem Ahmed Chemmengath (Bangalore, IN); Samarth Bharadwaj (Bangalore, IN); Karthik Sankaranarayanan (Bangalore, IN)
Assignee: International Business Machines Corporation
G10L15/01G10L15/16G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,859
App. No.
17/809,202
Granted
Jul 22, 2025
Kind
B2
Abstract

A method, system, and computer program product for automated artificial intelligence (AI) factsheet generation for modeling and model customization in speech to text (STT) services. The method receives audio data for a user. The audio data contains human speech. Text data is generated, using a first speech to text model, to represent the human speech of the audio data. A set of transcription errors of the first speech to text model are identified. A set of AI factsheets are generated to describe model metadata for the first speech to text model. Based on the set of transcription errors and the set of AI factsheets, the method generates a second speech to text model customized to the user.

Claims (47)

1. A computer-implemented method, comprising:

receiving audio data for a user, the audio data containing human speech;

generating, using a first speech to text model, text data representing the human speech of the audio data;

identifying a set of transcription errors of the first speech to text model;

generating a set of artificial intelligence (AI) factsheets describing model metadata for the first speech to text model, wherein at least one of the set of AI factsheets is a confidence sheet including a local explanation of the first speech to text model that provides values of audio quality, transcript quality, stylization and data redaction metrics, wherein the set of AI factsheets is a plurality of AI factsheets and generating the set of AI factsheets further comprises:

generating a first AI factsheet for the first speech to text model, the first AI factsheet being a model factsheet; and

generating a second AI factsheet for the first speech to text model, the second AI factsheet being a transcript factsheet; and

based on the first speech to text model, the set of transcription errors, the text data, the first AI factsheet, and the second AI factsheet, generating a second speech to text model customized to the user, wherein the first speech to text model identifies training data or subsets thereof for input to train the second speech to text model, and selecting the training data to address deficiencies represented by the set of transcription errors identified from the first speech to text model.

2. The method of claim 1 , wherein the set of transcription errors are a set of automatic speech recognition (ASR) errors and wherein identifying the set of ASR errors further comprises:

determining a set of characteristics for the set of ASR errors; and

attributing the set of ASR errors to a set of speech features of the audio data based on the set of characteristics.

3. The method of claim 2 , wherein identifying the set of ASR errors further comprises:

clustering ASR errors of the set of ASR errors to generate ASR error clusters, each ASR error cluster associated with at least one speech feature of the audio data.

4. The method of claim 1 , further comprising:

determining a customization level for the second speech to text model based on the set of AI factsheets.

5. A system, comprising:

one or more processors; and

a computer-readable storage medium, coupled to the one or more processors, storing program instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving audio data for a user, the audio data containing human speech;

generating, using a first speech to text model, text data representing the human speech of the audio data;

identifying a set of transcription errors of the first speech to text model;

generating a set of artificial intelligence (AI) factsheets describing model metadata for the first speech to text model, wherein at least one of the set of AI factsheets is a confidence sheet including a local explanation of the first speech to text model that provides values of audio quality, transcript quality, stylization and data redaction metrics, wherein the set of AI factsheets is a plurality of AI factsheets and generating the set of AI factsheets further comprises:

generating a first AI factsheet for the first speech to text model, the first AI factsheet being a model factsheet; and

generating a second AI factsheet for the first speech to text model, the second AI factsheet being a transcript factsheet; and

based on the first speech to text model, the set of transcription errors, the text data, the first AI factsheet, and the second AI factsheet, generating a second speech to text model customized to the user, wherein the first speech to text model identifies training data or subsets thereof for input to train the second speech to text model, and selecting the training data to address deficiencies represented by the set of transcription errors identified from the first speech to text model.

6. The system of claim 5 , wherein the set of transcription errors are a set of automatic speech recognition (ASR) errors and wherein identifying the set of ASR errors further comprises:

determining a set of characteristics for the set of ASR errors; and

attributing the set of ASR errors to a set of speech features of the audio data based on the set of characteristics.

7. The system of claim 6 , wherein identifying the set of ASR errors further comprises:

clustering ASR errors of the set of ASR errors to generate ASR error clusters, each ASR error cluster associated with at least one speech feature of the audio data.

8. The system of claim 5 , wherein the operations further comprise:

determining a customization level for the second speech to text model based on the set of AI factsheets.

9. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions being executable by one or more processors to cause the one or more processors to perform operations comprising:

receiving audio data for a user, the audio data containing human speech;

generating, using a first speech to text model, text data representing the human speech of the audio data;

identifying a set of transcription errors of the first speech to text model;

generating a set of artificial intelligence (AI) factsheets describing model metadata for the first speech to text model, wherein at least one of the set of AI factsheets is a confidence sheet including a local explanation of the first speech to text model that provides values of audio quality, transcript quality, stylization and data redaction metrics, wherein the set of AI factsheets is a plurality of AI factsheets and generating the set of AI factsheets further comprises:

generating a first AI factsheet for the first speech to text model, the first AI factsheet being a model factsheet; and

generating a second AI factsheet for the first speech to text model, the second AI factsheet being a transcript factsheet; and

based on the first speech to text model, the set of transcription errors, the text data, the first AI factsheet, and the second AI factsheet, generating a second speech to text model customized to the user, wherein the first speech to text model identifies training data or subsets thereof for input to train the second speech to text model, and selecting the training data to address deficiencies represented by the set of transcription errors identified from the first speech to text model.

10. The computer program product of claim 9 , wherein the set of transcription errors are a set of automatic speech recognition (ASR) errors and wherein identifying the set of ASR errors further comprises:

determining a set of characteristics for the set of ASR errors; and

attributing the set of ASR errors to a set of speech features of the audio data based on the set of characteristics.

11. The computer program product of claim 10 , wherein identifying the set of ASR errors further comprises:

clustering ASR errors of the set of ASR errors to generate ASR error clusters, each ASR error cluster associated with at least one speech feature of the audio data.

12. The computer program product of claim 9 , wherein the operations further comprise:

determining a customization level for the second speech to text model based on the set of AI factsheets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2022
From: KHARE, SHREYA; MITTAL, ASHISH R; CHEMMENGATH, SANEEM AHMED; BHARADWAJ, SAMARTH; SANKARANARAYANAN, KARTHIK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060325/0197 →
Continuity (1)
Related Publication 20230419950A1 · Dec 28, 2023
References Cited (23)
US 6622121B1 · Crepy · 2003 [cited by applicant]
US 7684988B2 · Barquilla · 2010 [cited by applicant]
US 10276163B1 · Lei · 2019 [cited by applicant]
US 11315570B2 · Allibhai · 2022 [cited by examiner]
US 20050049868A1 · Busayapongchai · 2005 [cited by applicant]
US 20060235687A1 · Carus · 2006 [cited by examiner]
US 20170329466A1 · Krenkler · 2017 [cited by applicant]
US 20190013038A1 · Thomson · 2019 [cited by examiner]
US 20190206389A1 · Kwon · 2019 [cited by examiner]
US 20190287519A1 · Ediz · 2019 [cited by examiner]
US 20210133162A1 · Arnold · 2021 [cited by applicant]
US 20210217403A1 · Chae · 2021 [cited by applicant]
CN 102723080B · 2014 [cited by applicant]
Bharadhwaj, Homanga, “Layer-wise relevance propagation for explainable deep learning based speech recognition,” 2018, 6 pages. [cited by applicant]
Calore, Michael “Watch People with Accents Confuse the Hell Out of Al Assistants,” https://www.wired.com/2017/05/ai-assistants-accented-english/, downloaded from the internet Jun. 24, 2022, 1 page. [cited by applicant]
Danilevsky et al., “A Survey of the State of Explainable Al for Natural Language Processing,” arXiv:2010.0071v1 [cs.CL], Oct. 1, 2020, 13 pages. [cited by applicant]
Errattahi, et al., “Automatic Speech Recognition Error Detection Using Supervised Learning Techniques,” 13th International Conference of Computer Systems and Applications, Nov. 29, 2016, 7 pages. [cited by applicant]
Errattahi, et al., “Automatic Speech Recognition Errors Detection and Correction: A Review,” ScienceDirect Procedia Computer Science, 2018, pp. 32-37, vol. 128. [cited by applicant]
Lin, Zhong Qiu, “Quantifying the Performance of Explainability Algorithms,” 2020, 72 pages. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing,” Recommendations of the National Institute of Standards and Technology, U.S. Department of Commerce, Special Publication 800-145, Sep. 2011, 7 pgs. [cited by applicant]
Mirzaei, et al., “Errors in automatic speech recognition versus difficulties in second language listening,” Critical CALL—Proceedings of the 2015 EUROCALL Conference, 2015, pp. 410-415, Research-publishing. net, Dublin … [cited by applicant]
Mirzaei, et al., “Exploiting Automatic Speech Recognition to Enhance Partial and Synchronized Caption for Facilitating second language Listening,” Computer Speech and Language, May 2018, 21 pgs. [cited by applicant]
“Human and Humanizing Speech Technology”, Interspeech, retrieved from web https://www.interspeech2022.org, Sep. 18-22, 2022, 75 pages. [cited by applicant]