IP Library Granted Patent US 12,451,124
Granted Patent B2
US 12,451,124 · App. 17/965,226 · Granted Oct 21, 2025

Domain adaptive speech recognition using artificial intelligence

Inventors: Tohru Nagano (Tokyo, JP); Gakuto Kurata (Tokyo, JP)
Assignee: International Business Machines Corporation
G10L15/16G10L15/02G10L15/063G10L15/30G10L2015/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,124
App. No.
17/965,226
Granted
Oct 21, 2025
Kind
B2
Abstract

Methods, systems, and computer program products for domain adaptive speech recognition using artificial intelligence are provided herein. A computer-implemented method includes generating a set of language data candidates, each language data candidate comprising one or more graphemes, by processing a sequence of phonemes related to input speech data using an artificial intelligence-based data conversion model; determining, for a target pair of phonemes and graphemes, a subset of graphemes from the set of language data candidates; generating a first speech recognition output by processing the subset of graphemes using at least one biasing language model and an artificial intelligence-based speech recognition model; generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one of the graphemes from the target pair; and performing automated actions based on the second speech recognition output.

Claims (44)

1. A computer-implemented method comprising:

generating a set of language data candidates, each language data candidate comprising one or more graphemes, by processing a sequence of phonemes related to user-provided domain-specific input speech data using an artificial intelligence-based data conversion model comprising a neural network model sharing a prediction network from a recurrent neural network transducer;

determining, for a target pair of one or more phonemes and one or more graphemes, a subset of graphemes from the set of language data candidates;

training at least one biasing language model using at least a portion of the subset of graphemes;

generating a first speech recognition output by processing the at least a portion of the subset of graphemes using the at least one biasing language model and an artificial intelligence-based speech recognition model comprising the recurrent neural network transducer, including the prediction network shared by the artificial intelligence-based data conversion model;

generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one of the one or more graphemes from the target pair; and

performing one or more automated actions based at least in part on the second speech recognition output;

wherein the method is carried out by at least one computing device.

2. The computer-implemented method of claim 1 , wherein the artificial intelligence-based data conversion model comprises at least one encoder-decoder neural network.

3. The computer-implemented method of claim 1 , wherein the artificial intelligence-based speech recognition model comprises at least one end-to-end automatic speech recognition model.

4. The computer-implemented method of claim 1 , further comprising:

training the artificial intelligence-based data conversion model using one or more graphemes-phonemes pairs.

5. The computer-implemented method of claim 1 , wherein performing one or more automated actions comprises automatically training at least one of the artificial intelligence-based data conversion model and the artificial intelligence-based speech recognition model using feedback data related to the second speech recognition output.

6. The computer-implemented method of claim 1 , wherein performing one or more automated actions comprises outputting, to at least one user, the second speech recognition output in response to the user-provided domain-specific input speech data.

7. The computer-implemented method of claim 1 , wherein the user-provided domain-specific input speech data comprise one or more of one or more individual words, one or more sentences, and one or more compound words.

8. The computer-implemented method of claim 1 , wherein software implementing the method is provided as a service in a cloud environment.

9. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to:

generate a set of language data candidates, each language data candidate comprising one or more graphemes, by processing a sequence of phonemes related to user-provided domain-specific input speech data using an artificial intelligence-based data conversion model comprising a neural network model sharing a prediction network from a recurrent neural network transducer;

determine, for a target pair of one or more phonemes and one or more graphemes, a subset of graphemes from the set of language data candidates;

train at least one biasing language model using at least a portion of the subset of graphemes;

generate a first speech recognition output by processing the at least a portion of the subset of graphemes using the at least one biasing language model and an artificial intelligence-based speech recognition model comprising the recurrent neural network transducer, including the prediction network shared by the artificial intelligence-based data conversion model;

generate a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one of the one or more graphemes from the target pair; and

perform one or more automated actions based at least in part on the second speech recognition output.

10. The computer program product of claim 9 , wherein the artificial intelligence-based data conversion model comprises at least one encoder-decoder neural network.

11. The computer program product of claim 9 , wherein the artificial intelligence-based speech recognition model comprises at least one end-to-end automatic speech recognition model.

12. The computer program product of claim 9 , wherein performing one or more automated actions comprises outputting, to at least one user, the second speech recognition output in response to the user-provided domain-specific input speech data.

13. The computer program product of claim 9 , wherein the program instructions executable by the computing device further cause the computing device to:

train the artificial intelligence-based data conversion model using one or more graphemes- phonemes pairs.

14. The computer program product of claim 9 , wherein performing one or more automated actions comprises automatically training at least one of the artificial intelligence-based data conversion model and the artificial intelligence-based speech recognition model using feedback data related to the second speech recognition output.

15. A system comprising:

a memory configured to store program instructions; and

a processor operatively coupled to the memory to execute the program instructions to:

generate a set of language data candidates, each language data candidate comprising one or more graphemes, by processing a sequence of phonemes related to user-provided domain-specific input speech data using an artificial intelligence-based data conversion model comprising a neural network model sharing a prediction network from a recurrent neural network transducer;

determine, for a target pair of one or more phonemes and one or more graphemes, a subset of graphemes from the set of language data candidates;

train at least one biasing language model using at least a portion of the subset of graphemes;

generate a first speech recognition output by processing the at least a portion of the subset of graphemes using the at least one biasing language model and an artificial intelligence-based speech recognition model comprising the recurrent neural network transducer, including the prediction network shared by the artificial intelligence-based data conversion model;

generate a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one of the one or more graphemes from the target pair; and

perform one or more automated actions based at least in part on the second speech recognition output.

16. The system of claim 15 , wherein the artificial intelligence-based data conversion model comprises at least one encoder-decoder neural network.

17. The system of claim 15 , wherein performing one or more automated actions comprises outputting, to at least one user, the second speech recognition output in response to the user-provided domain-specific input speech data.

18. The system of claim 15 , wherein the artificial intelligence-based speech recognition model comprises at least one end-to-end automatic speech recognition model.

19. The system of claim 15 , wherein the processor is operatively coupled to the memory to further execute the program instructions to:

train the artificial intelligence-based data conversion model using one or more graphemes-phonemes pairs.

20. The system of claim 15 , wherein performing one or more automated actions comprises automatically training at least one of the artificial intelligence-based data conversion model and the artificial intelligence-based speech recognition model using feedback data related to the second speech recognition output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: NAGANO, TOHRU; KURATA, GAKUTO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061412/0679 →
Continuity (1)
Related Publication 20240127801A1 · Apr 18, 2024
References Cited (30)
US 6910012B2 · Hartley et al. · 2005 [cited by applicant]
US 8751230B2 · Saffer · 2014 [cited by examiner]
US 10535339B2 · Fujimura · 2020 [cited by examiner]
US 11195513B2 · Kurata et al. · 2021 [cited by applicant]
US 11238845B2 · Chen · 2022 [cited by examiner]
US 12014725B2 · Huang · 2024 [cited by examiner]
US 20090150153A1 · Li et al. · 2009 [cited by applicant]
US 20180211652A1 · Mun · 2018 [cited by examiner]
US 20180286386A1 · Baughman et al. · 2018 [cited by applicant]
US 20190096390A1 · Kurata · 2019 [cited by examiner]
US 20190122654A1 · Song · 2019 [cited by examiner]
US 20200027445A1 · Raghunathan et al. · 2020 [cited by applicant]
US 20210295846A1 · Yang · 2021 [cited by examiner]
US 20210312294A1 · Kurata · 2021 [cited by examiner]
US 20220172706A1 · Hu · 2022 [cited by examiner]
US 20220310071A1 · Botros · 2022 [cited by examiner]
US 20220310077A1 · Tu · 2022 [cited by examiner]
US 20220319506A1 · Heikinheimo · 2022 [cited by examiner]
CN 108364651A · 2018 [cited by applicant]
CN 112309393A · 2021 [cited by applicant]
CN 113692616A · 2021 [cited by applicant]
CN 115132175A · 2022 [cited by applicant]
CN 120019431A · 2025 [cited by applicant]
DE 112023003661T5 · 2025 [cited by applicant]
WO 2024078565A1 · 2024 [cited by applicant]
Kanishka Rao, Has,im Sak, Rohit Prabhavalkar, “Exploring Architectures, Data and Units for Streaming End-To-End Speech Recognition With RNN-Transducer”, 2017 IEEE Automatic Speech Recognition and Understanding Workshop … [cited by examiner]
Rao et al., Exploring Architectures, Data and Units for Streaming End-To-End Speech Recognition with RNN-Transducer, IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 2017. [cited by applicant]
Masumura et al., Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition, INTERSPEECH 2020, Oct. 2020. [cited by applicant]
Samarakoon et al., Domain Adaptation of End-to-end Speech Recognition in Low-Resource Settings, IEEE Spoken Language Technology Workshop (SLT), Dec. 2018. [cited by applicant]
Decadt et al., Phoneme-to-Grapheme Conversion for Out-of-Vocabulary Words in Speech Recognition, IEEE Automatic Speech Recognition and Understanding Workshop, Oct. 2001. [cited by applicant]