IP Library › Granted Patent US 12,444,409
Granted Patent B2
US 12,444,409 · App. 18/468,086 · Granted Oct 14, 2025

Hybrid language models for conversational AI systems and applications

Inventors: Vladimir Bataev (Yerevan, AM); Roman Korostik (Yerevan, AM); Evgenii Shabalin (Moscow, RU); Vitaly Sergeyevich Lavrukhin (Campbell, CA); Boris Ginsburg (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G10L15/16G10L15/065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,409
App. No.
18/468,086
Granted
Oct 14, 2025
Kind
B2
Abstract

In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.

Claims (68)

1. A processor comprising:

one or more circuits to perform automatic speech recognition (ASR) using one or more ASR machine learning models (MLMs), the one or more ASR MLMs trained, at least, by:

generating, using one or more ASR MLMs and first textual data, one or more spectrograms;

generating, using the one or more ASR MLMs and the one or more spectrograms, output data indicating second textual data; and

updating one or more parameters of the one or more ASR MLMs based at least on the output data and ground truth data associated with the first textual data.

2. The processor of claim 1 , wherein the one or more ASR MLMs are initially trained using first training sets of audio data inputs and textual data ground truth, and the updating the one or more parameters is performed as part of adapting the one or more ASR MLMs to a target domain using second training sets of textual data inputs and textual data ground truth.

3. The processor of claim 1 , wherein the one or more ASR MLMs are further trained, at least, by:

generating one or more second spectrograms using audio data;

generating, using the one or more ASR MLMs and the one or more second spectrograms, second output data indicating third textual data; and

updating the one or more parameters of the one or more ASR MLMs based at least on the second output data and ground truth data associated with the audio data.

4. The processor of claim 1 , wherein the one or more ASR MLMs are further trained, at least, by:

providing text input to at least one first MLM of the one or more ASR MLMs;

generating, based at least on the text input, one or more initial spectrograms using the at least one first MLM; and

generating, based at least on the one or more initial spectrograms, the one or more spectrograms using one or more generators from one or more generative adversarial networks (GANs).

5. The processor of claim 1 , wherein the generating the one or more spectrograms includes applying an initial version of the one or more spectrograms as input to at least two generator blocks of one or more generators of the one or more ASR MLMs.

6. The processor of claim 1 , wherein the one or more ASR MLMs are trained as a text-to-speech (TTS) model.

7. The processor of claim 1 , wherein the processor is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing one or more simulation operations;

a system for performing one or more digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing one or more deep learning operations;

a system implementing one or more language models;

a system implementing one or more large language models (LLMs);

a system for performing one or more generative AI operations;

a system implemented using an edge device;

a system implemented using a machine;

a system for performing one or more conversational AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

8. A method comprising:

determining one or more first spectrograms corresponding to at least one of first audio data or first textual data;

generating one or more second spectrograms using the one or more first spectrograms and one or more generators from one or more generative adversarial networks (GANs); and

determining output data indicating at least one of second audio data or second textual data using one or more Machine Learning Models (MLMs) and the one or more second spectrograms.

9. The method of claim 8 , further comprising updating, based at least on the output data, one or more parameters of the one or more MLMs using ground truth data associated with the at least one of the first audio data or the first textual data.

10. The method of claim 8 , wherein the determining the output data is based at least on applying the one or more second spectrograms to the one or more MLMs.

11. The method of claim 8 , wherein the generating the one or more second spectrograms includes applying the one or more first spectrograms as input to at least two generator blocks of the one or more generators.

12. The method of claim 8 , wherein the one or more MLMs are initially trained using first training sets of audio data inputs and textual data ground truth, and the output data is used to adapt the one or more MLMs to a target domain.

13. The method of claim 8 , further comprising generating the one or more first spectrograms using one or more second MLMs.

14. The method of claim 8 , wherein the one or more MLMs are deployed as at least part of a language processing system.

15. A system comprising:

one or more processing units to perform one or more operations using one or more first machine learning models (MLMs), the one or more first MLMs trained, at least, by generating, using the one or more first MLMs and one or more audio representations, output data indicating first textual data, the one or more audio representations generated using second textual data applied to one or more second MLMs.

16. The system of claim 15 , wherein the one or more first MLMs are initially trained using first training sets of audio data inputs and textual data ground truth, and the output data is used to adapt the one or more first MLMs to a target domain.

17. The system of claim 15 , wherein the one or more audio representations are generated based at least on applying one or more initial audio representations to one or more generators from one or more generative adversarial networks (GANs).

18. The system of claim 15 , wherein the one or more audio representations are generated, at least, by applying an initial version of the one or more audio representations as input to at least two generator blocks of one or more generators.

19. The system of claim 15 , wherein the one or more second MLMs include one or more TTS models.

20. The system of claim 15 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing one or more simulation operations;

a system for performing one or more digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing one or more deep learning operations;

a system implementing one or more language models;

a system implementing one or more large language models (LLMs);

a system for performing one or more generative AI operations;

a system implemented using an edge device;

a system implemented using a machine;

a system for performing one or more conversational AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: BATAEV, VLADIMIR; KOROSTIK, ROMAN; SHABALIN, EVGENII; LAVRUKHIN, VITALY SERGEYEVICH; GINSBURG, BORIS
To: NVIDIA CORPORATION
Reel/Frame 064921/0357 →
Continuity (3)
Provisional Application 63417627 · Oct 19, 2022
Related Publication 20240135920A1 · Apr 25, 2024
Related Publication 20240233714A9 · Jul 11, 2024
References Cited (43)
US 11138471B2 · Park · 2021 [cited by examiner]
US 11605384B1 · Dalton et al. · 2023 [cited by applicant]
US 20130225240A1 · Largey et al. · 2013 [cited by applicant]
US 20200364303A1 · Liu et al. · 2020 [cited by applicant]
US 20210074316A1 · Souden · 2021 [cited by examiner]
US 20210150187A1 · Karras et al. · 2021 [cited by applicant]
US 20220012537A1 · Park · 2022 [cited by examiner]
US 20220028390A1 · Poznanski · 2022 [cited by examiner]
US 20220201121A1 · Kane · 2022 [cited by examiner]
US 20230326445A1 · Adam · 2023 [cited by examiner]
Karras, et al., “Analyzing and Improving the Image Quality of StyleGAN”, https://arxiv.org/abs/1912.04958, Mar. 23, 2020, 21 pgs. [cited by applicant]
Ioffe, S., and Szegedy, C., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, arXiv:1502.03167v3 [cs.LG], pp. 1-12 (Mar. 2, 2015), Available at: https://arxiv.org/abs/1502.0… [cited by applicant]
Kingma, D. P., and Ba, J. L., “Adam: A Method for Stochastic Optimization”, published as a conference paper at CLR 2015, arXiv:1412.6980v9 [cs.LG], pp. 1-15 (Jan. 30, 2017). [cited by applicant]
Gulati, et al.; “Conformer: Convolution-augmented Transformer for Speech Recognition,” https://arxiv.org/abs/2005.08100; May 16, 2020, 5 pgs. [cited by applicant]
Vaswani, et al.; “Attention is All You Need,” https://arxiv.org/abs/1706.03762; Aug. 2, 2023, 15 pgs. [cited by applicant]
Li, et al.; “Recent Advances in End-to-End Automatic Speech Recognition,” https://arxiv.org/abs/2111.01690, Feb. 2, 2022, 27 pgs. [cited by applicant]
Li, et al.; “Training neural speech recognition systems with synthetic speech augmentation,” https://arxiv.org/abs/1811.00707, Nov. 2, 2018, 5 pgs. [cited by applicant]
Laptev, et al.; “You do not need more data: Improving end-to-end speech recognition by text-to-speech data augmentation,” https://arxiv.org/abs/2005.07157, Jul. 30, 2020, 6 pgs. [cited by applicant]
Li, et al.; “Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability,” https://arxiv.org/abs/2007.15188, Jul. 30, 2020, 5 pgs. [cited by applicant]
Zheng, et al.; “Using Synthetic Audio to Improve the Recognition of Out-of-Vocabulary Words in End-to-End ASR Systems,” https://arxiv.org/abs/2011.11564, Feb. 10, 2021, 5 pgs. [cited by applicant]
Thomas, et al.; “Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models,” https://arxiv.org/abs/2202.13155, Feb. 26, 2022, 5 pgs. [cited by applicant]
Chen, et al.; “MAESTRO: Matched Speech Text Representations Through Modality Matching,”https://arxiv.org/abs/2204.03409, Jul. 1, 2022, 5 pgs. [cited by applicant]
Sato, et al.; “Text-only Domain Adaptation Based on Intermediate CTC,” Interspeech 2022, Sep. 2022, 5 pgs. [cited by applicant]
Meng, et al.; “Modular Hybrid Autoregressive Transducer,” https://arxiv.org/abs/2210.17049, Feb. 17, 2023, 8 pgs. [cited by applicant]
Ueno, et al.; “Phone-informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition,” in ICASSP, 2022, 5 pgs. [cited by applicant]
Hori, et al.; “Cycle-Consistency Training for End-to-End Speech Recognition,” https://arxiv.org/abs/1811.01690, May 23, 2019, 5 pgs. [cited by applicant]
Wang, et al.; “Improving Speech Recognition Using Consistent Predictions on Synthesized Speech,” ICASSP, 2020, 5 pgs. [cited by applicant]
Baskar, et al.; “Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition,” https://arxiv.org/abs/2104.07474, Apr. 13, 2021, 5 pgs. [cited by applicant]
Ueno, et al.; “Data Augmentation for ASR Using TTS Via A Discrete Representation,” 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, 8 pgs. [cited by applicant]
Kurata, et al.; “Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio,” in Interspeech, 2021, 5 pgs. [cited by applicant]
Lancucki; “FastPitch: Parallel Text-to-Speech with Pitch Prediction,” https://arxiv.org/abs/2006.06873, Feb. 16, 2021, 5 pgs. [cited by applicant]
Kuchaiev, et al.; “NeMo: A Toolkit for Building AI Applications Using Neural Modules,” https://arxiv.org/abs/1909.09577, Sep. 14, 2019, 8 pgs. [cited by applicant]
Lim, et al.; “Geometric GAN,” https://arxiv.org/abs/1705.02894, May 9, 2017, 17 pgs. [cited by applicant]
Mescheder, et al.; “Which Training Methods for GANs Do Actually Converge,” https://arxiv.org/abs/1801.04406, Jul. 31, 2018, 39 pgs. [cited by applicant]
Ba, et al.; “Layer Normalization,” https://arxiv.org/abs/1607.06450, Jul. 21, 2016, 14 pgs. [cited by applicant]
Loshchilov, et al.; “Decoupled Weight Decay Regularization,” https://arxiv.org/abs/1711.05101, Jan. 4, 2019, 19 pgs. [cited by applicant]
Loshchilov, et al.; “SGDR: Stochastic Gradient Descent With Warm Restarts,” https://arxiv.org/abs/1608.03983, May 3, 2017, 16 pgs. [cited by applicant]
Bastianelli, et al.; “SLURP: A Spoken Language Understanding Resource Package,” https://arxiv.org/abs/2011.13205, Nov. 26, 2020, 11 pgs. [cited by applicant]
Paul et al.; “The Design For The Wall Street Journal-Based CSR Corpus,” in Speech and Natural Language Workshop, 1992, 6 pgs. [cited by applicant]
Park, et al.; “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” https://arxiv.org/abs/1904.08779, Dec. 3, 2019, 6 pgs. [cited by applicant]
Jang, et al.; “UnivNet: A Neural Vocoder With Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation,” https://arxiv.org/abs/2106.07889, Jun. 15, 2021, 5 pgs. [cited by applicant]
Zen, et al.; “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” https://arxiv.org/abs/1904.02882, Apr. 5, 2019, 7 pgs. [cited by applicant]
Panayotov, et al.; “LibriSpeech: An ASR Corpus Based on Public Domain Audio Books,” in ICASSP, 2015, 5 pgs. [cited by applicant]