IP Library › Granted Patent US 12,548,555
Granted Patent B2
US 12,548,555 · App. 18/347,031 · Granted Feb 10, 2026

Joint training of speech recognition and speech synthesis models for conversational AI systems and applications

Inventors: Xianchao Wu (Tokyo, JP); Yi Dong (Lexington, MA); Scott Nunweiler (Yokohama, JP)
Assignee: NVIDIA Corporation
G10L15/063G10L13/047G10L15/16G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,555
App. No.
18/347,031
Granted
Feb 10, 2026
Kind
B2
Abstract

Disclosed are systems and techniques for training machine learning models. The techniques include providing a first data of a first modality as input to a first machine learning model to obtain a first output of a second modality, providing the first output of the second modality as input to a second machine learning model to obtain a second output of the first modality, providing the first data as input to a third machine learning model to obtain a first tensor, providing the second output as input to the third machine learning model to obtain a second tensor, calculating a first loss based on a comparison between the first tensor and the second tensor, and causing the first machine learning model to be modified based on the first loss.

Claims (68)

1 . A method comprising:

providing first data of a first modality as input to a first machine learning model to obtain a first output of a second modality, wherein the first modality and the second modality comprise at least one of a text-based modality or an audio-based modality and are distinct from one another;

providing the first output of the second modality as input to a second machine learning model to obtain a second output of the first modality, at least one of the first machine learning model or the second machine learning model being based at least on a diffusion generative adversarial network (diffusion-GAN) model that includes a discriminator comprising a plurality of timestep-dependent discriminators;

providing the first data as input to a third machine learning model to obtain a first tensor;

providing the second output as input to the third machine learning model to obtain a second tensor;

calculating a loss based at least on a comparison between the first tensor and the second tensor; and

updating one or more parameters of the first machine learning model based at least on the loss.

2 . The method of claim 1 , further comprising updating one or more parameters of the second machine learning model based at least on the loss.

3 . The method of claim 1 , further comprising:

receiving second data of the second modality;

providing the second data as input to the second machine learning model to obtain a third output of the first modality;

providing the third output as input to the first machine learning model to obtain a fourth output of the second modality;

providing the second data as input to a fourth machine learning model to obtain a third tensor;

providing the fourth output as input to the fourth machine learning model to obtain a fourth tensor;

calculating a second loss based at least on a comparison between the third tensor and the fourth tensor; and

updating the one or more parameters of the second machine learning model based at least on the second loss.

4 . The method of claim 3 , wherein the updating the one or more parameters of the first machine learning model is further based at least on a weighted combination of the loss and the second loss.

5 . The method of claim 1 , wherein the first machine learning model is an unsupervised text-to-speech model and the second machine learning model is an unsupervised automatic speech recognition model.

6 . The method of claim 5 , wherein the unsupervised text-to-speech model is based at least on the diffusion-GAN model, and wherein the diffusion-GAN model further comprises:

a Transformer-based language model;

a convolutional neural network for generating embeddings based at least on an audio input; and

a generator model.

7 . The method of claim 1 , wherein the first machine learning model is an unsupervised automatic speech recognition model and the second machine learning model is an unsupervised text-to-speech model.

8 . A method comprising:

applying a first deployed machine learning model to first data of a first modality to obtain a second data of a second modality, wherein the first deployed machine learning model is trained, at least in part, by:

providing a first training data of the first modality as input to a first machine learning model to obtain a first output of the second modality, wherein the first modality and the second modality comprise at least one of a text-based modality or an audio-based modality and are distinct from one another;

providing the first output of the second modality as input to a second machine learning model to obtain a second output of the first modality, at least one of the first machine learning model or the second machine learning model being based at least on a diffusion generative adversarial network (diffusion-GAN) model that includes a discriminator comprising a plurality of timestep-dependent discriminators;

providing the first training data as input to a third machine learning model to obtain a first tensor;

providing the second output as input to the third machine learning model to obtain a second tensor;

calculating a loss based at least on a comparison between the first tensor and the second tensor; and

causing the first machine learning model to be modified based at least on the loss, wherein the first machine learning model, after training, represents the first deployed machine learning model.

9 . The method of claim 8 , further comprising:

receiving third data of the second modality;

applying a second deployed machine learning model to the third data to obtain fourth data of the first modality, wherein the second deployed machine learning model is trained, at least in part, by:

receiving second training data of the second modality;

providing the second training data as input to the second machine learning model to obtain a third output of the first modality;

providing the third output as input to the first machine learning model to obtain a fourth output of the second modality;

providing the second training data as input to a fourth machine learning model to obtain a third tensor;

providing the fourth output as input to the fourth machine learning model to obtain a fourth tensor;

calculating a second loss based at least on a comparison between the third tensor and the fourth tensor; and

causing the second machine learning model to be modified based at least on the second loss, wherein the second machine learning model, after training, represents the second deployed machine learning model.

10 . The method of claim 9 , wherein the causing the first machine learning model to be modified is further based at least on a weighted combination of the loss and the second loss.

11 . The method of claim 8 , wherein the first machine learning model is an unsupervised text-to-speech model and the second machine learning model is an unsupervised automatic speech recognition model.

12 . The method of claim 11 , wherein the unsupervised automatic speech recognition model is based at least on the diffusion-GAN model, and wherein the diffusion-GAN model further comprises:

a Transformer-based language model;

a convolutional neural network for generating embeddings based at least on an audio input; and

a generator model.

13 . The method of claim 8 , wherein the first machine learning model is an unsupervised automatic speech recognition model and the second machine learning model is an unsupervised text-to-speech model.

14 . A system comprising:

one or more processing units to perform one or more operations using a deployed machine learning model, the deployed machine learning model trained, at least in part, using a loss function that penalizes a difference between a first tensor representation of first audio data and a second tensor representation of second audio data, the second audio data generated, at least in part, by converting the first audio data to text data and converting the text data to the second audio data, wherein the deployed machine learning model is based at least on a diffusion generative adversarial network (diffusion-GAN) model that includes a discriminator comprising a plurality of timestep-dependent discriminators.

15 . The system of claim 14 , wherein the converting the text data to the second audio data is performed using a text-to-speech model.

16 . The system of claim 14 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content;

a system implemented using a robot;

a system for performing conversational AI operations;

a system implementing one or more large language models (LLMs);

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2023
From: WU, XIANCHAO; DONG, YI; NUNWEILER, SCOTT
To: NVIDIA CORPORATION
Reel/Frame 064155/0432 →
Continuity (1)
Related Publication 20250014571A1 · Jan 9, 2025
References Cited (25)
US 20210350786A1 · Chen · 2021 [cited by examiner]
US 20220005457A1 · Balakrishnan · 2022 [cited by examiner]
US 20230023691A1 · Knerr · 2023 [cited by examiner]
US 20230298567A1 · Tan · 2023 [cited by examiner]
US 20240289563A1 · Ramanovich · 2024 [cited by examiner]
Ren, Yi, et al. “Almost unsupervised text to speech and automatic speech recognition.” International conference on machine learning. PMLR, 2019. (Year: 2019). [cited by examiner]
He, Di, et al. “Dual learning for machine translation.” Advances in neural information processing systems 29 (2016). (Year: 2016). [cited by examiner]
Karita, Shigeki, et al. “Semi-supervised end-to-end speech recognition using text-to-speech and autoencoders.” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 20… [cited by examiner]
Wang, Zhendong, et al. “Diffusion-gan: Training gans with diffusion.” arXiv preprint arXiv:2206.02262 (2022). (Year: 2022). [cited by examiner]
Baevski, et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, 12 pages. https://proceedin… [cited by applicant]
Baevski, et al., “Unsupervised speech recognition,” 35th Conference on Neural Processing Systems (NeurIPS 2021), 14 pages. https://proceedings.neurips.cc/paper/2021/file/ea159dc9788ffac311592613b7f71fbb-Paper.pdf. [cited by applicant]
Blei, et al., “Variational Inference: A Review for Statisticians,” Journal of the American Statistical Association (2017), 112:518, 859-877. DOI: 10.1080/01621459.2017.1285773. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Proceedings of NAACL-IILT 2019, pp. 4171-4186. doi: 10.18653/v1/N19-1423. https://aclanthology.org/N19-1423. [cited by applicant]
Endres and Schindelin, “A New Metric for Probability Distributions,” IEEE Transactions on Information Theory, 49(7):1858-1860, 2003. doi: 10.1109/TIT.2003.813506. [cited by applicant]
Gulrajani, et al., “Improved Training of Wasserstein GANS,” Proceedings of the 31st NeurIPS, NIPS'17, pp. 5769-5779, 2017. ISBN 9781510860964. [cited by applicant]
Ho, et al., “Denoising Diffusion Probabilistic Models,” arXiv:2006.11239v2 [cs.LG] Dec. 16, 2020, 25 pages. https://arxiv.org/abs/2006.11239. [cited by applicant]
Ho, et al., “Denoising Diffusion Probabilistic Models,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, 12 pages. https://proceedings.neurips.cc/paper/2020/file/4c5bcfec8584af… [cited by applicant]
Ni, et al., “Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition,” arXiv:2203.15796v2 [eess.AS] Aug. 15, 2022, 5 pages. https://arxiv.org/abs/2203.15796. [cited by applicant]
Pratap et al., “MLS: A Large-Scale Multilingual Dataset for Speech Research,” Interspeech 2020, Oct. 25-29, 2020, Shanghai, China, 2757-2761. doi:10.21437/interspeech.2020-2826. http://dx.doi.org/10.21437/Interspeech.20… [cited by applicant]
Ronneberger, et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” arXiv:1505.04597v1 [cs.CV] May 18, 2015, 8 pages. http://arxiv.org/abs/1505.04597. [cited by applicant]
Sohl-Dickstein, et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics,” Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 2015, JMLR: W&CP vol. 37, 18 pages. http://a… [cited by applicant]
Song and Ermon, “Generative Modeling by Estimating Gradients of the Data Distribution,” 33rd Conference on Neural Information Processing Systems (NEURIPS 2019), Vancouver, Canada, 13 pages. https://proceedings.neurips.c… [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 11 pages. https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c… [cited by applicant]
Wang and Cho, “BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model,” Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation (NeuralGen), pp. 30-36, … [cited by applicant]
Wang, et al., “Diffusion-GAN: Training GANs with Diffusion,” arXiv preprint arXiv:2206.02262v3 [cs.LG] Oct. 9, 2022, 28 pages. https://arxiv.org/abs/2206.02262.12. [cited by applicant]