IP Library Granted Patent US 12,499,326
Granted Patent B2
US 12,499,326 · App. 18/278,364 · Granted Dec 16, 2025

Multi-model joint denoising training

Inventors: Linjun Shou (Beijing, CN); Ming Gong (Beijing, CN); Xuanyu Bai (Beijing, CN); Xuguang Wang (Beijing, CN); Daxin Jiang (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
G06F40/44G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,326
App. No.
18/278,364
Granted
Dec 16, 2025
Kind
B2
Abstract

The present disclosure proposes a method and apparatus for multi-model joint denoising training. Multiple models may be obtained. A set of training samples may be denoised through the multiple models. The multiple models may be trained with the set of denoised training samples.

Claims (47)

1 . A method for multi-model joint denoising training, the method comprising:

obtaining multiple machine learning models having identical neural network architectures, wherein each machine learning model is initialized with a different combination of training corpora selected from a source training corpus containing reliable labels, at least one translated training corpus generated through machine translation, and at least one generated training corpus created through neural language generation;

jointly denoising a set of training samples through the multiple models by:

for each model, selecting a subset of training samples from the set of training samples by filtering out a predetermined proportion of training samples having highest prediction losses as determined by other models in the multiple models, and

for each training sample, determining a weight for calculating training loss based on consistency of prediction results obtained from the multiple models, wherein inconsistent prediction results indicate noisy training samples;

training the multiple models simultaneously with their respective selected subsets of denoised training samples using the determined weights; and

iteratively updating labels of training samples from translated and generated training corpora based on consensus predictions from the trained multiple models.

2 . The method of claim 1 , wherein the set of training samples comes from a training data set, the training data set including two or two of a source training corpus, at least one translated training corpus, and at least one generated training corpus.

3 . The method of claim 1 , further comprising:

initializing the multiple models with multiple initialization data groups, and the multiple initialization data groups are formed by at least one training corpus in a training data set including the set of training samples.

4 . The method of claim 1 , wherein the denoising a set of training samples comprises:

for each model in the multiple models, selecting training samples for the model from the set of training samples through other models in the multiple models.

5 . The method of claim 4 , wherein the selecting training samples for the model comprises:

filtering out a predetermined proportion of training samples from the set of training samples through the other models; and

determining remaining training samples in the set of training samples as the training samples for the model.

6 . The method of claim 1 , wherein the denoising a set of training samples comprises:

for each training sample in the set of training samples, determining, through the multiple models, a weight of the training sample for calculating a training loss.

7 . The method of claim 6 , wherein the determining a weight of the training sample comprises:

determining the weight according to consistency of multiple prediction results obtained by the multiple models based on the training sample.

8 . The method of claim 1 , further comprising:

further denoising the set of training samples through the trained multiple models.

9 . The method of claim 8 , wherein the further denoising the set of training samples comprises:

updating labels of one or more training samples in the set of training samples through the trained multiple models.

10 . The method of claim 9 , wherein the updating labels comprises, for each training sample in the one or more training samples:

acquiring multiple prediction results obtained by the multiple trained models based on the training sample; and

updating a label of the training sample based on the multiple prediction results.

11 . The method of claim 9 , wherein the one or more training samples come from a translated training corpus and/or a generated training corpus.

12 . The method of claim 1 , wherein the set of training samples corresponds to a training data set, and the training data set is used for performing the denoising and the training for multiple rounds.

13 . The method of claim 1 , wherein:

the set of training samples comes from a training data set, and the training data set is used for performing the denoising and the training for multiple rounds, and

in each round, the set of training samples corresponds to a training data subset among a plurality of training data subsets in the training data set, and the plurality of training data subsets are used for performing the denoising and the training iteratively.

14 . An apparatus for multi-model joint denoising training, comprising:

at least one processor; and

a memory storing computer-executable instructions that, when executed, cause the at least one processor to:

obtain multiple machine learning models having identical neural network architectures, wherein each machine learning model is initialized with a different combination of training corpora selected from a source training corpus containing reliable labels, at least one translated training corpus generated through machine translation, and at least one generated training corpus created through neural language generation;

jointly denoise a set of training samples through the multiple models by:

for each model, select a subset of training samples from the set of training samples by filtering out a predetermined proportion of training samples having highest prediction losses as determined by other models in the multiple models, and

for each training sample, determine a weight for calculating training loss based on consistency of prediction results obtained from the multiple models, wherein inconsistent prediction results indicate noisy training samples;

train the multiple models simultaneously with their respective selected subsets of denoised training samples using the determined weights; and

iteratively update labels of training samples from translated and generated training corpora based on consensus predictions from the trained multiple models.

15 . A computer program product for multi-model joint denoising training, comprising a computer program that is executed by at least one processor for:

obtaining multiple machine learning models having identical neural network architectures, wherein each machine learning model is initialized with a different combination of training corpora selected from a source training corpus containing reliable labels, at least one translated training corpus generated through machine translation, and at least one generated training corpus created through neural language generation;

jointly denoising a set of training samples through the multiple models by:

for each model, selecting a subset of training samples from the set of training samples by filtering out a predetermined proportion of training samples having highest prediction losses as determined by other models in the multiple models, and

for each training sample, determining a weight for calculating training loss based on consistency of prediction results obtained from the multiple models, wherein inconsistent prediction results indicate noisy training samples;

training the multiple models simultaneously with their respective selected subsets of denoised training samples using the determined weights; and

iteratively updating labels of training samples from translated and generated training corpora based on consensus predictions from the trained multiple models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: SHOU, LINJUN; GONG, MING; BAI, XUANYU; WANG, XUGUANG; JIANG, DAXIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064872/0645 →
Priority Claims (1)
CN 202110338761.8 · Mar 30, 2021 · national
Continuity (1)
Related Publication 20240184997A1 · Jun 6, 2024
References Cited (67)
US 20190043516A1 · Germain · 2019 [cited by applicant]
US 20200074997A1 · Jankowski, Jr. · 2020 [cited by examiner]
US 20200226212A1 · Tan · 2020 [cited by applicant]
US 20200311599A1 · Chen · 2020 [cited by applicant]
US 20200334539A1 · Wang · 2020 [cited by applicant]
CN 109784391A · 2019 [cited by applicant]
CN 110276248A · 2019 [cited by applicant]
CN 110598224A · 2019 [cited by applicant]
CN 111340233A · 2020 [cited by applicant]
CN 111859994A · 2020 [cited by applicant]
WO 2004036546A1 · 2004 [cited by applicant]
Office Action Received for Chinese Application No. 202110338761.8, mailed on Sep. 12, 2024, 17 pages (English Translation Provided). [cited by applicant]
Notice of Grant Received for Chinese Application No. 202110338761.8, mailed on Feb. 19, 2024, 4 pages (English Translation Provided). [cited by applicant]
Third Office Action Received for Chinese Application No. 202110338761.8, mailed on Nov. 25, 2024, 14 pages (English Translation Provided). [cited by applicant]
Zhao, et al., “Data Augmentation with Atomic Templates for Spoken Language Understanding”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference … [cited by applicant]
Zoph, et al., “Rethinking Pre-training and Self-training”, In Repository of arXiv:2006.06882v1, Jun. 11, 2020, pp. 1-16. [cited by applicant]
Office Action Received for Chinese Application No. 202110338761.8, mailed on Apr. 26, 2024, 19 pages (English Translation Provided). [cited by applicant]
Anaby-Tavor, et al., “Do Not Have Enough Data? Deep Learning to the Rescue!”, In the Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence, Feb. 7, 2020, pp. 7383-7390. [cited by applicant]
Bari, et al., “XLA: A Robust Unsupervised Data Augmentation Framework for Cross-Lingual NLP”, In Proceedings of The International Conference on Learning Representations, May 4, 2021, pp. 1-29. [cited by applicant]
Bian, et al., “Learning to Match Jobs with Resumes from Sparse Interaction Data using Multi-View Co-Teaching Network”, In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, Oct. … [cited by applicant]
Conneau, et al., “Cross-lingual Language Model Pretraining”, In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 8, 2019, pp. 1-11. [cited by applicant]
Conneau, et al., “Unsupervised Cross-lingual Representation Learning at Scale”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5, 2020, pp. 8440-8451. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: H… [cited by applicant]
Doersch, Carl, “Tutorial on variational autoencoders”, In Repository of arXiv:1606.05908v1, Jun. 19, 2016, pp. 1-22. [cited by applicant]
Dyer, et al., “A Simple, Fast, and Effective Reparameterization of IBM Model 2”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi… [cited by applicant]
Einolghozati, et al., “El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing”, In Repository of arXiv.2101.10524v2, Jan. 27, 2021, 13 Pages. [cited by applicant]
Fan, et al., “Hierarchical Neural Story Generation”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 15, 2018, pp. 889-898. [cited by applicant]
Gao, et al., “Paraphrase Augmented Task-Oriented Dialog Generation”, In Repository of arXiv:2004.07462v1, Apr. 16, 2020, 10 Pages. [cited by applicant]
Goodfellow, et al., “Generative Adversarial Networks”, In the Journal of Communications of the ACM, vol. 63, Issue 11, Nov. 2020, pp. 139-144. [cited by applicant]
Guo, et al., “Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding”, In Repository of arXiv:2109.01583v1, Sep. 3, 2021, 12 Pages. [cited by applicant]
Han, et al., “Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels”, In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Dec. 3, 2018, pp. 1-11. [cited by applicant]
Huang, et al., “Federated Learning for Spoken Language Understanding”, In Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8, 2020, pp. 3467-3478. [cited by applicant]
Huang, et al., “Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International J… [cited by applicant]
Kumar, et al., “Data Augmentation Using Pre-trained Transformer Models”, In Repository of arXiv:2003.02245v1, Mar. 4, 2020, 8 Pages. [cited by applicant]
Lample, et al., “Unsupervised Machine Translation Using Monolingual Corpora Only”, In Repository of arXiv:1711.00043v1, Oct. 31, 2017, pp. 1-12. [cited by applicant]
Li, et al., “DivideMix: Learning with Noisy Labels as Semi-supervised Learning”, In Repository of arXiv:2002.07394v1, Feb. 18, 2020, pp. 1-14. [cited by applicant]
Li, et al., “MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark”, In Repository of arXiv:2008.09335v1, Aug. 21, 2020, 10 Pages. [cited by applicant]
Li, et al., “Unsupervised Cross-lingual Adaptation for Sequence Tagging and Beyond”, In Repository of arXiv:2010.12405v1, Oct. 23, 2020, 14 Pages. [cited by applicant]
Liu, et al., “Attention-Based Recurrent Neural Network Models for Joint Intent Detection and Slot Filling”, In Proceedings of the Interspeech, Sep. 8, 2016, pp. 685-689. [cited by applicant]
Liu, et al., “Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue Systems”, In Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence, Feb. 7, 2020, pp. 843… [cited by applicant]
Liu, et al., “Butterfly: One-step Approach towards Wildly Unsupervised Domain Adaptation”, In Repository of arXiv:1905.07720v3, Feb. 18, 2021, pp. 1-23. [cited by applicant]
Liu, et al., “Cross-lingual Spoken Language Understanding with Regularized Representation Alignment”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 16, 2020, pp. 7241-7251. [cited by applicant]
Liu, et al., “Multilingual Denoising Pre-training for Neural Machine Translation”, In Journal of Transactions of the Association for Computational Linguistics, vol. 8, Jan. 24, 2020, pp. 726-742. [cited by applicant]
Liu, et al., “Unsupervised Anomaly Detection by Robust Collaborative Autoencoders”, In Proceedings of The International Conference on Learning Representations, May 4, 2021, 19 Pages. [cited by applicant]
Liu, et al., “Zero-shot Cross-lingual Dialogue Systems with Transferable Latent Variables”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference… [cited by applicant]
Loshchilov, et al., “Decoupled Weight Decay Regularization”, In Proceedings of the 7th International Conference on Learning Representations, May 6, 2019, 18 Pages. [cited by applicant]
Marivate, et al., “Improving Short Text Classification Through Global Augmentation Methods”, In Proceedings of Machine Learning and Knowledge Extraction: 4th IFIP TC 5, TC 12, WG 8.4, WG 8.9, WG 12.9 International Cross… [cited by applicant]
Mccann, et al., “Learned in Translation: Contextualized Word Vectors”, In Proceedings of the 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, pp. 1-12. [cited by applicant]
Och, et al., “A Systematic Comparison of Various Statistical Alignment Models”, In Journal of Computational Linguistics, vol. 29, Issue 1, Mar. 1, 2003, pp. 19-51. [cited by applicant]
Ott, Myle, “MBART: Multilingual Denoising Pre-training for Neural Machine Translation”, Retrieved From: https://github.com/facebookresearch/fairseq/tree/main/examples/mbart, Jan. 29, 2021, 5 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/019217”, Mailed Date: Jul. 15, 2022, 14 Pages. [cited by applicant]
Peng, et al., “Data Augmentation for Spoken Language Understanding via Pretrained Language Models”, In Repository of arXiv:2004.13952v1, Apr. 29, 2020, 6 Pages. [cited by applicant]
Qin, et al., “CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLP”, In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, 2020 , pp. 3853-3860. [cited by applicant]
Ramshaw, et al., “Text Chunking Using Transformation-Based Learning”, In Book of Natural Language Processing Using Very Large Corpora, 1999, pp. 157-176. [cited by applicant]
Ruder, et al., “A Survey of Cross-lingual Word Embedding Models”, In Journal of Artificial Intelligence Research, vol. 65, Aug. 12, 2019, pp. 569-630. [cited by applicant]
Russo, et al., “Control, Generate, Augment: A Scalable Framework for Multi-Attribute Text Generation”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 16, 2020, pp. 351… [cited by applicant]
Schuster, et al., “Cross-lingual Transfer Learning for Multilingual Task Oriented Dialog”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language … [cited by applicant]
Shakeri, et al., “Multilingual Synthetic Question and Answer Generation for Cross-Lingual Reading Comprehension”, In Repository of arXiv:2010.12008v1, Oct. 22, 2020, 7 Pages. [cited by applicant]
Tanaka, et al., “Data Augmentation Using GANs”, In Repository of arXiv:1904.09135v1, Apr. 19, 2019, pp. 1-16. [cited by applicant]
Tur, et al., “Spoken Language Understanding: Systems for Extracting Semantic Information from Speech”, In Book Spoken Language Understanding: Systems for Extracting Semantic Information from Speech, John Wiley And Sons … [cited by applicant]
Upadhyay, et al., “(Almost) Zero-Shot Cross-Lingual Spoken Language Understanding”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, pp. 6034-6038. [cited by applicant]
Wang, et al., “Spoken Language Understanding”, In Journal of IEEE Signal Processing Magazine, vol. 22, Issue 5, Sep. 26, 2005, pp. 16-31. [cited by applicant]
Wang, et al., “That's So Annoying !!!: A Lexical and Frame-Semantic Embedding Based Data Augmentation Approach to Automatic Categorization of Annoying Behaviors using #petpeeve Tweets”, In Proceedings of the Conference … [cited by applicant]
Wu, et al., “Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Na… [cited by applicant]
Wu, et al., “Conditional BERT Contextual Augmentation”, In Journal of Computational Science, Jun. 8, 2019, pp. 84-95. [cited by applicant]
Xu, et al., “End-to-End Slot Alignment and Recognition for Cross-Lingual NLU”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 16, 2020, pp. 5052-5063. [cited by applicant]
Yang, et al., “Asymmetric Co-Teaching for Unsupervised Cross Domain Person Re-Identification”, In Repository of arXiv:1912.01349v1, Dec. 3, 2019, 9 Pages. [cited by applicant]