IP Library Granted Patent US 12,530,581
Granted Patent B2
US 12,530,581 · App. 18/351,397 · Granted Jan 20, 2026

Contrastive sequence-to-sequence data selector

Inventors: Wei Wang (Sunnyvale, CA); Bowen Liang (Mountain View, CA); Macduff Hughes (Los Gatos, CA); Taro Watanabe (Mountain View, CA); Tetsuji Nakagawa (Tokyo, JP); Alexander Rudnick (Mountain View, CA)
Assignee: Google LLC
G06N20/00G06N7/01H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,581
App. No.
18/351,397
Granted
Jan 20, 2026
Kind
B2
Abstract

A method includes generating a base model by training with a first dataset of data pairs and generating an adapted model by training the base model on a second dataset of data pairs. The method also includes determining a contrastive score for each data pair of a third dataset of data pairs using the base model and the adapted model. The contrastive score is indicative of a probability of quality of the respective data pair. The method also includes training a target model using the data pairs of the third dataset and the contrastive scores.

Claims (34)

1 . A computer-implemented method executed by data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a dataset comprising a plurality of data pairs;

for each respective data pair of the plurality of data pairs:

determining a first probability score for the respective data pair using a first trained machine learning model;

determining a second probability score for the respective data pair using a second trained machine learning model, the second trained machine learning model different from the first trained machine learning model; and

determining, based on the first probability score and the second probability score, a respective contrastive score for the respective data pair, the respective contrastive score indicative of a probability of quality of the respective data pair;

selecting, based on the respective contrastive scores, a batch of data pairs from the dataset, wherein selecting the batch of data pairs comprises:

determining a selection ratio based on a current training step of a plurality of training steps, the selection ratio decreasing after each training step;

determining a batch size based on the selection ratio and a size of the dataset;

selecting a number of data pairs from the dataset that corresponds with the determined batch size;

sorting, based on the respective contrastive scores, the selected data pairs; and

removing, from the batch, a quantity of the selected data pairs with the lowest contrastive scores based on a removal ratio, the removal ratio an inverse of the selection ratio and increasing after each training step; and

at each training step of the plurality of training steps, training, using the sorted batch of data pairs of the batch of data pairs with the removed data pairs with the lowest contrastive scores, a target machine learning model, the target machine learning model different from the first trained machine learning model and the second trained machine learning model.

2 . The method of claim 1 , wherein a first model size of the first trained machine learning model is larger than a second model size of the second trained machine learning model.

3 . The method of claim 1 , wherein the first trained machine learning model is trained on a first dataset and the second trained machine learning model is trained on a second dataset different from the first dataset.

4 . The method of claim 1 , wherein the batch size is equal to a fixed batch size divided by the selection ratio.

5 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a dataset comprising a plurality of data pairs;

for each respective data pair of the plurality of data pairs:

determining a first probability score for the respective data pair using a first trained machine learning model;

determining a second probability score for the respective data pair using a second trained machine learning model, the second trained machine learning model different from the first trained machine learning model; and

determining, based on the first probability score and the second probability score, a respective contrastive score for the respective data pair, the respective contrastive score indicative of a probability of quality of the respective data pair;

selecting, based on the respective contrastive scores, a batch of data pairs from the dataset, wherein selecting the batch of data pairs comprises:

determining a selection ratio based on a current training step of a plurality of training steps, the selection ratio decreasing after each training step;

determining a batch size based on the selection ratio and a size of the dataset;

selecting a number of data pairs from the dataset that corresponds with the determined batch size;

sorting, based on the respective contrastive scores, the selected data pairs; and

removing, from the batch, a quantity of the selected data pairs with the lowest contrastive scores based on a removal ratio, the removal ratio an inverse of the selection ratio and increasing after each training step; and

at each training step of the plurality of training steps, training, using the sorted batch of data pairs of the batch of data pairs with the removed data pairs with the lowest contrastive scores, a target machine learning model, the target machine learning model different from the first trained machine learning model and the second trained machine learning second model.

6 . The system of claim 5 , wherein a first model size of the first trained machine learning model is larger than a second model size of the second trained machine learning model.

7 . The system of claim 5 , wherein the first trained machine learning model is trained on a first dataset and the second trained machine learning model is trained on a second dataset different from the first dataset.

8 . The system of claim 5 , wherein the batch size is equal to a fixed batch size divided by the selection ratio.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2023
From: WANG, WEI; LIANG, BOWEN; HUGHES, MACDUFF; WATANABE, TARO; NAKAGAWA, TETSUJI; RUDNICK, ALEXANDER
To: GOOGLE LLC
Reel/Frame 064233/0083 →
Continuity (3)
Continuation 16376254 · Apr 5, 2019
Provisional Application 62668650 · May 8, 2018
Related Publication 20230359938A1 · Nov 9, 2023
References Cited (31)
US 5768603A · Brown et al. · 1998 [cited by applicant]
US 8122026B1 · Laroco, Jr. et al. · 2012 [cited by applicant]
US 8538946B1 · Thakur et al. · 2013 [cited by applicant]
US 10002188B2 · Misra · 2018 [cited by examiner]
US 10885900B2 · Li et al. · 2021 [cited by applicant]
US 11410029B2 · Fukuda · 2022 [cited by examiner]
US 20030009322A1 · Marcu · 2003 [cited by applicant]
US 20060206308A1 · Moore · 2006 [cited by applicant]
US 20080162111A1 · Bangalore et al. · 2008 [cited by applicant]
US 20080306725A1 · Moore · 2008 [cited by applicant]
US 20090112573A1 · He · 2009 [cited by applicant]
US 20090164206A1 · Zhanyi et al. · 2009 [cited by applicant]
US 20160071526A1 · Wingate · 2016 [cited by examiner]
US 20160078339A1 · Li et al. · 2016 [cited by applicant]
US 20170011738A1 · Senior · 2017 [cited by examiner]
US 20170169015A1 · Huang · 2017 [cited by applicant]
US 20180150609A1 · Kim et al. · 2018 [cited by applicant]
US 20190188584A1 · Rao et al. · 2019 [cited by applicant]
US 20200311572A1 · Baker · 2020 [cited by examiner]
WO 2009120449A1 · 2009 [cited by applicant]
Van Der Wees M, Bisazza A, Monz C. Dynamic data selection for neural machine translation. arXiv preprint arXiv:1708.00712. Aug. 2, 2017. (Year: 2017). [cited by examiner]
Dakwale P, Monz C. Fine-tuning for neural machine translation with limited degradation across in-and out-of-domain data. InProceedings of Machine Translation Summit XVI: Research Track 2017 (pp. 156-169). (Year: 2017). [cited by examiner]
Mandal A, Vergyri D, Wang W, Zheng J, Stolcke A, Tur G, Hakkani-Tur D, Ayan NF. Efficient data selection for machine translation. In2008 IEEE Spoken Language Technology Workshop Dec. 15, 2008 (pp. 261-264). IEEE. (Year:… [cited by examiner]
Tuan LA, Kim JJ, Ng SK. Incorporating trustiness and collective synonym/contrastive evidence into taxonomy construction. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing Sep. 2015… [cited by examiner]
Hinton GE. Training products of experts by minimizing contrastive divergence. Neural computation. Aug. 1, 2002;14(8):1771-800. (Year: 2022). [cited by examiner]
Smith NA, Eisner J. Contrastive estimation: Training log-linear models on unlabeled data. InProceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05) Jun. 2005 (pp. 354-362). (Year… [cited by examiner]
Marlies Van Der Wees et al.: “Dynamic Data Selection for Neural Machine Translati on”, a rxiv .org,Corn Ell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 2, 2017 (Aug. 2, 2017), XP080951… [cited by applicant]
Mandala et al: “Efficient data selection for machine translation”, Spoken Language Technology Workshop, 2008. SLT 2008. IEEE, IEEE, Piscataway, NJ′ USA, Dec. 15, 2008 (Dec. 15, 2008), pp. 261-264, XP031421141,ISBN: 978-… [cited by applicant]
Eetemadi Sauleh et al: “Survey of data-selection methods in statistical machine translation”, Machine Translation, Kluwer Academic Publishers, Dordrecht, N L, vol. 29, No. 3, Dec. 28, 2015 (Dec. 28, 2015), pp. 189-223, … [cited by applicant]
International Search Report and Written Opinion of PCT/US2019/026003 dated Jul. 22, 2019. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 16/376,254, dated Sep. 29, 2022. [cited by applicant]