IP Library › Granted Patent US 12,670,898
Granted Patent B1
US 12,670,898 · App. 18/473,712 · Granted Jun 30, 2026

Training and testing machine learning models using biased samples

Inventors: Yunlong Jiao (Hatfield, GB); Anisha Garg (Seattle, WA); Emine Yilmaz (London, GB); Gabriella Kazai (Bishops Stortford, GB); Liu Yang (Seattle, WA); Wenbo Yan (Redmond, WA); Liane Lewin-Eytan (Binyamina, IL); Prathap Ramachandra (Kirkland, WA); Prasanna Soundararajan (Redmond, WA)
Assignee: Amazon Technologies, Inc.
G10L15/063G10L15/26G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,898
App. No.
18/473,712
Filed
Sep 25, 2023
Granted
Jun 30, 2026
Kind
B1
Art Unit
2655
USPC
704/200
Abstract

Techniques for training and testing machine learning models using biased samples are described. In some examples, a sampled test set is generated by estimating a propensity score for each annotation of the set of annotations, wherein a propensity score quantifies a likelihood of being human generated using the set of data, estimating a confidence score for each annotation of the set of annotations, wherein a confidence score quantifies a confidence in a correctness of the annotation, mapping each annotation of the set of annotations to a multi-dimensional space based at least in part on the propensity score, stratifying, based on the propensity score, the mapped annotations, and sampling each stratum according to a request to generate a sampled test set.

Claims (50)

1 . A computer-implemented method comprising:

receiving audio data to be annotated;

receiving text derived from the audio data;

receiving annotations associated with the text derived from the audio data, wherein at least a proper subset of the annotations is human-generated using the data according to an opt-in preference;

generating a sampled test set by:

estimating, by an audio annotation service implemented as audio annotation service code executed out of memory by one or more processors, a propensity score for each annotation, wherein the propensity score quantifies a likelihood of each annotation being human-generated using the audio data;

estimating a confidence score for each annotation, wherein the confidence score quantifies a confidence in a correctness of the annotation;

mapping each annotation to a multi-dimensional space based at least in part on the propensity score to generate mapped annotations;

binning the mapped annotations; and

sampling each bin of the mapped annotations according to a request to generate the sampled test set; and

training a machine learning (ML) model using the sampled test set.

2 . The computer-implemented method of claim 1 , wherein the request to generate a sampled test set indicates the sampling is to only use the proper subset of the annotations that are human-generated using the data according to an opt-in preference.

3 . The computer-implemented method of claim 1 , wherein the request to generate a sampled test set indicates the sampling is to use the proper subset of the annotations that are human-generated using the data and a proper subset of the annotations that are human-generated without using the data.

4 . A computer-implemented method comprising:

receiving data to be annotated;

receiving annotations associated with the data, wherein at least a proper subset of the annotations is human-generated using the data according to an opt-in preference;

sampling the annotations to generate a sampled test set by:

estimating, by an audio annotation service implemented as audio annotation service code executed out of memory by one or more processors, a propensity score for each annotation, wherein the propensity score quantifies a likelihood of each annotation being human-generated using the data,

estimating a confidence score for each annotation, wherein the confidence score quantifies a confidence in a correctness of the annotation,

mapping each annotation to a multi-dimensional space based at least in part on the propensity score to generate mapped annotations,

stratifying, based on the propensity score, the mapped annotations into strata, and

sampling each stratum of the mapped annotations according to a request to generate the sampled test set; and

performing at least one of evaluating a machine learning (ML) model using the sampled test set or training another ML model using the sampled test set.

5 . The computer-implemented method of claim 4 , wherein the training is performed using a service of a provider network.

6 . The computer-implemented method of claim 5 , wherein the request to generate a sampled test set indicates the sampling is to only use the proper subset of the annotations that are human-generated using the data according to an opt-in preference.

7 . The computer-implemented method of claim 5 , wherein the request to generate a sampled test set indicates the sampling is to use the proper subset of the annotations that are human-generated using the data and a proper subset of the annotations that are human-generated without using the data.

8 . The computer-implemented method of claim 7 , wherein low confidence score annotations from the proper subset of the annotations that are human-generated without using the data are swapped with annotations that are human-generated using the data of a same stratum.

9 . The computer-implemented method of claim 8 , wherein the low confidence score annotations are below a confidence score threshold.

10 . The computer-implemented method of claim 9 , wherein the confidence score threshold is configurable.

11 . The computer-implemented method of claim 5 , wherein evaluating a ML model uses a statistical estimator.

12 . The computer-implemented method of claim 11 , wherein the statistical estimator is to re-weight samples in the sampled test set based at least in part on the propensity scores.

13 . The computer-implemented method of claim 4 , wherein the data to be annotated is audio data.

14 . The computer-implemented method of claim 13 , wherein text is generated from the audio data using at least one automatic speech recognition machine learning model.

15 . A system comprising:

a first one or more electronic devices to implement a data storage service in a multi-tenant provider network; and

a second one or more electronic devices to implement an annotation service in the multi-tenant provider network, the annotation service including instructions that upon execution cause the annotation service to:

receive data to be annotated from the data storage service;

receive annotations associated with the data, wherein at least a proper subset of the annotations is human-generated using the data according to an opt-in preference;

sample the annotations to generate a sampled test set by

estimating, by an audio annotation service implemented as audio annotation service code executed out of memory by one or more processors, a propensity score for each annotation, wherein the propensity score quantifies a likelihood of each annotation being human-generated using the data,

estimating a confidence score for each annotation, wherein the confidence score quantifies a confidence in a correctness of the annotation,

mapping each annotation to a multi-dimensional space based at least in part on the propensity score to generate mapped annotations,

stratifying, based on the propensity score, the mapped annotations into strata, and

sampling each stratum of the mapped annotations according to a request to generate the sampled test set; and

perform at least one of evaluating a machine learning (ML) model using the sampled test set or training another ML model using the sampled test set.

16 . The system of claim 15 , further comprising a model training service to train the ML model using the sampled test set.

17 . The system of claim 15 , further comprising a model hosting service to host a ML model trained using the sampled test set.

18 . The system of claim 15 , wherein the request to generate a sampled test set indicates the sampling is to only use the proper subset of the annotations that are human-generated using the data according to an opt-in preference.

19 . The system of claim 15 , wherein the data to be annotated is audio data.

20 . The system of claim 15 , wherein text is generated from the audio data using at least one automatic speech recognition machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2026
From: JIAO, YUNLONG; GARG, ANISHA; YILMAZ, EMINE; KAZAI, GABRIELLA; YANG, LIU; YAN, WENBO; LEWIN-EYTAN, LIANE; RAMACHANDRA, PRATHAP; SOUNDARARAJAN, PRASANNA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 074798/0790 →
References Cited (28)
US 11195522B1 · Makashir · 2021 [cited by examiner]
US 20140223284A1 · Rankin, Jr. · 2014 [cited by examiner]
US 20160006744A1 · Du · 2016 [cited by examiner]
US 20160162569A1 · Erle · 2016 [cited by examiner]
US 20180293908A1 · Wang · 2018 [cited by examiner]
US 20220013023A1 · Hellman · 2022 [cited by examiner]
US 20220050955A1 · Awadalla · 2022 [cited by examiner]
US 20230067976A1 · Bhusan · 2023 [cited by examiner]
US 20230110027A1 · Bajpayee · 2023 [cited by examiner]
US 20240062745A1 · Smyth · 2024 [cited by examiner]
US 20240070516A1 · Patel · 2024 [cited by examiner]
US 20240126822A1 · Hamilton · 2024 [cited by examiner]
Alam et al., “Domain Adaptation with Adversarial Training and Graph Embeddings”, May 14, 2018, 11 pages. [cited by applicant]
Aslam et al., “A Practical Sampling Strategy for Efficient Retrieval Evaluation”, 2007, pp. 1-10. [cited by applicant]
Autenrieth et al., “A general-purpose statistical method for improved learning under Covariate Shift”, Oct. 20, 2021, 13 pages. [cited by applicant]
Bianca Zadrozny, “Learning and Evaluating Classifiers under Sample Selection Bias”, Proceedings of the 21 st International Conference on Machine Learning, Banff, Canada, 2004, 8 pages. [cited by applicant]
Blitzer et al., “Domain Adaptation with Structural Correspondence Learning”, Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing (EMNLP 2006), pp. 120-128. [cited by applicant]
Charles Elkan, “The Foundations of Cost-Sensitive Learning”, Proceedings of the Seventeenth International Joint Conference on Artificial Intelligence (IJCAI'01), May 2001, 6 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, May 24, 2019, 16 pages. [cited by applicant]
Efron et al., “An Introduction to the Bootstrap”, 1994, 11 pages. [cited by applicant]
Ganin et al., “Domain-Adversarial Training of Neural Networks”, Journal of Machine Learning Research 17 (2016), pp. 1-35. [cited by applicant]
Han et al., “Unsupervised Domain Adaptation of Contextualized Embeddings for Sequence Labeling”, Sep. 5, 2019, 12 pages. [cited by applicant]
Horvitz et al., “A Generalization of Sampling Without Replacement From a Finite Universe”, Journal of the American Statistical Association, vol. 47, No. 260 (Dec. 1952), pp. 663-685. [cited by applicant]
Mehrabi et al., “A Survey on Bias and Fairness in Machine Learning”, Jan. 25, 2022, 34 pages. [cited by applicant]
Pan et al., “Transferrable Prototypical Networks for Unsupervised Domain Adaptation”, 2019, pp. 2239-2247. [cited by applicant]
Rosenbaum et al., “The Central Role of the Propensity Score in Observational Studies for Causal Effects”, vol. 70, No. 1. (Apr. 1983), pp. 41-55. [cited by applicant]
Timothy Miller et al., “Simplified Neural Unsupervised Domain Adaptation”, May 22, 2019, 6 pages. [cited by applicant]
Vu et al., “Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models”, Oct. 5, 2020, 11 pages. [cited by applicant]