IP Library Granted Patent US 12670701
Granted Patent B1
US 12670701 · App. 19/397,798 · Granted Jun 30, 2026

Systems and methods for labeling training data for information extraction systems

Inventors: Lei Zhang (Rego Park, NY); Chuanni He (Chamblee, GA)
Assignee: American International Group, Inc.
G06V10/7753G06V30/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670701
App. No.
19/397,798
Granted
Jun 30, 2026
Kind
B1
Abstract

A system for extracting a number of data elements from one or more data sources. The system increases the size of training examples that can be used to test, score, and generate the extraction procedure by generating additional training examples. The additional training examples are generated by automatically labeling unlabeled examples and augmented the labeled training examples with the unlabeled examples for which a ground truth value has been estimated. The system queries a number of language models to extract the information from the unlabeled examples and uses an algorithm to estimate the ground truth value from the values estimated by the ensemble of language models. A flag is also generated indicating those unlabeled examples of particular difficulty which may have high uncertainty and require supplemental validation of the estimated ground truth value. The system can populate an ontological data store using the extraction procedure developed using the additional training examples.

Claims (57)

1 . A method for labeling ground truth values for data fields within submission content, the method comprising:

prompting, by one or more processors, a plurality of language models to extract a value for a data field from the submission content for a plurality of labeled submissions having a ground truth value for the data field;

generating, by the one or more processors, a data structure comprising an entry for each combination of a submission from the plurality of labeled submissions and a language model of the plurality of language models, the entry indicating whether the value for the data field extracted by the language model from the submission content of the submission matches the ground truth value corresponding to the data field for the submission;

determining, by the one or more processors, parameters for a labeling model configured to generate an estimated ground truth value and an uncertainty metric for the data field based on inputs comprising the value extracted by each of the plurality of language models;

generating, by the one or more processors, the estimated ground truth value and the uncertainty metric for an unlabeled submission by applying the labeling model to values for the data field extracted by each language model of the plurality of language models from the submission content for the unlabeled submission; and

responsive to the uncertainty metric for the unlabeled submission satisfying an uncertainty threshold, generating, by the one or more processors, an indication to validate the estimated ground truth value for the unlabeled submission.

2 . The method of claim 1 , wherein the unlabeled submission is one of a plurality of unlabeled training submissions used to determine an extraction prompt to extract the value for the data field.

3 . The method of claim 2 , further comprising:

generating, by the one or more processors, the estimated ground truth value for each training submission of the plurality of unlabeled training submissions;

generating, by the one or more processors, an extraction training set comprising the estimated ground truth value corresponding to each training submission of the plurality of unlabeled training submissions, the plurality of unlabeled training submissions, the plurality of labeled submissions, and the ground truth value corresponding to each of the plurality of labeled submissions; and

determining, by the one or more processors, a score indicating a fraction of submissions of the extraction training set for which the extraction prompt causes an extraction language model to extract the value for the submission of the extraction training set matching the ground truth value corresponding to the submission of the extraction training set.

4 . The method of claim 3 , further comprising adjusting, by the one or more processors, the extraction prompt to improve the score.

5 . The method of claim 3 , wherein the extraction language model is not included in the plurality of language models.

6 . The method of claim 2 , wherein:

the extraction prompt is a first extraction prompt; and

prompting, by the one or more processors, the plurality of language models to extract the value for the data field from the submission content for the plurality of labeled submissions uses a second extraction prompt different than the first extraction prompt.

7 . The method of claim 1 , wherein:

the labeling model is configured to generate the estimated ground truth value based on a weighted voting algorithm; and

determining the parameters for the labeling model comprises determining a weight for each language model of the plurality of language models in the weighted voting algorithm.

8 . The method of claim 1 , wherein the labeling model is configured to generate the uncertainty metric based on a number of language models from the plurality of language models.

9 . The method of claim 1 , wherein the uncertainty metric for the unlabeled submission is either a Gini impurity or an entropy based on a distribution of the values extracted by the plurality of language models for the unlabeled submission.

10 . The method of claim 1 , wherein:

the estimated ground truth value is a value extracted by a largest number of language models of the plurality of language models; and

the uncertainty metric is related to the largest number of language models that extracted the estimated ground truth value.

11 . A system for labeling ground truth values for data fields within submission content, the system comprising one or more processing circuits configured to:

prompt a plurality of language models to extract a value for a data field from the submission content for a plurality of labeled submissions having a ground truth value for the data field;

generate a data structure comprising an entry for each combination of a submission from the plurality of labeled submissions and a language model of the plurality of language models, the entry indicating whether the value for the data field extracted by the language model from the submission content of the submission matches the ground truth value corresponding to the data field for the submission;

determine parameters for a labeling model configured to generate an estimated ground truth value and an uncertainty metric for the data field based on inputs comprising the value extracted by each of the plurality of language models;

generate the estimated ground truth value and the uncertainty metric for an unlabeled submission by applying the labeling model to values for the data field extracted by each language model of the plurality of language models from the submission content for the unlabeled submission; and

responsive to the uncertainty metric for the unlabeled submission satisfying an uncertainty threshold, generate an indication to validate the estimated ground truth value for the unlabeled submission.

12 . The system of claim 11 , wherein the unlabeled submission is one of a plurality of unlabeled training submissions used to determine an extraction prompt to extract the value for the data field.

13 . The system of claim 12 , wherein the one or more processing circuits are configured to:

generate the estimated ground truth value for each training submission of the plurality of unlabeled training submissions;

generate an extraction training set comprising the estimated ground truth value corresponding to each training submission of the plurality of unlabeled training submissions, the plurality of unlabeled training submissions, the plurality of labeled submissions, and the ground truth value corresponding to each of the plurality of labeled submissions; and

determine a score indicating a fraction of submissions of the extraction training set for which the extraction prompt causes an extraction language model to extract the value for the submission of the extraction training set matching the ground truth value corresponding to the submission of the extraction training set.

14 . The system of claim 13 , wherein the one or more processing circuits are configured to adjust the extraction prompt to improve the score.

15 . The system of claim 13 , wherein the extraction language model is not included in the plurality of language models.

16 . The system of claim 12 , wherein:

the extraction prompt is a first extraction prompt; and

prompt the plurality of language models to extract the value for the data field from the submission content for the plurality of labeled submissions uses a second extraction prompt different than the first extraction prompt.

17 . The system of claim 11 , wherein:

the labeling model is configured to generate the estimated ground truth value based on a weighted voting algorithm; and

determining the parameters for the labeling model comprises determining a weight for each language model of the plurality of language models in the weighted voting algorithm.

18 . The system of claim 11 , wherein the labeling model is configured to generate the uncertainty metric based on a number of language models from the plurality of language models.

19 . The system of claim 11 , wherein:

the estimated ground truth value is a value extracted by a largest number of language models of the plurality of language models; and

the uncertainty metric is related to the largest number of language models that extracted the estimated ground truth value.

20 . A system for labeling ground truth values for data fields within submission content, the system comprising one or more processing circuits configured to:

prompt a plurality of language models to extract a value for a data field from the submission content for a plurality of labeled submissions having a ground truth value for the data field;

generate a data structure comprising an entry for each combination of a submission from the plurality of labeled submissions and a language model of the plurality of language models, the entry indicating whether the value for the data field extracted by the language model from the submission content of the submission matches the ground truth value corresponding to the data field for the submission;

generate a labeling model configured to, for a corresponding submission, generate:

an estimated ground truth value based on the value for the data field extracted for the corresponding submission by a largest number of language models; and

an uncertainty metric for the corresponding submission based on the largest number of language models;

generate the estimated ground truth value and the uncertainty metric for each training submission of a plurality of unlabeled training submissions using the labeling model;

responsive to the uncertainty metric for a training submission satisfying an uncertainty threshold, generate an indication to validate the estimated ground truth value for the training submission;

generate an extraction training set comprising the estimated ground truth value corresponding to each training submission of the plurality of unlabeled training submissions, the plurality of unlabeled training submissions, the plurality of labeled submissions, and the ground truth value corresponding to each of the plurality of labeled submissions; and

adjust an extraction prompt to improve a score indicating a fraction of submissions of the extraction training set for which the extraction prompt causes an extraction language model to extract the value for the submission of the extraction training set matching the ground truth value corresponding to the submission of the extraction training set.