Training data synthesis for machine learning
A method can include generating a plurality of synthetic objects and associated labels using a trained first machine learning system that is trained to generate a synthetic object based at least in part on a feature of a labeled object, an assigned label that represents the feature, and stochastic variation input; training a second machine learning model to predict labels for features of objects based at least in part on the plurality of synthetic objects and associated labels; and predicting a label for an unlabeled feature of an object using the second machine learning model.
1 . A method comprising:
generating a plurality of synthetic objects and associated labels using a trained first machine learning system that is trained to generate a synthetic object based at least in part on a feature of a labeled object, an assigned label that represents the feature, and stochastic variation input, wherein the synthetic objects include a stochastic variation output, wherein the stochastic variation input comprises single channel images including uncorrelated Gaussian noise, and wherein the stochastic variation output comprises one or more of:
erased or partially erased gridlines;
width and intensity variations in the gridlines;
noise on the gridlines;
intensity variation in curves; and
width variation in the curves;
training a second machine learning model to predict labels for features of objects based at least in part on the plurality of synthetic objects and associated labels; and
predicting a label for an unlabeled feature of an object using the second machine learning model.
2 . The method of claim 1 , further comprising:
determining an uncertainty associated with labeling the unlabeled feature in the object using the trained second machine learning model;
determining that the uncertainty is greater than a predetermined value;
in response to determining that the uncertainty is greater than the predetermined value, soliciting an input of one or more training pairs of objects having the unlabeled feature of the object and an assigned label associated therewith;
generating a new plurality of synthetic objects and associated labels using the trained first machine learning system and the one or more training pairs of objects; and
training the second machine learning model to predict labels for features of objects based at least in part on the plurality of new synthetic objects and associated labels.
3 . The method of claim 1 , wherein the labeled object comprises a well log, wherein the feature comprises one or more of a header section, a depth track, and a plot segment, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the header section, the depth track, and the plot segment based on the labeled object.
4 . The method of claim 3 , wherein the variations include different relative locations for one or more of the header section, the depth track, and the plot segment in the individual synthetic objects.
5 . The method of claim 1 , wherein the labeled object comprises a plot of a well log or a seismic survey log, wherein the feature comprises one or more of a curve shape, a number of curves, a range of values, and a line style, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the curve shape, the number of curves, the range of values, and the line style.
6 . The method of claim 1 , wherein the labeled object comprises a header section of a well log, wherein the feature comprises a line style, units, or a scale in the header section, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the line style, the units, or the scale.
7 . The method of claim 6 , wherein the variations include different relative locations for display of the line style, the units, or the scale in the individual synthetic objects.
8 . The method of claim 1 , wherein the labeled object comprises a natural language search query, wherein the feature comprises one or more of a country, a state, an operator identity, and a field need, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the country, the state, the operator, and the field need, and wherein training the second machine learning model comprises training the second machine learning model to label natural language search queries as database-specific language search queries.
9 . The method of claim 1 , wherein the first machine learning system comprises a generative adversarial network.
10 . The method of claim 9 , wherein the generative adversarial network comprises a generator and a discriminator.
11 . A non-transitory, computer-readable medium storing instructions that, when executed by at least one processor of a computing system, cause the computing system to perform operations, the operations comprising:
generating a plurality of synthetic objects and associated labels using a trained first machine learning system that is trained to generate a synthetic object based at least in part on a feature of a labeled object, an assigned label that represents the feature, and stochastic variation input, wherein the synthetic objects include a stochastic variation output, wherein the stochastic variation input comprises single channel images including uncorrelated Gaussian noise, and wherein the stochastic variation output comprises one or more of:
erased or partially erased gridlines;
width and intensity variations in the gridlines;
noise on the gridlines;
intensity variation in curves; and
width variation in the curves;
training a second machine learning model to predict labels for features of objects based at least in part on the plurality of synthetic objects and associated labels; and
predicting a label for an unlabeled feature of an object using the second machine learning model.
12 . The medium of claim 11 , wherein the operations further comprise:
determining an uncertainty associated with labeling the unlabeled feature in the object using the trained second machine learning model;
determining that the uncertainty is greater than a predetermined value;
in response to determining that the uncertainty is greater than the predetermined value, soliciting an input of one or more training pairs of objects having the feature of the unlabeled object and an assigned label associated therewith;
generating a new plurality of synthetic objects and associated labels using the trained first machine learning system and the one or more training pairs of objects; and
training the second machine learning model to predict labels for features of objects based at least in part on the plurality of new synthetic objects and associated labels.
13 . The medium of claim 11 , wherein the labeled object comprises a well log, wherein the feature comprises one or more of a header section, a depth track, and a plot segment, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the header section, the depth track, and the plot segment based on the labeled object.
14 . The medium of claim 11 , wherein the labeled object comprises a plot of a well log or a seismic survey log, wherein the feature comprises one or more of a curve shape, a number of curves, a range of values, and a line style, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the curve shape, the number of curves, the range of values, and the line style.
15 . The medium of claim 11 , wherein the labeled object comprises a header section of a well log, wherein the feature comprises a line style, units, or a scale in the header section, wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the line style, the units, and the scale, and wherein the variations include different relative locations for display of the line style, the units, and the scale in the individual synthetic objects.
16 . The medium of claim 11 , wherein the labeled object comprises a natural language search query, wherein the feature comprises one or more of a country, a state, an operator identity, and a field need, and wherein training the first machine learning model comprises training the first machine learning model to generate synthetic objects having variations of the one or more of the country, the state, the operator, and the field need, and wherein training the second machine learning model comprises training the second machine learning model to label natural language search queries as database-specific language search queries.
17 . A computing system, comprising:
one or more processors; and
a memory system including one or more non-transitory computer-readable media storing instructions that, when executed by at least one of the one or more processors, cause the computing system to perform operations, the operations comprising:
generating a plurality of synthetic objects and associated labels using a trained first machine learning system that is trained to generate a synthetic object based at least in part on a feature of a labeled object, an assigned label that represents the feature, and stochastic variation input, wherein the synthetic objects include a stochastic variation output, wherein the stochastic variation input comprises single channel images including uncorrelated Gaussian noise, and wherein the stochastic variation output comprises one or more of:
erased or partially erased gridlines;
width and intensity variations in the gridlines;
noise on the gridlines;
intensity variation in curves; and
width variation in the curves;
training a second machine learning model to predict labels for features of objects based at least in part on the plurality of synthetic objects and associated labels; and
predicting a label for an unlabeled feature of an object using the second machine learning model.
18 . The computing system of claim 17 , wherein the operations further comprise:
determining an uncertainty associated with labeling the unlabeled feature in the object using the trained second machine learning model;
determining that the uncertainty is greater than a predetermined value;
in response to determining that the uncertainty is greater than the predetermined value, soliciting an input of one or more training pairs of objects having the unlabeled feature of the object and an assigned label associated therewith;
generating a new plurality of synthetic objects and associated labels using the trained first machine learning system and the one or more training pairs of objects; and
training the second machine learning model to predict labels for features of objects based at least in part on the plurality of new synthetic objects and associated labels.