Learning unified embedding
A computer-implemented method for generating a unified machine learning model using a neural network on a data processing apparatus is described. The method includes the data processing apparatus determining respective learning targets for each of a plurality of object verticals. The data processing apparatus determines the respective learning targets based on two or more embedding outputs of the neural network. The method also includes the data processing apparatus training the neural network to identify data associated with each of the plurality of object verticals. The data processing apparatus trains the neural network using the respective learning targets and based on a first loss function. The data processing apparatus uses the neural network trained to generate a unified machine learning model, where the model is configured to identify particular data items associated with each of the plurality of object verticals.
1 . A computer-implemented method for generating a unified machine learning computing model using a neural network on a data processing apparatus, the method comprising:
determining, by the data processing apparatus and for the neural network, respective learning targets for each of a plurality of object verticals, wherein each object vertical defines a distinct category for an object that belongs to the vertical, wherein each distinct category comprises a particular object class, wherein each respective learning target comprises an embedding output generated by a respective specialized model of a plurality of specialized models, wherein the plurality of specialized models are individually trained, wherein the respective specialized model is associated with a particular object vertical of the plurality of object verticals, wherein each distinct specialized model of the plurality of specialized models was trained for a different object class;
training, by the data processing apparatus and based on a first loss function, the neural network to identify data associated with each of the plurality of object verticals, wherein the neural network is trained to generate one or more feature embeddings associated with the embedding output of the respective learning targets of a respective specialized model, wherein the first loss function differs from a second loss function, wherein the plurality of specialized models are trained with the second loss function, wherein training with the first loss function comprises:
training the neural network to generate one or more particular feature embeddings that emulate the embedding output of the respective specialized model associated with a particular object vertical associated with a subset of training data being processed by the neural network, wherein the respective specialized model differs based on the particular object vertical and the respective portion of training data; and
generating, by the data processing apparatus and using the neural network trained based on the first loss function, a unified machine learning model configured to identify items that are included in the data associated with each of the plurality of object verticals.
2 . The method of claim 1 , wherein determining respective learning targets for the neural network further comprises:
training, by the data processing apparatus and based on the second loss function, at least one other neural network to identify data associated with each of the plurality of object verticals,
in response to training, generating, by the data processing apparatus, two or more respective embedding outputs, where each respective embedding output indicates a particular learning target and includes a vector of parameters that correspond to the data associated with a particular object vertical; and
generating, by the data processing apparatus and using the at least one other neural network trained based on the second loss function, respective machine learning models, each machine learning model being configured to use a particular embedding output.
3 . The method of claim 2 , wherein determining respective learning targets for the neural network further comprises:
providing, for training the neural network, the respective learning targets generated from respective separate models.
4 . The method of claim 2 , wherein each of the plurality of object verticals corresponds to particular category of items and the data associated with each of the plurality of object verticals includes image data of an item in the particular category of items.
5 . The method of claim 4 , wherein the particular category is an apparel category and items of the particular category include at least one of: handbags, shoes, dresses, pants, or outerwear; and
wherein the image data indicates an image of at least one of: a particular handbag, a particular shoe, a particular dress, a particular pant, or particular outerwear.
6 . The method of claim 5 , wherein:
each of the respective machine learning models are configured to identify data associated with a particular object vertical and within a first degree of accuracy; and
the unified machine learning model is configured to identify data associated with each of the plurality of object verticals and within a second degree of that exceeds the first degree of accuracy.
7 . The method of claim 2 , wherein determining the respective learning targets for each of the plurality of object verticals comprises:
analyzing the two or more respective embedding outputs, each respective embedding output corresponding to a particular object vertical of the plurality of object verticals; and
based on the analyzing, determining the respective learning targets for each of the plurality of object verticals.
8 . The method of claim 2 , wherein the first loss function is an L2-loss function and generating the unified machine learning model includes:
generating a particular unified machine learning model that minimizes a computational output associated with the L2-loss function.
9 . The method of claim 2 , wherein the neural network includes a plurality of neural network layers that receive multiple layer inputs, and where training the neural network based on the first loss function includes:
performing batch normalization to normalize layer inputs to a particular neural network layer; and
minimizing covariate shift in response to performing the batch normalization.
10 . The method of claim 2 , wherein the second loss function is a triplet loss function and generating the respective machine learning models includes:
generating a particular machine learning model based on associations between an anchor image, a positive image and a negative image.
11 . A system for generating a unified machine learning model using a neural network, the system comprising:
a data processing apparatus configured to implement the neural network, the data processing apparatus including one or more processing devices; and
one or more non-transitory machine-readable storage devices storing instructions that are executable by the one or more processing devices to cause performance of operations comprising:
determining, by the data processing apparatus and for the neural network, respective each learning targets for each of a plurality of object verticals, wherein each object vertical defines a distinct category for an object that belongs to the vertical, wherein each distinct category comprises a particular object class, wherein each respective learning target comprises an embedding output generated by a respective specialized model of a plurality of specialized models, wherein the plurality of specialized models are individually trained, wherein the respective specialized model is associated with a particular object vertical of the plurality of object verticals, wherein each distinct specialized model of the plurality of specialized models was trained for a different object class;
training, by the data processing apparatus and based on a first loss function, the neural network to identify data associated with each of the plurality of object verticals, wherein the neural network is trained to generate one or more feature embeddings associated with the embedding output of the respective learning targets generated by a respective specialized model, wherein the first loss function differs from a second loss function, wherein the plurality of specialized models are trained with the second loss function, wherein training with the first loss function comprises:
training the neural network to generate one or more particular feature embeddings that emulate the embedding output of the respective specialized model associated with a particular object vertical associated with a subset of training data being processed by the neural network, wherein the respective specialized model differs based on the particular object vertical and the respective portion of training data; and
generating, by the data processing apparatus and using the neural network trained based on the first loss function, a unified machine learning model configured to identify items that are included in the data associated with each of the plurality of object verticals.
12 . The system of claim 11 , wherein determining respective learning targets for the neural network further comprises:
training, by the data processing apparatus and based on the second loss function, at least one other neural network to identify data associated with each of the plurality of object verticals;
in response to training, generating, by the data processing apparatus, two or more embedding outputs, where each embedding output indicates a particular learning target and includes a vector of parameters that correspond the data associated with particular object vertical; and
generating, by the data processing apparatus and using the at least one other neural network trained based on the second loss function, respective machine learning models, each machine learning mode being configured to use a particular embedding output.
13 . The system of claim 12 , wherein each of the plurality of object verticals corresponds to particular category of items and the data associated with each of the plurality of object verticals includes image data of an item in the particular category of items.
14 . The system of claim 13 , wherein the particular category is an apparel category and items of the particular category include at least one of: handbags, shoes, dresses, pants, or outerwear; and
wherein the image data indicates an image of at least one of: a particular handbag, a particular shoe, a particular dress, a particular pant, or particular outerwear.
15 . The system of claim 14 , wherein:
each of the respective machine learning models are configured to identify data associated with a particular object vertical and within a first degree of accuracy; and
the unified machine learning model is configured to identify data associated with each of the plurality of object verticals and within a second degree of that exceeds the first degree of accuracy.
16 . The system of claim 12 , wherein determining the respective learning targets for each of the plurality of objects verticals, comprises:
analyzing the two or more embedding outputs, each embedding output corresponding to a particular object vertical of the plurality of object verticals; and
based on the analyzing, determining the respective learning targets for each of the plurality of object verticals.
17 . The system of claim 12 , wherein the first loss function is an L2-loss function and generating the unified machine learning model includes:
generating a particular unified machine learning model that minimizes a computational output associated with the L2-loss function.
18 . The system of claim 12 , wherein the neural network includes a plurality of neural network layers that receive multiple layer inputs, and where training the neural network based on the first loss function includes:
performs batch normalization to normalize layer inputs to a particular neural network layer; and
minimizing covariate shift in response to performing the batch normalization.
19 . The system of claim 12 , wherein the second loss function is a triplet loss function and generating the respective machine learning models includes:
generating a particular machine learning model based on association between an anchor image, a positive image, and a negative image.
20 . One or more non-transitory machine-readable storage devices storing instructions that are executable by the one or more processing devices to generate a unified machine learning model and to cause performance of operations comprising:
determining, by a data processing apparatus configured to implement a neural network, respective learning targets for each of a plurality of object verticals, wherein each object vertical defines a distinct category for an object that belongs to the vertical, wherein each object vertical comprises a different type of apparel item of a plurality of various types of apparel items, wherein each respective learning target comprises an embedding output generated by a respective specialized model of a plurality of specialized models, wherein the plurality of specialized models are individually trained, wherein the respective specialized model is associated with a particular object vertical of the plurality of object verticals, wherein each distinct specialized model of the plurality of specialized models was trained for a different object class, wherein determining respective learning targets for each of a plurality of object verticals comprises:
processing a first set of training data with a first specialized model associated with a first object class to generate a first set of embedding outputs; and
processing a second set of training data with a second specialized model associated with a second object class to generate a second set of embedding outputs;
training, by the data processing apparatus and based on a first loss function, the neural network to identify data associated with each of the plurality of object verticals, wherein the neural network is trained to generate one or more feature embeddings associated with the embedding output of the respective learning targets of a respective specialized model, wherein the first loss function differs from a second loss function, wherein the plurality of specialized models are trained with the second loss function, wherein the first loss function comprises an L2 loss function, and wherein the second loss function comprises a triplet loss function, wherein training the neural network comprises:
training the neural network to generate one or more first feature embeddings to emulate the first set of embedding outputs based on the first set of embedding outputs being the respective learning target when processing the first set of training data; and
training the neural network to generate one or more second feature embeddings to emulate the second set of embedding outputs based on the second set of embedding outputs being the respective learning target when processing example from the second set of training data; and
generating, by the data processing apparatus and using the neural network trained based on the first loss function, a unified machine learning model configured to identify items that are included in the data associated with each of the plurality of object verticals.