Deep model integration techniques for machine learning entity interpretation
Various embodiments of the present disclosure provide machine learning training techniques for implementing a multi-modal interpretation process to generate holistic outputs for an event. The techniques may include generating, using first layers of a multi-modal machine learning model, text-based intermediate representations for an entity based on textual input data. The techniques include generating, using second layers of the multi-modal machine learning model, image-based intermediate representations for the entity based on the text-based intermediate representations and input images for the entity. The techniques include generating, using one or more third layers of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representations and an image narrative summary for the input images. The techniques include initiating the performance of a prediction-based action based on the entity representation summary.
1 . A computer-implemented method comprising:
generating, by one or more processors and using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;
generating, by the one or more processors and using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images;
generating, by the one or more processors and using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and
initiating, by the one or more processors, performance of a prediction-based action based on the entity representation summary.
2 . The computer-implemented method of claim 1 , wherein:
the text-based intermediate representation comprises one or more of an initial structured textual output or an initial weight matrix for the entity, and
the image-based intermediate representation comprises one or more of an augmented structured textual output or an augmented weight matrix for the entity.
3 . The computer-implemented method of claim 2 , wherein the entity representation summary comprises a third structured textual output that is based on the augmented structured textual output and the augmented weight matrix.
4 . The computer-implemented method of claim 1 , wherein the first layer, the second layer, and the third layer of the multi-modal machine learning model are trained end-to-end using a labeled training dataset.
5 . The computer-implemented method of claim 1 , further comprising:
receiving a user summary for the plurality of input images; and
generating, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.
6 . The computer-implemented method of claim 5 , further comprising:
generating a plurality of performance insights based on a comparison between the entity representation summary and the user summary.
7 . The computer-implemented method of claim 6 , wherein the plurality of performance insights are indicative of a confidence score for the user summary or the entity representation summary.
8 . The computer-implemented method of claim 6 , wherein initiating the performance of the prediction-based action based on the entity representation summary comprises:
generating a performance alert based on the plurality of performance insights; and
providing the performance alert to a user associated with the user summary.
9 . The computer-implemented method of claim 1 , further comprising:
augmenting the textual input data with the entity representation summary.
10 . A system comprising:
one or more processors; and
one or more memories storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
generate, using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;
generate, using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images for the entity;
generate, using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and
initiate performance of a prediction-based action based on the entity representation summary.
11 . The system of claim 10 , wherein:
the text-based intermediate representation comprises one or more of an initial structured textual output or an initial weight matrix for the entity, and
the image-based intermediate representation comprises one or more of an augmented structured textual output or an augmented weight matrix for the entity.
12 . The system of claim 11 , wherein the entity representation summary comprises a third structured textual output that is based on the augmented structured textual output and the augmented weight matrix.
13 . The system of claim 10 , wherein the first layer, the second layer, and the third layer of the multi-modal machine learning model are trained end-to-end using a labeled training dataset.
14 . The system of claim 10 , wherein the one or more processors are further configured to:
receive a user summary for the plurality of input images; and
generate, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.
15 . The system of claim 14 , further comprising:
generating a plurality of performance insights based on a comparison between the entity representation summary and the user summary.
16 . The system of claim 15 , wherein initiating the performance of the prediction-based action based on the entity representation summary comprises:
generating a performance alert based on the plurality of performance insights; and
providing the performance alert to a user associated with the user summary.
17 . One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
generating, using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;
generating, using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images;
generating, using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and
initiating performance of a prediction-based action based on the entity representation summary.
18 . The one or more non-transitory computer-readable media of claim 17 , wherein the one or more processors are caused to perform operations further comprising:
receive a user summary for the plurality of input images; and
generate, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.
19 . The one or more non-transitory computer-readable media of claim 18 , wherein the one or more processors are caused to perform operations further comprising:
generate a plurality of performance insights based on a comparison between the entity representation summary and the user summary.
20 . The one or more non-transitory computer-readable media of claim 19 , wherein the plurality of performance insights are indicative of a confidence score for the user summary or the entity representation summary.