Positionless encoding for classification and analysis
A system may receive a request to determine an information item associated with a plurality of media received from a requesting system. The plurality of media may be unordered positionless media. The system may generate a multi-media embedding comprising information from each media item of the plurality of media. The system may determine the information item based on the multi-media embedding using a machine learning model configured to process information from unordered positionless media. The machine learning model may be acausal and may be prevented from applying an ordering or position to a portion of the multi-media embedding.
1 . A system comprising:
a non-transitory computer-readable memory that stores computer-executable instructions;
a multi-media encoder that generates a multi-media embedding based on a plurality of media associated with an object; and
one or more processors in communication with the computer-readable memory, wherein the computer-executable instructions, when executed by the one or more processors, causes the one or more processors to at least:
retrieve training information comprising the plurality of media and a plurality of text information associated with the plurality of media, wherein the text information indicates an object information item, wherein the plurality of media is unordered, and wherein the plurality of media lacks position information;
generate, by the multi-media encoder, a multi-media embedding based on at least a portion of the plurality of media, wherein the multi-media encoder comprises a plurality of image encoders that each generate a respective media embedding of a plurality of media embeddings in parallel, and wherein the multi-media encoder combines the generated media embeddings to generate the multi-media embedding;
provide the multi-media embedding as input to a machine learning model configured to determine object information based on the multi-media embedding, wherein the machine learning model is configured to maintain information of the multi-media embedding as positionless, and wherein the machine learning model is configured to maintain information of the multi-media embedding as unordered;
compare the determined object information to a text information of the plurality of text information associated with the portion of the plurality of media;
based on the comparison of the determined object information to the text information, adjust a parameter of the machine learning model;
store the machine learning model.
2 . The system of claim 1 , wherein the machine learning model is a transformer-based machine learning model.
3 . The system of claim 1 , wherein the plurality of media comprises at least one of: textual media, video media, audio media, or image media.
4 . The system of claim 1 , wherein the object is a property.
5 . The system of claim 1 , wherein the computer-executable instructions, when executed, further cause the one or more processors to:
receive, from a requesting system, a request to determine a second object information associated with a second object;
retrieve a second plurality of media associated with the second object, wherein the second plurality of media is unordered, wherein the second plurality of media lacks position information, and wherein the second plurality of media is associated with the second object;
generate, by the multi-media encoder, a second multi-media embedding based on at least a portion of the second plurality of media;
determine the second object information associated with the second object based on the second multi-media embedding using the machine learning model to generate a response, wherein the response comprises the second object information; and
provide the response to the requesting system.
6 . The system of claim 1 , wherein to retrieve the training information, the computer-executable instructions, when executed, further cause the one or more processors to:
transmit a request for the plurality of media and the plurality of text information to a content source; and
receive the plurality of media and the plurality of text information from the content source.
7 . The system of claim 1 , wherein the machine learning model is a neural network machine learning model, wherein the machine learning model comprises a first node, wherein the machine learning model comprises a second node, and wherein to adjust the parameter of the machine learning model the computer-executable instructions, when executed by the one or more processors, causes the one or more processors to update a weight of a connection between the first node and the second node.
8 . A method comprising:
receiving, from a requesting system, a request to determine a property information associated with a property and a plurality of media, wherein the plurality of media is associated with the property, wherein the plurality of media lacks position information, and wherein the plurality of media is unordered;
generating, by a multi-media encoder, a multi-media embedding based on the plurality of media, wherein the multi-media encoder comprises a plurality of image encoders that each generate a respective media embedding of a plurality of media embeddings in parallel, and wherein the multi-media encoder combines the generated media embeddings to generate the multi-media embedding;
determining the property information associated with the property using a machine learning model configured to generate a response based on the multi-media embedding, wherein the multi-media embedding is generated based on unordered positionless media information, wherein the machine learning model operates on acausal information, and wherein the response comprises the determined property information;
providing the response to the requesting system.
9 . The method of claim 8 , wherein the property information is at least one of: Uniform Appraisal Dataset condition rating, a number of rooms, a housing style, a storm damage assessment, a determination of a condition of a foundation, a roof type, an approximate square footage, a number of stories of a building, a quality score, a property rating metric used by one or more businesses, a foundation type, an approximate age of a building, or a facing type.
10 . The method of claim 8 , wherein the plurality of media is associated with a media type, and wherein the media type is one of: video, audio, image, or text.
11 . The method of claim 8 , further comprising:
retrieving training information comprising a second plurality of media and a plurality of text information associated with the second plurality of media, wherein the text information indicates a training property information item, wherein the second plurality of media is unordered, and wherein the second plurality of media lacks position information;
generating, by the multi-media encoder, a second multi-media embedding based on at least a portion of the second plurality of media;
providing the second multi-media embedding as input to the machine learning model to generate a second response, the second response comprising second determined property information;
comparing the second determined property information to a text information of the plurality of text information associated with the portion of the plurality of media;
based on the comparison of the second determined property information to the text information, adjusting a parameter of the machine learning model; and
storing the machine learning model.
12 . The method of claim 11 , wherein the machine learning model is a neural network machine learning model, wherein the machine learning model comprises a first node, wherein the machine learning model comprises a second node, and wherein adjusting the parameter of the machine learning model comprises adjusting a weight of a connection between the first node and the second node.
13 . The method of claim 8 , further comprising:
associating the determined property information with the property; and
storing the determined property information and the association between the determined property information and the property.
14 . The method of claim 8 , wherein the multi-media encoder comprises a plurality of media encoders, and wherein the multi-media embedding is generated based at least in part on a plurality of outputs of the plurality of media encoders.
15 . A non-transitory, computer-readable medium encoded with computer-executable instructions executable by a processor of a computing device, wherein the computer-executable instructions, when executed by the processor, cause the computing device to:
receive, from a requesting system, a request to determine a property information associated with a property;
generate, by a multi-image encoder, the multi-image embedding based on a plurality of images associated with the property, wherein the multi-image encoder comprises a plurality of image encoders, wherein the plurality of image encoders generate a plurality of image embeddings at least partly in parallel, and wherein the multi-image encoder is configured to generate the multi-image embedding based on the plurality of image embeddings;
determine the property information associated with the property based on a multi-image embedding using a machine learning model configured to generate a response based on the multi-image embedding, wherein the multi-image embedding comprises unordered positionless image information, wherein the machine learning model is configured to operate on acausal information, and wherein the response comprises the property information; and
provide the response to the requesting system.
16 . The non-transitory, computer-readable medium of claim 15 , wherein the request comprises a sample image representative of the property information, and wherein the response comprises a property comprising a similar property information associated with the property information.
17 . The non-transitory, computer-readable medium of claim 15 , wherein the response further comprises an indication of a confidence in the determined property information.
18 . The non-transitory, computer-readable medium of claim 15 , wherein the property information is at least one of: Uniform Appraisal Dataset condition rating, a number of rooms, a housing style, a storm damage assessment, a determination of a condition of a foundation, a roof type, an approximate square footage, a number of stories of a building, a quality score, a property rating metric used by one or more businesses, a foundation type, an approximate age of a building, or a facing type.