Systems and methods of generating and transforming data sets for entity comparison
In an illustrative embodiment, automated systems and methods identify relationships between property features and property quality using customized features sets. Filtered feature sets for properties may be generated, containing feature values representing physical attributes and neighborhood attributes. Machine learning model(s) may be trained with the filtered feature data sets to identify correlations between individual features and relative importance to the feature in determining property quality. The model(s) may be used to identify a set of significant features applicable to performing comparisons between comparable properties.
1 . A system for identifying relationships between property features and property quality, the system comprising:
a non-transitory computer readable medium storing a plurality of feature data sets, each feature data set corresponding to a respective property of a plurality of properties, wherein
each respective feature data set of the plurality of feature data sets comprises a plurality of feature values corresponding to a plurality of features, wherein
the plurality of feature values comprises a cost value and a geographic area value,
a first subset of the plurality of features represents physical attributes of the respective property, and
a second subset of the plurality of features represents neighborhood attributes of a surrounding region of the respective property, and
each given feature data set of at least a portion of the plurality of feature data sets includes one or more feature values of the plurality of feature values representing an unknown value due to lack of information pertaining to a corresponding feature of the plurality of features in relation to the respective property; and
one or more processors configured to perform a plurality of operations, the plurality of operations comprising
identifying, among the plurality of feature data sets, a first at least one respective feature of the plurality of features as having the unknown value in at least a threshold portion of the plurality of feature data sets,
identifying, among the plurality of feature data sets, a second at least one respective feature of the plurality of features as having less than a threshold percentage of variation in a respective feature value of the second at least one respective feature across the plurality of feature data sets,
creating, from the plurality of feature data sets, a plurality of filtered feature data sets, each feature set of the plurality of filtered feature sets comprising a plurality of retained features of the plurality of features by filtering the first at least one respective feature and the second at least one respective feature from the plurality of features of each feature set of the plurality of feature sets,
training a first at least one machine learning model with the plurality of filtered feature data sets, wherein the first at least one machine learning model is configured to identify a first set of correlation weights comprising a respective weight corresponding to each respective feature of the plurality of retained features, wherein the respective weight represents relative importance of the respective feature to determining property quality,
obtaining, from the first at least one machine learning model, the first set of correlation weights,
refining the first set of correlation weights by
i) selecting a subset of the plurality of retained features, each retained feature of the subset of the plurality of retained features having a corresponding correlation weight of the first set of correlation weights above a predetermined threshold value,
ii) filtering the plurality of filtered feature data sets to remove values corresponding to all features excluded from the subset of the plurality of retained features,
iii) training a second at least one machine learning model with the plurality of filtered feature data sets, wherein the second at least one machine learning model is configured to identify a second set of correlation weights comprising a respective weight corresponding to each respective feature of the subset of the plurality of retained features,
iv) obtaining, from the second at least one machine learning model, a refined set of correlation weights, and
v) repeating steps i) through iv) until all correlation weights of the refined set of correlation weights are above the predetermined threshold value,
conducting a model testing phase, comprising
identifying a subject property and a plurality of potentially comparable properties,
for each respective comparable property of the plurality of potentially comparable properties, calculating, using the respective correlation weight of the refined set of correlation weights corresponding to each feature of the subset of the plurality of retained features, a plurality of subject property feature values corresponding to the subset of the plurality of retained features of the subject property, and a respective plurality of comparable property feature values corresponding to the subset of the plurality of retained features of the respective comparable property, at least one respective similarity score between the subject property and the respective comparable property,
based on the respective similarity scores between the subject property and the plurality of potentially comparable properties, selecting at least one comparable property of the plurality of potentially comparable properties for presentation to a user of a remote computing device, and
responsive to presenting the at least one comparable property at the remote computing device, receiving, from the remote computing device, feedback regarding a level of acceptability of one or more comparable properties of the at least one comparable property,
determining, based on the feedback, whether a recommendation accuracy is below a predetermined threshold, and
responsive to the recommendation accuracy being below the predetermined threshold, applying the feedback to retrain the second at least one machine learning model, thereby customizing the second at least one machine learning model according to the model testing phase to improve the recommendation accuracy.
2 . The system of claim 1 , wherein training the first at least one machine learning model comprises:
training a first one or more machine learning models using a first portion of the plurality of filtered feature data sets representing rental properties of the plurality of properties; and
training a second one or more machine learning models using a second portion of the plurality of filtered feature data sets representing cost properties of the plurality of properties.
3 . The system of claim 1 , wherein identifying the subject property comprises receiving, from the remote computing device, a query for property analysis identifying the subject property.
4 . The system of claim 1 , wherein identifying the plurality of potentially comparable properties comprises identifying, from the plurality of properties, the plurality of potentially comparable properties within a threshold geographic distance from the subject property.
5 . The system of claim 1 , wherein calculating the at least one respective similarity score comprises:
calculating, using a portion of the subset of the plurality of retained features representing the physical attributes, a respective physical similarity score of the at least one respective similarity score; and
calculating, using a portion of the subset of the plurality of retained features representing the neighborhood attributes, a respective neighborhood similarity score of the at least one respective similarity score.
6 . The system of claim 1 , wherein the plurality of operations further comprises generating the plurality of feature data sets by:
obtaining, from a plurality of data sources, a plurality of raw data values representing characteristics of the plurality of properties; and
transforming, for at least a portion of the plurality of features, a respective raw data value of each respective property of at least a portion of the plurality of properties to the respective feature value of the respective property.
7 . The system of claim 6 , wherein transforming the respective raw data value for one or more features of the portion of the plurality of features comprises normalizing, across the plurality of properties, at least one expense value, wherein:
the at least one expense value comprises at least one of a rent or a sales price; and
normalizing the at least one expense value comprises deriving an expense per square foot value.
8 . The system of claim 6 , wherein transforming the respective raw data value the respective raw data value for the portion of the plurality of features comprises classifying at least one continuous feature of the portion of the plurality of features to a set of categorical divisions.
9 . The system of claim 8 , wherein the at least one continuous feature comprises one or more of a renovation date or a square footage.
10 . The system of claim 6 , wherein transforming the respective raw data value for the portion of the plurality of features comprises re-classifying at least one category-based feature of the portion of the plurality of features to a reduced set of categorical divisions.
11 . The system of claim 10 , wherein re-classifying the at least one category-based feature comprises performing a correlation analysis to group highly correlated feature values into a same categorical division of the reduced set of categorical divisions.
12 . The system of claim 1 , wherein training the first at least one machine learning model comprises training at least one machine learning model per respective geographic region of a plurality of geographic regions using a respective portion of the plurality of filtered feature data sets having the respective geographic area value corresponding to the respective geographic region.
13 . The system of claim 1 , wherein the plurality of operations further comprises generating a plurality of missing data values, each respective missing data value corresponding to a respective unknown value in a respective feature data set of the plurality of feature data sets.
14 . The system of claim 13 , wherein generating the plurality of missing data values comprises deriving a portion of the plurality of missing data values using, for each respective feature data set of a subset of the plurality of feature data sets, one or more known values of the respective feature data set.
15 . The system of claim 1 , wherein the plurality of operations further comprises illustrating a respective location of each comparable property of the at least one comparable property on an interactive graphical map presented at the remote computing device.
16 . A method for identifying relationships between property features and property quality, the method comprising:
accessing, by processing circuitry, a plurality of feature data sets, wherein
each feature data set corresponds to a respective property of a plurality of properties, wherein
each respective feature data set of the plurality of feature data sets comprises a plurality of feature values corresponding to a plurality of features, wherein
the plurality of feature values comprises a cost value and a geographic area value,
a first subset of the plurality of features represents physical attributes of the respective property, and
a second subset of the plurality of features represents neighborhood attributes of a surrounding region of the respective property, and
each given feature data set of at least a portion of the plurality of feature data sets includes one or more feature values of the plurality of feature values representing an unknown value due to lack of information pertaining to a corresponding feature of the plurality of features in relation to the respective property;
by the processing circuitry, identifying, among the plurality of feature data sets, a first at least one respective feature of the plurality of features as having the unknown value in at least a threshold portion of the plurality of feature data sets;
by the processing circuitry, identifying, among the plurality of feature data sets, a second at least one respective feature of the plurality of features as having less than a threshold percentage of variation in a respective feature value of the second at least one respective feature across the plurality of feature data sets;
creating, by the processing circuitry from the plurality of feature data sets, a plurality of filtered feature data sets, each feature set of the plurality of filtered feature data sets comprising a plurality of retained features of the plurality of features by filtering the first at least one respective feature and the second at least one respective feature from the plurality of features of each feature set of the plurality of feature data sets;
training, by the processing circuitry, a first at least one machine learning model with the plurality of filtered feature data sets, wherein the first at least one machine learning model is configured to identify a first set of correlation weights comprising a respective weight corresponding to each respective feature of the plurality of retained features, wherein the respective weight represents relative importance of the respective feature to determining property quality;
obtaining, from the first at least one machine learning model, the first set of correlation weights;
refining, by the processing circuitry, the first set of correlation weights by
i) selecting a subset of the plurality of retained features, each retained feature of the subset of the plurality of retained features having a corresponding correlation weight of the first set of correlation weights above a predetermined threshold value,
ii) filtering the plurality of filtered feature data sets to remove values corresponding to all features excluded from the subset of the plurality of retained features,
iii) training a second at least one machine learning model with the plurality of filtered feature data sets, wherein the second at least one machine learning model is configured to identify a second set of correlation weights comprising a respective weight corresponding to each respective feature of the subset of the plurality of retained features,
iv) obtaining, from the second at least one machine learning model, a refined set of correlation weights, and
v) repeating steps i) through iv) until all correlation weights of the refined set of correlation weights are above the predetermined threshold value;
conducting, by the processing circuitry, a model testing phase, comprising
identifying a subject property and a plurality of potentially comparable properties
for each respective comparable property of the plurality of potentially comparable properties, calculating, using the respective correlation weight of the refined set of correlation weights corresponding to each feature of the subset of the plurality of retained features, a plurality of subject property feature values corresponding to the subset of the plurality of retained features of the subject property, and a respective plurality of comparable property feature values corresponding to the subset of the plurality of retained features of the respective comparable property, at least one respective similarity score between the subject property and the respective comparable property,
based on the respective similarity scores between the subject property and the plurality of potentially comparable properties, selecting at least one comparable property of the plurality of potentially comparable properties for presentation to a user of a remote computing device, and
responsive to presenting the at least one comparable property at the remote computing device, receiving, from the remote computing device, feedback regarding a level of acceptability of one or more comparable properties of the at least one comparable property;
determining, by the processing circuitry based on the feedback, whether a recommendation accuracy is below a predetermined threshold; and
responsive to the recommendation accuracy being below the predetermined threshold, applying, by the processing circuitry, the feedback to retrain the second at least one machine learning model, thereby customizing the second at least one machine learning model according to the model testing phase to improve the recommendation accuracy.
17 . The method of claim 16 , wherein training the first at least one machine learning model comprises:
training a first one or more machine learning models using a first portion of the plurality of filtered feature data sets representing rental properties of the plurality of properties; and
training a second one or more machine learning models using a second portion of the plurality of filtered feature data sets representing cost properties of the plurality of properties.
18 . The method of claim 16 , wherein calculating the at least one respective similarity score comprises:
calculating, using a portion of the subset of the plurality of retained features representing the physical attributes, a respective physical similarity score of the at least one respective similarity score; and
calculating, using a portion of the subset of the plurality of retained features representing the neighborhood attributes, a respective neighborhood similarity score of the at least one respective similarity score.
19 . The method of claim 16 , further comprising generating, by the processing circuitry, the plurality of feature data sets by:
obtaining, from a plurality of data sources, a plurality of raw data values representing characteristics of the plurality of properties; and
transforming, for at least a portion of the plurality of features, a respective raw data value of each respective property of at least a portion of the plurality of properties to the respective feature value of the respective property.
20 . The method of claim 16 , further comprising generating, by the processing circuitry, a plurality of missing data values, each respective missing data value corresponding to a respective unknown value in a respective feature data set of the plurality of feature data sets.