Computer processes for clustering properties into neighborhoods and generating neighborhood-specific models
A computer system and associated processes are disclosed for grouping similar real estate properties into contiguous neighborhoods, and for generating neighborhood-specific models capable of estimating property values within their respective neighborhoods. A clustering component uses various sources of property-level data to group properties based on measures of property similarity. For example, the clustering component may measure property similarity based on how frequently specific properties are designated as comparable in appraisal reports. A model generator uses a machine learning process to determine, for specific neighborhoods, correlations between property attributes and values, and uses these correlations to generate the neighborhood specific models.
1 . A computing system comprising one or more hardware processors to execute program code stored in computer memory, the one or more hardware processors programmed to implement the program code to execute a process that comprises:
determining a plurality of blocks within a geographic region, each block comprising a plurality of residential properties;
forming links between pairs of the blocks based on measures of similarity between particular residential properties in the blocks;
aggregating the links to generate a normalized score for each pair of the blocks representing a degree of similarity between the blocks;
applying a clustering algorithm to the blocks to form neighborhoods, said clustering algorithm constrained such that each neighborhood has a contiguous boundary, wherein the clustering algorithm merges neighboring blocks based at least partly on numbers of said links formed between the blocks and the normalized scores;
assigning a unique identifier to each neighborhood; and
generating, for a neighborhood identified by the unique identifier and formed by application of the clustering algorithm, a neighborhood-specific model configured to estimate values of properties in the neighborhood.
2 . The computing system of claim 1 , wherein forming links between pairs of the blocks comprises forming a link between a first block and a second block based at least partly on a residential property in the first block being designated in an appraisal report as comparable to a residential property in the second block.
3 . The computing system of claim 2 , wherein the first block is a first multi-residence development and the second block is a second multi-residence development.
4 . The computing system of claim 1 , wherein at least some of the blocks are census tracks.
5 . The computing system of claim 1 , wherein applying the clustering algorithm comprises:
generating, for each of a plurality of pairs of neighboring blocks, a normalized score representing a degree of connectivity or similarity between the neighboring blocks of the respective pair, said normalized score eliminating an effect of block size; and
determining whether to merge particular neighboring blocks based at least partly on the normalized scores.
6 . The computing system of claim 1 , wherein the process comprises applying the clustering algorithm iteratively until a stopping criterion is met.
7 . The computing system of claim 6 , wherein applying the clustering algorithm iteratively comprises merging a first block with a second block to form a third block based at least partly on a number of links between the first and second blocks, and thereafter merging the third block with a fourth block based at least partly on a number of links between the third and fourth blocks.
8 . The computing system of claim 6 , wherein the stopping criterion comprises a comparison of strengths of linkages within a given block to strengths of linkages between the given block and other blocks.
9 . The computing system of claim 1 , wherein the clustering algorithm determines whether a first block and a second block are neighbors by expanding geographic boundaries of the first and second blocks while determining whether the first and second blocks touch each other without first touching any other blocks.
10 . The computing system of claim 1 , wherein the process further comprises generating, for a neighborhood formed by application of the clustering algorithm, a neighborhood-specific model configured to estimate values of properties in the neighborhood.
11 . The computing system of claim 10 , wherein generating the neighborhood-specific model comprises using a machine learning process to determine, for the neighborhood, correlations between property attributes and property values in the neighborhood, and generating weight values specifying amounts of weight to give to specific property attributes in estimating property values in the neighborhood.
12 . A computing system comprising one or more computing devices programmed with executable program code stored in computer memory, the computing system programmed to implement a process that comprises:
calculating measures of similarity between real estate properties, said measures of similarity based at least partly on designations in appraisal reports of particular real estate properties as comparable; and
forming links between pairs of the real estate properties based on the measures of similarity;
aggregating the links to generate a normalized score between each pair of real estate properties representing a degree of similarity between the real estate properties;
grouping the real estate properties into neighborhoods based on the normalized scores such that each neighborhood has a contiguous geographic boundary, wherein grouping the real estate properties into neighborhoods comprises executing a clustering algorithm; and
generating, for each neighborhood formed by application of the clustering algorithm, a neighborhood-specific model configured to estimate values of properties in the neighborhood.
13 . The computing system of claim 12 , wherein the process comprises generating, for a property pair composed of a first real estate property and a second real estate property, an occurrence count value representing a number of times the first and second real estate properties have been designated as comparable to each other in the appraisal reports, and using said occurrence count value as a factor for measuring property similarity of the first and second real estate properties.
14 . The computing system of claim 12 , wherein the process comprises measuring similarity between first and second neighbor blocks, each of which comprises a plurality of real estate properties, based at least in part on a number of times a real estate property in the first block is designated in an appraisal report as comparable to a real estate property in the second block.
15 . The computing system of claim 14 , wherein measuring said similarity comprises generating a normalized score that eliminates an effect of block size.
16 . The computing system of claim 15 , wherein the clustering algorithm is configured to use the normalized score to determine whether to merge the first and second blocks.
17 . The computing system of claim 12 , wherein the process further comprises generating, for a neighborhood formed by application of the clustering algorithm, a neighborhood-specific model configured to estimate values of properties in the neighborhood.
18 . The computing system of claim 17 , wherein generating the neighborhood-specific model comprises using a machine learning process to determine, for the neighborhood, correlations between property attributes and property values in the neighborhood, and generating weight values specifying amounts of weight to give to specific property attributes in estimating property values in the neighborhood.
19 . A computing system comprising one or more computing devices programmed with executable program code stored in computer memory, the computing system programmed to implement a process that comprises:
determining a plurality of blocks within a geographic area, each block comprising a plurality of residential properties;
determining a neighborhood that comprises a subset of the blocks, said neighborhood having a contiguous geographic boundary, wherein determining the neighborhood comprises:
forming links between pairs of the blocks based on measures of similarity between residential properties;
aggregating the links to generate a normalized score for each pair of the blocks representing a degree of similarity between the blocks;
merging neighboring blocks of said plurality of blocks based on the normalized scores; and
determining correlations between property attributes and property values of residential properties in the neighborhood based on a neighborhood-specific model, wherein determining said correlations comprises executing the neighborhood-specific model that learns said correlations.
20 . The computing system of claim 19 , further comprising generating, based on said correlations, weight values specifying amounts of weight to give to specific property attributes in estimating property values of residential properties in the neighborhood.
21 . The computing system of claim 19 , wherein merging neighboring blocks comprises:
for a first block and second block that are neighboring, generating a normalized score representing a degree to which residential properties in the first block are similar to residential properties in the second block; and
using the normalized score to determine whether to merge the first block with the second block.