Computer Methods and Interfaces for Efficient Categorization of Voluminous Data
Computer-aided categorization or classification of numerous data records can be controlled and guided through a user interface that accepts user input to produce clustering training data, and that conveys the improved automatic classification results efficiently to the user. Features that facilitate working with thousands or millions of data records are described and claimed.
1 . A user interface for training a machine-learning algorithm to classify a multitude of data records into a plurality of similar clusters, comprising:
preparing a two-dimensional display area;
selecting a representative subset of the multitude of data records;
displaying representative images corresponding to the representative subset on the two-dimensional display area;
receiving user input to indicate that a first representative image should be clustered with a second representative image;
amending a clustering algorithm according to the user input to produce an amended clustering algorithm;
applying the amended clustering algorithm to the representative subset to produce an improved clustering;
adjusting a position of the representative images besides the first representative image and the second representative image to reflect the improved clustering.
2 . The user interface of claim 1 , wherein the plurality of similar clusters is two similar clusters.
3 . The user interface of claim 1 , wherein a count of the plurality of similar clusters is between three similar clusters and ten similar clusters.
4 . The user interface of claim 1 , further comprising:
displaying abridged symbols on the two-dimensional display area, each abridged symbol to represent at least one data record of the multitude of data records that is not a member of the representative subset;
applying the amended clustering algorithm to data records represented by the abridged symbols; and
adjusting a position of the abridged symbols to reflect the improved clustering.