Apparatus and method for clustering related tuples derived from content in an unstructured database
A computer implemented method includes receiving a baseline dataset divisible by temporal or logical criteria. A target dataset representing a small fraction of the baseline dataset is received. The target dataset is segmented by the temporal or logical criteria. Numbers of documents containing individual terms within the baseline dataset are identified to form baseline singles. Numbers of documents containing common combinations of individual terms within the target dataset are identified to form target tuples. Combinations of the target tuples are coalesced based upon baseline singles criteria to form a first coalesced dataset representing significant content in the target dataset.
1 . A computer implemented method, comprising:
receiving at a server a baseline dataset from a network connected content source machine, where the baseline dataset is divisible by temporal or logical criteria;
receiving at the server a target dataset from the network connected content source machine, where the target dataset represents a small fraction of the baseline dataset, the target dataset being segmented by the temporal or logical criteria;
identifying, without user direction, numbers of documents containing individual terms within the baseline dataset to form baseline singles;
identifying, without user direction, numbers of documents containing common combinations of individual terms within the target dataset to form target tuples, where each common combination of individual terms provides context for each term in the common combination; and
coalescing, without user direction, combinations of the target tuples based upon baseline singles criteria to form a first coalesced dataset representing content in the target dataset, where the content includes thematic conceptual clusters in the target dataset.
2 . The computer implemented method of claim 1 wherein a temporal criterion is a 24-hour period.
3 . The computer implemented method of claim 1 wherein a logical criterion is segments within a publication.
4 . The computer implemented method of claim 1 further comprising pre-processing the baseline dataset and the target dataset by removing punctuation, stemming terms, and connecting terms that belong together.
5 . The computer implemented method of claim 1 wherein coalescing includes ranking target tuples by numbers of instances and then filtering and thereby skimming target tuples to a fraction of all target tuples.
6 . The computer implemented method of claim 1 wherein the number of tuples subject to further processing is less than 2% of the number of tuples formed.
7 . The computer implemented method of claim 6 wherein the number of tuples coalesced is less than 2% of the number of tuples subject to further processing.
8 . The computer implemented method of claim 7 wherein the number of tuples after final filtering is less than 2% of the number of tuples coalesced.