IP Library Granted Patent US 8,402,029
Granted Patent B2
US 8,402,029 · App. 13/194,609 · Granted Mar 19, 2013

Clustering system and method

Inventors: Raul Valdes-Perez (Pittsburgh, PA); Andre dos Santos Lessa (Pittsburgh, PA); Christopher Palmer (Pittsburgh, PA); Jerome Pesenti (Pittsburgh, PA)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,402,029
App. No.
13/194,609
Granted
Mar 19, 2013
Kind
B2
Abstract

An increase in information available to a user of computing technologies has a tendency to increase the number of topics that are similarly related. Given the large amount of information that is now available, it is increasingly likely that a first set of search results generated in response to an initial search query will contain information that is not of interest to the user. What is needed in the art is a technique to enable a search query to be conducted by taking advantage of linguistic feedback. Furthermore, what is needed is a technique to enable the presentation of search results to be refined in a manner based on what is not of interest to a user, either intrinsically or because the user has already seen and evaluated certain information and next wants to see more or different information.

Claims (64)

1. A non-transitory computer-readable medium comprising instructions for execution by a computer, the instructions implementing steps in a method comprising:

storing search results at a server based on a search engine query, wherein said search results comprise a plurality of items;

generating at the server a first set of clusters responsive to the search engine query, wherein each of said items is associated with at least one cluster in said first set of clusters;

sending the first set of clusters to a terminal configured to display the first set of clusters;

receiving at the server user input consisting of an indication to recluster the search results;

generating at the server a second set of clusters, wherein said second set of clusters excludes one or more clusters from said first set of clusters, and wherein each of said items is associated with at least one cluster in said second set of clusters; and

sending the second set of clusters to the terminal configured to display the second set of clusters,

wherein each cluster is defined by a cluster title, and wherein generating at the server a second set of clusters comprises excluding from the second set of clusters one or more cluster titles used in said first set of clusters, excluding the literal phraseology of at least one cluster title in the first set of clusters from the second set of clusters, and excluding a linguistic equivalence class corresponding to at least one cluster title in the first set of clusters from the second set of clusters, and

wherein generating the second set of clusters comprises excluding each displayed cluster of the first set of clusters from the second set of clusters.

2. The non-transitory computer readable medium of claim 1 , wherein each generating step comprises:

determining one or more linguistic equivalence classes, each linguistic equivalence class identifying a primary term and one or more corresponding linguistically similar terms, and

for each linguistic equivalence class, treating all linguistically similar terms within the search results as identical to the corresponding primary term.

3. The non-transitory computer readable medium of claim 1 , wherein generating the second set of clusters comprises allowing the second set of clusters to use a portion of a title of at least one cluster within the first set of clusters.

4. The non-transitory computer readable medium of claim 1 , wherein the steps of generating each of the first and second set of clusters excludes clusters that would otherwise only include stopwords.

5. A method, implemented in a computer, comprising:

determining in the computer a first set of clusters to display from a plurality of search query results, each of the plurality of search query results being associated with at least one cluster in the first set of clusters;

determining, in the computer, a title to display of at least one cluster in the first set of clusters;

after determining the title to display of the at least one cluster in the first set of clusters, identifying an input indication to recluster the plurality of search query results;

determining in the computer a second set of clusters to display, which excludes one or more clusters from the first set of clusters, each of the plurality of search query results being associated with at least one cluster in the second set of clusters.

6. The method of claim 5 , further comprising:

determining an elapsed amount of time between a time when the title of at least one cluster in the first set of clusters is displayed and a time when the input indication to recluster the plurality of search query results is identified, wherein

an amount of clusters of the one or more clusters excluded from the second set of clusters increases as the elapsed amount of time increases.

7. The method of claim 5 , further comprising:

determining a linguistic equivalence class of each cluster in the first set of clusters; and

when determining the second set of clusters, further excluding the linguistic equivalence class of each cluster in the first set of clusters.

8. The method of claim 5 , further comprising:

determining a title to display of at least one cluster in the second set of clusters, wherein

identifying an input indication, determining the second set of clusters, and determining the title to display of a least one cluster in the second set of cluster are continuously repeated as a loop.

9. The method of claim 5 , wherein:

identifying the input indication to recluster the plurality of search query results includes identifying a reclustering trigger that is automatically generated, without a user interaction, after a predetermined amount of time has elapsed after determining the title to display of at least one cluster in the first set of clusters.

10. A non-transitory computer-readable medium comprising instructions for execution by a computer, the instructions implementing the method comprising:

determining a first set of clusters to display from a plurality of search query results, each of the plurality of search query results being associated with at least one cluster in the first set of clusters;

determining a title to display of at least one cluster in the first set of clusters;

after determining the title to display of the at least one cluster in the set of clusters, identifying an input indication to recluster the plurality of search query results;

determining a second set of clusters to display, which excludes one or more clusters from the first set of clusters, each of the plurality of search query results being associated with at least one cluster in the second set of clusters.

11. The non-transitory computer-readable medium of claim 10 , wherein the method further comprises:

determining an elapsed amount of time between a time when the title of at least one cluster in the first set of clusters is determined and a time when the input indication to recluster the plurality of search query results is identified, wherein

an amount of clusters of the one or more clusters excluded from the second set of clusters increases as the elapsed amount of time increases.

12. The non-transitory computer-readable medium of claim 10 , wherein the method further comprises:

determining a linguistic equivalence class of each cluster in the first set of clusters; and

when determining the second set of clusters to display, further excluding the linguistic equivalence class of each cluster in the first set of clusters.

13. The non-transitory computer-readable medium of claim 10 , wherein the method further comprises:

determining a title to display of at least one cluster in the second set of clusters, wherein

identifying an input indication, determining the second set of clusters, and determining the to display title of a least one cluster in the second set of cluster are continuously repeated as a loop.

14. The non-transitory computer-readable medium of claim 10 , wherein:

identifying the input indication to recluster the plurality of search query results includes identifying a reclustering trigger that is automatically generated, without a user interaction, after a predetermined amount of time has elapsed after determining the title to display of at least one cluster in the first set of clusters.

15. An apparatus comprising:

a processor; and

memory storing computer executable instructions that, when executed by the processor, perform a method of clustering, comprising:

determining a first set of clusters to display from a plurality of search query results, each of the plurality of search query results being associated with at least one cluster in the first set of clusters;

determining a title to display of at least one cluster in the first set of clusters;

after determining the title to display of the at least one cluster in the first set of clusters, identifying an input indication to recluster the plurality of search query results;

determining a second set of clusters to display, which excludes one or more clusters from the first set of clusters, each of the plurality of search query results being associated with at least one cluster in the second set of clusters.

16. An apparatus according to claim 15 , wherein the method of clustering further comprises:

determining an elapsed amount of time between a time when the title of at least one cluster in the first set of clusters is displayed and a time when the input indication to recluster the plurality of search query results is identified, wherein

an amount of clusters of the one or more clusters excluded from the second set of clusters increases as the elapsed amount of time increases.

17. An apparatus according to claim 15 , wherein the method of clustering further comprises:

determining a linguistic equivalence class of each cluster in the first set of clusters; and

when determining the second set of clusters, further excluding the linguistic equivalence class of each cluster in the first set of clusters.

18. An apparatus according to claim 15 , wherein the method of clustering further comprises:

determining a title to display of at least one cluster in the second set of clusters, wherein

identifying an input indication, determining the second set of clusters, and determining the title to display of a least one cluster in the second set of cluster are continuously repeated as a loop.

19. An apparatus according to claim 15 , wherein the method of clustering further comprises:

identifying the input indication to recluster the plurality of search query results includes identifying a reclustering trigger that is automatically generated, without a user interaction, after a predetermined amount of time has elapsed after determining the title to display of at least one cluster in the first set of clusters.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2013
From: VIVISIMO, INC.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029709/0365 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2011
From: VALDES-PEREZ, RAUL; DOS SANTOS LESSA, ANDRE; PALMER, CHRISTOPER; PESENTI, JEROME
To: VIVISIMO, INC.
Reel/Frame 026946/0855 →
Continuity (2)
Division 11774908 · Jul 9, 2007
Related Publication 20110313990A1 · Dec 22, 2011