IP Library Granted Patent US 7,558,769
Granted Patent B2
US 7,558,769 · App. 11/241,694 · Granted Jul 7, 2009

Identifying clusters of similar reviews and displaying representative reviews from multiple clusters

Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,558,769
App. No.
11/241,694
Granted
Jul 7, 2009
Kind
B2
Abstract

A method and system of selecting reviews for display are described. Reviews for a subject are identified and organized into two or more clusters. Reviews are selected from each cluster. A response that includes content from the selected reviews is generated. The content may include the full content or snippets of at least some of the selected reviews.

Claims (74)

1. A method of processing reviews, comprising:

at a server:

identifying a plurality of reviews in a corpus of reviews;

organizing the plurality of reviews into a plurality of clusters based on terms

in the reviews and importance of the terms in the corpus of reviews, wherein a result of the organizing is that each review of the plurality of reviews is assigned to a single one of the plurality of clusters;

receiving a review summary request from a user;

selecting a subset of reviews from each cluster;

determining at least one quality score for each review in the selected subset;

selecting content from the selected subset of reviews in accordance with the determined quality scores;

generating a response that includes the selected content from the selected subset of reviews; and

transmitting the response to the user.

2. The method of claim 1 , wherein selecting a subset of reviews comprises:

identifying a number of reviews in each cluster, the number of reviews determining a size of the respective cluster; and

selecting reviews from each cluster in proportion to the cluster sizes.

3. The method of claim 1 , wherein selecting reviews comprises selecting reviews from each cluster based on one or more predefined criteria and in proportion to cluster sizes, wherein the size of each cluster is determined by a number of reviews in each cluster.

4. The method of claim 1 , wherein organizing comprises:

generating a vector for each of the plurality of reviews, each vector comprising elements corresponding to terms in a respective review; and

organizing the plurality of reviews into the plurality of clusters based on the vectors.

5. The method of claim 4 , wherein organizing comprises organizing the plurality of reviews into one or more clusters based on the vectors and vectors associated with one or more canonical reviews.

6. The method of claim 1 , wherein generating a response comprises generating snippets of at least a subset of the selected reviews.

7. The method of claim 6 , wherein generating a snippet of a review comprises:

partitioning the review into one or more partitions;

selecting a subset of the partitions based on predefined quality score criteria; and

generating the snippet including content from the selected subset of the partitions.

8. The method of claim 1 , wherein the importance of the terms in the corpus of reviews is based at least in part on frequency of occurrence of the terms in the corpus of reviews.

9. The method of claim 1 , wherein the importance of a respective term in a respective review is based at least in part on an inverse document frequency value that is inversely related to frequency of occurrence of the respective term in the corpus of reviews.

10. A system for processing reviews, comprising:

one or more processors; and

memory storing one or more programs to be executed by the one or more processors, the one or more programs including instructions:

to identify a plurality of reviews; to organize the plurality of reviews into a plurality of clusters based on terms in the reviews and importance of the terms in the corpus of reviews, wherein a result of the organizing is that each review of the plurality of reviews is assigned to a single one of the plurality of clusters;

to receive a review summary request from a user;

to select a subset of reviews from each cluster;

to determine at least one quality score for each review in the selected subset;

to select content from the selected subset of reviews in accordance with the determined quality scores;

to generate a response that includes the selected content from the selected subset of reviews; and

to transmit the response to the user.

11. The system of claim 10 , wherein the one or more programs include instructions:

to identify a number of reviews in each cluster, the number of reviews determining a size of the respective cluster; and

to select reviews from each cluster in proportion to the cluster sizes.

12. The system of claim 10 , wherein the one or more programs include instructions to select reviews from each cluster based on one or more predefined criteria and in proportion to cluster sizes, wherein the size of each cluster is determined by a number of reviews in each cluster.

13. The system of claim 10 , wherein the one or more programs include instructions:

to generate a vector for each of the plurality of reviews, each vector comprising elements corresponding to terms in a respective review; and

to organize the plurality of reviews into the plurality of clusters based on the vectors.

14. The system of claim 13 , wherein the one or more programs include instructions to organize the plurality of reviews into one or more clusters based on the vectors and vectors associated with one or more canonical reviews.

15. The system of claim 10 , wherein the one or more programs include instructions to generate snippets of at least a subset of the selected reviews.

16. The system of claim 15 , wherein the one or more programs include instructions:

to partition the review into one or more partitions;

to select a subset of the partitions based on predefined quality score criteria; and

to generate the snippet including content from the selected subset of the partitions.

17. A computer readable storage medium, storing one or more programs for execution by one or more processors at a server computer, the one or more programs comprising instructions for:

identifying a plurality of reviews;

organizing the plurality of reviews into a plurality of clusters based on terms in the reviews and importance of the terms in the corpus of reviews, wherein a result of the

organizing is that each review of the plurality of reviews is assigned to a single one of the plurality of clusters;

receiving a review summary request from a user;

selecting a subset of reviews from each cluster;

determining at least one quality score for each review in the selected subset;

selecting content from the selected subset of reviews in accordance with the determined quality scores;

generating a response that includes content from the selected subset of reviews; and

transmitting the response to the user.

18. The computer readable storage medium of claim 17 , wherein the importance of a respective term in a respective review is based at least in part on an inverse document frequency value that is inversely related to frequency of occurrence of the respective term in the corpus of reviews.

19. The computer readable storage medium of claim 17 , wherein the importance of the terms in the corpus of reviews is based at least in part on frequency of occurrence of the terms in the corpus of reviews.

20. The computer readable storage medium of claim 19 , wherein the instructions for selecting comprise instructions for:

identifying a number of reviews in each cluster, the number of reviews determining a size of the respective cluster; and

selecting reviews from each cluster in proportion to the cluster sizes.

21. The computer readable storage medium of claim 19 , wherein the instructions for organizing comprise instructions for:

generating a vector for each of the plurality of reviews, each vector comprising elements corresponding to terms in a respective review; and

organizing the plurality of reviews into the plurality of clusters based on the vectors.

22. The computer readable storage medium of claim 19 , wherein the instructions for generating a response comprise instructions for generating snippets of at least a subset of the selected reviews.

23. The computer readable storage medium of claim 22 , wherein the instructions for generating a snippet of a review comprise instructions for:

partitioning the review into one or more partitions;

selecting a subset of the partitions based on predefined quality score criteria; and

generating the snippet including content from the selected subset of the partitions.

24. The system of claim 10 , wherein the importance of the terms in the corpus of reviews is based at least in part on frequency of occurrence of the terms in the corpus of reviews.

25. The system of claim 10 , wherein the importance of a respective term in a respective review is based at least in part on an inverse document frequency value that is inversely related to frequency of occurrence of the respective term in the corpus of reviews.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0610 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2005
From: SCOTT, JAMES KEVIN; DAVE, KUSHAL B.; HYLTON, JEREMY A.
To: GOOGLE INC.
Reel/Frame 016946/0745 →
Continuity (1)
Related Publication 20070078845A1 · Apr 5, 2007