IP Library Granted Patent US 10,095,752
Granted Patent B1
US 10,095,752 · App. 15/145,486 · Granted Oct 9, 2018

Methods and apparatus for clustering news online content based on content freshness and quality of content source

Inventors: Michael Schmitt (Mountain View, CA); Krishna Bharat (San Jose, CA); Michael Curtiss (Sunnyvale, CA)
Assignee: Google LLC
G06F17/3053G06F17/3071H04L67/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,095,752
App. No.
15/145,486
Granted
Oct 9, 2018
Kind
B1
Abstract

Methods and apparatus are described for scoring documents in response, in part, to parameters related to the document, source, and/or cluster score. Methods and apparatus are also described for scoring a cluster in response, in part, to parameters related to documents within the cluster and/or sources corresponding to the documents within the cluster. In one embodiment, the invention may identify the source; detect a plurality of documents published by the source; analyze the plurality of documents with respect to at least one parameter, and determine a source score for the source in response, in part, to the parameter. In another embodiment, the invention may identify a topic; identify a plurality of clusters in response to the topic; analyze at least one parameter corresponding to each of the plurality of clusters; and calculate a cluster score for each of the plurality of clusters in response, in part, to the parameter.

Claims (38)

1. A computer-implemented method comprising:

identifying, by a processor, online documents published online by one or more sources;

calculating, by the processor, a first score based on a measure of freshness of a first online document of the online documents, the measure of freshness being based on an amount of time between a first time when the first online document of the online documents was published and a second time when an event described by the first online document occurred;

calculating, by the processor, a second score based on a quantity of the online documents that have a relationship to the first online document;

ranking, by the processor, the first online document based on the first score and the second score; and

providing, by the processor, the first online document for display based on the ranking of the first online document.

2. The computer-implemented method of claim 1 , wherein a centroid is calculated for the quantity of the online documents, the centroid uniquely describing a subject matter of the quantity of the online documents.

3. The computer-implemented method of claim 2 , wherein the second score is further based on whether a title of the first online document includes one or more words that match the centroid.

4. The computer-implemented method of claim 1 , wherein the relationship includes a common relationship to a subject matter.

5. The computer-implemented method of claim 1 , wherein the second score is further based on a number of views that the first online document received within a time frame.

6. The computer-implemented method of claim 1 , wherein the second score is based on circulation statistics of a first source of the one or more sources, the first source having published the first online document.

7. The computer-implemented method of claim 1 , wherein the online documents are each weighted for calculating the first score, the weighting based on a respective time when each of the online documents is published.

8. A system comprising:

a processor; and

a non-transitory computer readable medium storing instructions that, when executed by the processor, cause the processor to perform operations comprising:

identifying online documents published online by one or more sources;

calculating a first score based on a measure of freshness of a first online document of the online documents, the measure of freshness being based on an amount of time between a first time when the first online document of the online documents was published and a second time when an event described by the first online document occurred;

calculating a second score based on a quantity of the online documents that have a relationship to the first online document;

ranking the first online document based on the first score and the second score; and

providing the first online document for display based on the ranking of the first online document.

9. The system of claim 8 , wherein a centroid is calculated for the quantity of the online documents, the centroid uniquely describing a subject matter of the quantity of the online documents.

10. The system of claim 9 , wherein the second score is further based on whether a title of the first online document includes one or more words that match the centroid.

11. The system of claim 8 , wherein the relationship includes a common relationship to a subject matter.

12. The system of claim 8 , wherein the second score is further based on a number of views that the first online document received within a time frame.

13. The system of claim 8 , wherein the second score is based on circulation statistics of a first source of the one or more sources, the first source having published the first online document.

14. The system of claim 8 , wherein the online documents are each weighted for calculating the first score, the weighting based on a respective time when each of the online documents is published.

15. A non-transitory computer-readable medium having computer executable instructions for performing a method comprising:

identifying, by a processor, online documents published online by one or more sources;

calculating, by the processor, a first score based on a measure of freshness of a first online document of the online documents, the measure of freshness being based on an amount of time between a first time when the first online document of the online documents was published and a second time when an event described by the first online document occurred:

calculating, by the processor, a second score based on a quantity of the online documents that have a relationship to the first online document;

ranking, by the processor, the first online document based on the first score and the second score; and

providing, by the processor, the first online document for display based on the ranking of the first online document.

16. The non-transitory computer-readable medium of claim 15 , wherein a centroid is calculated for the quantity of the online documents, the centroid uniquely describing a subject matter of the quantity of the online documents.

17. The non-transitory computer-readable medium of claim 16 , wherein the second score is further based on whether a title of the first online document includes one or more words that match the centroid.

18. The non-transitory computer-readable medium of claim 15 , wherein the relationship includes a common relationship to a subject matter.

19. The non-transitory computer-readable medium of claim 15 , wherein the second score is further based on a number of views that the first online document received within a time frame.

20. The non-transitory computer-readable medium of claim 15 , wherein the second score is based on circulation statistics of a first source of the one or more sources, the first source having published the first online document.

21. The non-transitory computer-readable medium of claim 15 , the online documents are each weighted for calculating the first score, the weighting based on a respective time when each of the online documents is published.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2016
From: SCHMITT, MICHAEL; BHARAT, KRISHNA; CURTISS, MICHAEL
To: GOOGLE INC.
Reel/Frame 039216/0275 →
Continuity (4)
Continuation 13548930 · Jul 13, 2012
Continuation 12344153 · Dec 24, 2008
Continuation 10611269 · Jun 30, 2003
Provisional Application 60412287 · Sep 20, 2002