IP Library Granted Patent US 10,970,353
Granted Patent B1
US 10,970,353 · App. 16/250,143 · Granted Apr 6, 2021

Ranking content using content and content authors

Inventors: Douwe Osinga (Sydney, AU); Stefan Christoph (Uster, CH)
Assignee: Google LLC
G06F16/955G06F16/382G06F16/93G06K9/00456G06K9/00463
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,970,353
App. No.
16/250,143
Granted
Apr 6, 2021
Kind
B1
Abstract

Methods, systems, and apparatus, including computer program products for identifying original content. In one aspect a method is described that includes identifying a first document in a collection of documents. The first document contains a content piece and the content piece does not occur in any earlier document in the collection. The first document is associated with a first author and the first author associated with a first rank. The first rank of the first author is determined using a score of the content piece. The score is a figure of merit of the content piece.

Claims (64)

1. A computer-implemented method comprising:

accessing, by one or more processors, a corpus of documents;

determining, by the one or more processors, content of a particular document in the corpus of documents;

determining, by the one or more processors, that a first group of documents in the corpus of documents are from a particular source;

determining, by the one or more processors, that a second group of documents in the corpus of documents includes content from the particular source, wherein the second group of documents does not include any documents from the first group;

comparing, by the one or more processors, the content of the particular document to the content from the particular source;

based on comparing the content of the particular document to the content from the particular source, determining, by the one or more processors, an amount of shared content between the content of the particular document and the content from the particular source;

based on the amount of shared content between the content of the particular document and the content, adjusting, by the one or more processors, a rank of the particular document in relation to other document in the corpus of documents; and

configuring, by the one or more processors, a web crawling process or search result ranking process for the particular document based on the adjusted rank.

2. The method of claim 1 , comprising:

fragmenting the particular document into multiple content pieces,

wherein comparing the content of the particular document to the content from the particular source comprises comparing the multiple content pieces of the particular document to the content form the particular source.

3. The method of claim 2 , wherein multiple content pieces represents a number of non-adjacent words.

4. The method of claim 1 , wherein the corpus of documents comprises documents that are classifed as news-related documents.

5. The method of claim 1 , wherein configuring the web crawling process comprises configuring a frequency with which a web crawler crawls a web server associated with the particular document based on the adjusted rank.

6. The method of claim 1 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document does not include content from the particular source, and

adjusting the rank of the particular document comprises decreasing the rank of the particular document based on determining that the content of the particular document does not include content from the particular source.

7. The method of claim 1 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document includes content from the particular source, and

adjusting the rank of the particular document comprises increasing the rank of the particular document based on determining that the content of the particular document includes content from the particular source.

8. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

accessing, by one or more processors, a corpus of documents;

determining, by the one or more processors, content of a particular document in the corpus of documents;

determining, by the one or more processors, that a first group of documents in the corpus of documents are from a particular source;

determining, by the one or more processors, that a second group of documents in the corpus of documents includes content from the particular source, wherein the second group of documents does not include any documents from the first group;

comparing, by the one or more processors, the content of the particular document to the content from the particular source;

based on comparing the content of the particular document to the content from the particular source, determining, by the one or more processors, an amount of shared content between the content of the particular document and the content from the particular source;

based on the amount of shared content between the content of the particular document and the content, adjusting, by the one or more processors, a rank of the particular document in relation to other document in the corpus of documents; and

configuring, by the one or more processors, a web crawling process or search result ranking process for the particular document based on the adjusted rank.

9. The system of claim 8 , wherein the operations comprise:

fragmenting the particular document into multiple content pieces,

wherein comparing the content of the particular document to the content from the particular source comprises comparing the multiple content pieces of the particular document to the content form the particular source.

10. The system of claim 9 , wherein multiple content pieces represents a number of non-adjacent words.

11. The system of claim 8 , wherein the corpus of documents comprises documents that are classifed as news-related documents.

12. The system of claim 8 , wherein configuring the web crawling process comprises configuring a frequency with which a web crawler crawls a web server associated with the particular document based on the adjusted rank.

13. The system of claim 8 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document does not include content from the particular source, and

adjusting the rank of the particular document comprises decreasing the rank of the particular document based on determining that the content of the particular document does not include content from the particular source.

14. The system of claim 8 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document includes content from the particular source, and

adjusting the rank of the particular document comprises increasing the rank of the particular document based on determining that the content of the particular document includes content from the particular source.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

accessing, by one or more processors, a corpus of documents;

determining, by the one or more processors, content of a particular document in the corpus of documents;

determining, by the one or more processors, that a first group of documents in the corpus of documents are from a particular source;

determining, by the one or more processors, that a second group of documents in the corpus of documents includes content from the particular source, wherein the second group of documents does not include any documents from the first group;

comparing, by the one or more processors, the content of the particular document to the content from the particular source;

based on comparing the content of the particular document to the content from the particular source, determining, by the one or more processors, an amount of shared content between the content of the particular document and the content from the particular source;

based on the amount of shared content between the content of the particular document and the content, adjusting, by the one or more processors, a rank of the particular document in relation to other document in the corpus of documents; and

configuring, by the one or more processors, a web crawling process or search result ranking process for the particular document based on the adjusted rank.

16. The medium of claim 15 , wherein the operations comprise:

fragmenting the particular document into multiple content pieces,

wherein comparing the content of the particular document to the content from the particular source comprises comparing the multiple content pieces of the particular document to the content form the particular source.

17. The medium of claim 15 , wherein the corpus of documents comprises documents that are classifed as news-related documents.

18. The medium of claim 15 , wherein configuring the web crawling process comprises configuring a frequency with which a web crawler crawls a web server associated with the particular document based on the adjusted rank.

19. The medium of claim 15 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document does not include content from the particular source, and

adjusting the rank of the particular document comprises decreasing the rank of the particular document based on determining that the content of the particular document does not include content from the particular source.

20. The medium of claim 15 , wherein:

determining an amount of shared content between the content of the particular document and the content from the particular source comprises determining, by the one or more processors, that the content of the particular document includes content from the particular source, and

adjusting the rank of the particular document comprises increasing the rank of the particular document based on determining that the content of the particular document includes content from the particular source.

Assignments (2)
CHANGE OF NAME Recorded Jul 23, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 049857/0839 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: OSINGA, DOUWE; CHRISTOPH, STEFAN
To: GOOGLE INC.
Reel/Frame 048253/0991 →
Continuity (4)
Continuation 15395316 · Dec 30, 2016
Continuation 14633377 · Feb 27, 2015
Continuation 13447806 · Apr 16, 2012
Continuation 11608189 · Dec 7, 2006