IP Library Granted Patent US 7,167,871
Granted Patent B2
US 7,167,871 · App. 10/232,714 · Granted Jan 23, 2007

Systems and methods for authoritativeness grading, estimation and sorting of documents in large heterogeneous document collections

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,167,871
App. No.
10/232,714
Granted
Jan 23, 2007
Kind
B2
Abstract

Systems and methods for determining the authoritativeness of a document based on textual, non-topical cues. The authoritativeness of a document is determined by evaluating a set of document content features contained within each document to determine a set of document content feature values, processing the set of document content feature values through a trained document textual authority model, and determining a textual authoritativeness value and/or textual authority class for each document evaluated using the predictive models included in the trained document textual authority model. Estimates of a document's textual authoritativeness value and/or textual authority class can be used to re-rank documents previously retrieved by a search, to expand and improve document query searches, to provide a more complete and robust determination of a document's authoritativeness, and to improve the aggregation of rank-ordered lists with numerically-ordered lists.

Claims (40)

1. A method for determining an authoritativeness of a document having a plurality of document content features, the method comprising:

determining a set of document content feature values of a document based on textual contents in the document, the document providing information regarding a subject;

determining an authoritativeness for the document based on the determined set of document content feature values using a trained document textual authority model, wherein determining the authoritativeness comprises determining a reliability of the document, the reliability indicative of whether the information, as provided in the document, is reliable regarding the subject; and

outputting the determined authoritativeness in association with the document.

2. The method of claim 1 , wherein determining the set of document content feature values comprises extracting a subset of document content features from the plurality of document content features.

3. The method of claim 2 , wherein extracting a subset of document content features is performed using one or more regression techniques or methods.

4. The method of claim 3 , wherein one or more regression techniques or methods comprises a stepwise regression technique.

5. The method of claim 2 , wherein extracting a subset of document content features is performed using one or more variable selection techniques or methods.

6. The method of claim 5 , wherein one or more variable selection techniques or methods comprises one or more of mutual information technique and AdaBoost technique.

7. The method of claim 1 , wherein determining the set of document content feature values comprises determining the set of document content feature values using one or more parsing techniques or methods.

8. The method of claim 1 , wherein determining the authoritativeness for the document comprises:

providing the set of document content feature values to the trained document textual authority model; and

determining a document textual authoritativeness value based at least on the set of document content feature values determined.

9. The method of claim 8 , wherein determining a document textual authoritativeness value is performed by processing the set of document content feature values using one or more statistical processes or techniques.

10. The method of claim 8 , wherein determining a document textual authoritativeness value is performed by processing the set of document content feature values using one or more metric-regression algorithms or methods.

11. The method of claim 8 , wherein determining a document textual authoritativeness value is performed by processing the set of document content feature values using an AdaBoost algorithm model or method.

12. The method of claim 1 , wherein determining an authoritativeness for the document further comprising determining a textual authority class for the document.

13. The method of claim 1 , wherein the plurality of document content features includes at least one or more question marks, semicolons, numerals, words with learned prefixes, words with learned suffixes, words in certain grammatical locations, HTML features, abbreviations and classes of abbreviations, text characteristics features, speech tagging features or readability indices features.

14. The method of claim 1 , wherein determining the set of document content feature values comprises determining the set of document content feature values based solely on the information provided in the document.

15. The method of claim 1 , wherein determining the reliability of the document comprises determining, from the information within the document, a background of an author of the document, an institutional affiliation of the author, whether the document reads as if the document is well-researched, and whether the document has been reviewed or examined by others.

16. A machine-readable medium that provides instructions for determining the authority of a document having a plurality of document content features, instructions, which when executed by a processor, cause the processor to perform operations comprising:

determining a set of document content feature values of a document based on textual contents in the document, the document providing information regarding a subject; and

determining at least one of textual authoritativeness value or textual authority class for the document based on the determined set of document content feature values using a trained document textual authority model, wherein determining the authoritativeness comprises determining a reliability of the document, the reliability indicative of whether the information, as provided in the document, is reliable regarding the subject; and

outputting the determined authoritativeness in association with the document.

17. The machine-readable medium according to claim 16 , wherein the plurality of document content features includes at least one or more question marks, semicolons, numerals, words with learned prefixes, words with learned suffixes, words in certain grammatical locations, HTML features, abbreviations and classes of abbreviations, text characteristics features, speech tagging features or readability indices features.

18. The machine-readable medium according to claim 16 , wherein determining the textual authoritativeness value or a textual authority class for the document comprises:

extracting a plurality of document content features from each document;

determining a set of document content feature values for each document using one or more parsing techniques or methods; and

determining a textual authoritativeness value or a textual authority class for the document by using one or more of metric regression or boosted decision tree algorithms or methods.

19. The machine-readable medium according to claim 16 , wherein determining the set of document content feature values comprises determining the set of document content feature values based solely on the information provided in the document.

20. The machine-readable medium according to claim 16 , wherein determining the reliability of the document comprises determining, from the information within the document, a background of an author of the document, an institutional affiliation of the author, whether the document reads as if the document is well-researched, and whether the document has been reviewed or examined by others.

21. A textual authority determining system that determines an authority of a document having a plurality of document content features, comprising:

a memory; and

a document textual authoritativeness value determination circuit or routine that:

determines at least a textual authoritativeness value for the document based on textual contents in the document by processing a set of document content feature values determined for a subset of document content features extracted from the plurality of document content features using one or more of metric regression or boosted decision tree algorithms or methods, the document providing information regarding a subject, wherein determining the authoritativeness comprises determining a reliability of the document, the authoritativeness indicative of whether the information, as provided in the document, is reliable regarding the subject; and

outputs the determined authoritativeness in association with the document.

22. The textual authority determining system of claim 21 , wherein the plurality of document content features includes at least one or more question marks, semicolons, numerals, words with learned prefixes, words with learned suffixes, words in certain grammatical locations, HTML features, abbreviations and classes of abbreviations, text characteristics features, speech tagging features or readability indices features.

23. The textual authority determining system of claim 21 , further comprising a document content features extraction circuit or routine that determines a subset of document content features from the plurality of document content features using a stepwise regression process.

24. The textual authority determining system of claim 21 , wherein the document textual authoritativeness value determination circuit or routine determines the document content feature values based solely on the information provided in the document.

25. The textual authority determining system of claim 21 , wherein determining the reliability of the document comprises determining, from the information within the document, a background of an author of the document, an institutional affiliation of the author, whether the document reads as if the document is well-researched, and whether the document has been reviewed or examined by others.

Assignments (11)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 062740/0214 Recorded May 18, 2023
From: CITIBANK, N.A., AS AGENT
To: XEROX CORPORATION
Reel/Frame 063694/0122 →
SECURITY INTEREST Recorded Nov 10, 2022
From: XEROX CORPORATION
To: CITIBANK, N.A., AS AGENT
Reel/Frame 062740/0214 →
RELEASE OF SECURITY INTEREST Recorded Sep 7, 2022
From: JPMORGAN CHASE BANK, N.A. AS SUCCESSOR-IN-INTEREST ADMINISTRATIVE AGENT AND COLLATERAL AGENT TO JPMORGAN CHASE BANK
To: XEROX CORPORATION
Reel/Frame 066728/0193 →
RELEASE OF SECURITY INTEREST Recorded Feb 13, 2014
From: JPMORGAN CHASE BANK, N.A.
To: XEROX CORPORATION
Reel/Frame 032210/0001 →
RELEASE OF SECURITY INTEREST Recorded Feb 3, 2014
From: BANK ONE, NA
To: XEROX CORPORATION
Reel/Frame 032119/0854 →
SECURITY AGREEMENT Recorded Oct 31, 2003
From: XEROX CORPORATION
To: JPMORGAN CHASE BANK, AS COLLATERAL AGENT
Reel/Frame 015134/0476 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2002
From: FARAHAT, AYMAN O.; CHEN, FRANCINE R.; MATHIS, CHARLES R.; NUNBERG, GEOFFREY D.
To: XEROX CORPORATION
Reel/Frame 013263/0228 →