IP Library Granted Patent US 7,587,381
Granted Patent B1
US 7,587,381 · App. 10/350,869 · Granted Sep 8, 2009

Method for extracting a compact representation of the topical content of an electronic text

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,587,381
App. No.
10/350,869
Granted
Sep 8, 2009
Kind
B1
Abstract

An electronic document is parsed to remove irrelevant text and to identify the significant elements of the retained text. The elements are assigned scores representing their significance to the topical content of the document. A matrix of element-pairs is constructed such that the matrix nodes represent the result of one or more functions of the scores and other attributes of the paired elements. The resulting matrix is a compact representation of topical content that affords great precision in information retrieval applications that depend on measurements of the relatedness of topical content.

Claims (33)

1. A computer-implemented method comprising:

analyzing a document to identify elements within the document, wherein elements are words or phrases that explicitly appear in the document;

assigning an element score to each of one or more of the elements, the element score indicating the element's significance to topical content of the document;

generating a matrix of element-pairs from the elements identified within the document, and for one or more of the element-pairs, calculating an element-pair score indicating the element-pair's significance to topical content of the document; and

generating a sentence set for one or more elements identified within the document, the sentence set comprising ordinal values indicating sentences in which the elements occur, wherein said calculating an element-pair score indicating the element-pair's significance to topical content of the document includes weighting the element-pair score based, in part, on a measure of a shortest distance between occurrences of the elements of the element-pair, as indicated by the sentence sets associated with each element-pair, whereby the sentence set is stored in memory and used by information retrieval systems,

whereby element-pair scores are stored as a data structure and executable by information retrieval systems.

2. The computer-implemented method of claim 1 , wherein calculating an element-pair score indicating the element-pair's significance to topical content of the document includes calculating the element-pair score as a function of the individual element scores associated with the elements of the element-pair.

3. The computer-implemented method of claim 1 , wherein assigning the element score to each of the one or more elements includes generating the element score based, in part, on a measure of the frequency with which the element occurs in the document.

4. The computer-implemented method of claim 1 , wherein assigning the element score to each of the one or more elements includes weighting the element score based, in part, on determining if the element occurs in a particular section or sentence of the document.

5. The computer-implemented method of claim 1 , wherein assigning the element score to each of the one or more elements includes weighting the element score based, in part, on determining if the element is included in a predetermined list of words and/or phrases.

6. The computer-implemented method of claim 1 , wherein analyzing a document to identify elements within the document includes processing the document to select words and/or phrases, which are representative of topical content of the document, in accordance with one or more predetermined criteria such as discarding sections of the document except a title and a body, discarding navigation elements of the document on an HTML page, and discarding components of the document whose size falls below a configurable threshold.

7. The computer-implemented method of claim 1 , wherein analyzing a document to identify elements within the document includes eliminating text determined not to be representative of topical content of the document.

8. The computer-implemented method of claim 7 , wherein eliminating text determined not to be representative of topical content of the document includes eliminating text based on a determination that the text corresponds to parts of speech not representative of topical content of the document.

9. The computer-implemented method of claim 7 , wherein eliminating text determined not to be representative of topical content of the document includes eliminating text determined to be present in a predefined list of words.

10. The computer-implemented method of claim 1 , wherein analyzing a document to identify elements within the document includes constructing a phrase from two or more words according to one or more predetermined criteria such as constructing sequences of adjectives and nouns, as determined by a setting of a configuration parameter.

11. The computer-implemented method of claim 1 wherein analyzing a document to identify elements within the document includes:

parsing the document into sections representing structural components of the document; and

analyzing only those sections determined to contain words and/or phrases representative of topical content of the document.

12. The computer-implemented method of claim 11 , wherein assigning an element score to one or more elements includes generating a score on a per section basis, and combining scores from one or more sections to calculate the element score assigned to a particular element.

13. The computer-implemented method of claim 1 , wherein analyzing a document to identify elements within the document includes analyzing stem forms of words so that two words with a common stem form are considered as the same element.

14. The computer-implemented method of claim 1 , further comprising:

eliminating from the matrix those element-pairs with an element-pair score lower than a predetermined threshold.

15. The computer-implemented method of claim 1 , wherein the matrix is manipulated to form a vector such that the dimensions of the vector are the element-pairs ordered by the element-pair score in descending order, and the magnitude in each dimension is the element-pair score, wherein element-pairs with an element-pair score lower than a predetermined threshold are eliminated from the vector.

16. A computer-implemented method comprising:

analyzing a document to identify textual elements within one or more sections of the document in accordance with one or more predetermined criteria such as discarding sections of the document except a title and a body, discarding navigation elements of the document on an HTML page, and discarding components of the document whose size falls below a configurable threshold, wherein textual elements are words or phrases that explicitly appear in the document;

assigning a base score to each textual element based on a measure of the frequency with which the textual element occurs in the document;

determining a sentence set for each textual element, the sentence set comprising ordinal values associated with sentences in which the textual element occurs within a section of the document; and

generating a matrix by pairing textual elements identified within the document to create element-pairs and assigning an element-pair a significance score calculated as a function of the base scores of each textual element in the element pair, wherein the significance score is weighted based on a shortest distance between textual elements within a section of the document, as indicated by the sentence set for each textual element in the element pair,

whereby significance scores are stored as a data structure and executable by information retrieval systems.

17. The computer-implemented method of claim 16 , wherein analyzing a document to identify textual elements within one or more sections of the document in accordance with one or more predefined document processing rules includes processing the document to select words and/or phrases, which are representative of topical content of the document, in accordance with the one or more predefined document processing rules.

18. The computer-implemented method of claim 16 , wherein the base score is weighted based on the occurrence of the textual element in a particular section or sentence of the document.

19. The computer-implemented method of claim 16 , wherein the base score is weighted based on the inclusion of the textual element in a predetermined list of textual elements.

20. The computer-implemented method of claim 16 , wherein assigning a base score to each textual element based on a measure of the frequency with which the textual element occurs in the document includes generating a section score on a per section basis, and combining section scores from one or more sections to calculate the base score assigned to each textual element.

Assignments (11)
CHANGE OF NAME Recorded Aug 22, 2025
From: OUTBRAIN INC.
To: TEADS HOLDING CO.
Reel/Frame 072558/0062 →
RELEASE OF SECURITY INTEREST (REEL 034315, FRAME 0093) Recorded Feb 3, 2025
From: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY
To: OUTBRAIN INC.
Reel/Frame 070096/0212 →
SECURITY INTEREST Recorded Feb 3, 2025
From: OUTBRAIN INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 070096/0384 →
CHANGE OF NAME Recorded Feb 24, 2015
From: SPHERE SOURCE, INC.
To: SURPHACE INC.
Reel/Frame 035090/0228 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2015
From: SURPHACE ACQUISITION, INC.
To: OUTBRAIN, INC.
Reel/Frame 035017/0809 →
SECURITY AGREEMENT Recorded Nov 21, 2014
From: OUTBRAIN INC.
To: SILICON VALLEY BANK
Reel/Frame 034315/0093 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2011
From: SURPHACE, INC.
To: SURPHACE ACQUISITION, INC.
Reel/Frame 025743/0896 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Nov 16, 2010
From: BANK OF AMERICA, N A
To: AOL INC; AOL ADVERTISING INC; GOING INC; LIGHTNINGCAST LLC; MAPQUEST, INC; NETSCAPE COMMUNICATIONS CORPORATION; QUIGO TECHNOLOGIES LLC; SPHERE SOURCE, INC; TACODA LLC; TRUVEO, INC; YEDDA, INC
Reel/Frame 025323/0416 →
SECURITY AGREEMENT Recorded Dec 14, 2009
From: AOL INC.; AOL ADVERTISING INC.; BEBO, INC.; ICQ LLC; GOING, INC.; LIGHTNINGCAST LLC; MAPQUEST, INC.; NETSCAPE COMMUNICATIONS CORPORATION; QUIGO TECHNOLOGIES LLC; SPHERE SOURCE, INC.; TACODA LLC; TRUVEO, INC.; YEDDA, INC.
To: BANK OF AMERICAN, N.A. AS COLLATERAL AGENT
Reel/Frame 023649/0061 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2007
From: NIEKER, STEVEN; REMY, MARTIN
To: THINK TANK 23 LLC
Reel/Frame 020281/0027 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2007
From: THINK TANK 23 LLC
To: SPHERE SOURCE, INC.
Reel/Frame 020281/0063 →