IP Library Granted Patent US 9,152,730
Granted Patent B2
US 9,152,730 · App. 13/563,060 · Granted Oct 6, 2015

Extracting principal content from web pages

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,152,730
App. No.
13/563,060
Granted
Oct 6, 2015
Kind
B2
Abstract

Extracting principal content from Web pages includes identifying and classifying items on the Web page, building a list of candidates, calculating candidate scores, selecting a top score candidate, performing clean up processing for the top score candidate, and performing final page processing for the top score candidate. Candidate scores may vary according to a number of paragraphs and images grouped according to size. A word length of CJK (Chinese-Japanese-Korean) text may be determined according to punctuation therein. Candidate scores may be modified according to a number of containers and pieces and wherein a container is a Web page element that is associated with tags ‘body’, ‘div’, ‘td’, ‘li’, ‘article/section’ and pieces are candidates that do not include other candidates. Candidate scores may be modified according to a number of ratios corresponding to text and link density.

Claims (59)

1. A method of extracting principal content from Web pages, comprising:

identifying and classifying items on the Web page;

building a list of candidates;

calculating candidate scores that vary according to a plurality of weights assigned to paragraphs and images of the Web page, wherein a particular paragraph is provided with a first weight based on a number of lines in the particular paragraph and is provided with a second weight based on a number of words in the particular paragraph, the first weight being independent of the second weight and a particular image is provided with a weight based on a size of the particular image;

selecting a top score candidate;

performing clean up processing for the top score candidate; and

performing final page processing for the top score candidate.

2. A method, according to claim 1 , where a word length of CJK (Chinese-Japanese-Korean) text is determined according to punctuation therein.

3. A method, according to claim 1 , wherein calculating candidate scores includes determining an Initial Score using the formula:

Initial Score=1000×(1.5×( N 5-line-paragraphs +N 80-word-paragraphs )+ N 3-line-paragraphs +N 50-word-paragraphs +3 ×N large-images −0.5×( N small-images +N skipped-images ))

where N 5-line-paragraphs is a number of five line paragraphs in the candidate, N 80-word-paragraphs is a number of 80 word paragraphs in the candidate, N 3-line-paragraphs is a number of 3 line paragraphs in the candidate, N 50-word-paragraphs is a number of 50 word paragraphs in the candidate, N large-images is a number of images that contain at least 50K pixels or have a width of at least 350 pixels and a height of at least 75 pixels, N skipped-images is a number of images that have a size not exceeding 5×5 or are present on Web page as references to at least one of: blacklisted and quarantined sites, and all other images are count toward N small-images .

4. A method, according to claim 1 , wherein candidate scores are modified according to a number of containers and pieces and wherein a container is a Web page element that is associated with tags ‘body’, ‘div’, ‘td’, ‘li’, ‘article/section’ and pieces are candidates that do not include other candidates.

5. A method, according to claim 4 , wherein candidate scores are modified using a formula:

Modified Initial Score=Initial Score×(⅓+⅔×⅓×(1 /N pieces +1 /N candidates +1 /N containers )).

6. A method, according to claim 4 , wherein candidate scores are modified according to a number of ratios corresponding to text and link density.

7. A method, according to claim 6 , wherein the ratios include a ratio of a length of regular text for the candidate to a total length of regular text, a number of words in regular text of the candidate to a total number of words in regular text, a length of regular text above the candidate to a total length of regular text, and a number of words in regular text above the candidate to a total number of words in regular text.

8. A method, according to claim 7 , wherein candidate scores are modified using a formula:

Modified Initial Score n+1 =Modified Initial Score n ×(Percentage n ×(ratio degree ) n +(100−Percentage n ))/100

wherein Percentage n and ratio degree are predetermined values that are empirically determined for each of n ratios.

9. A method, according to claim 1 , further comprising:

following selecting the top score candidate, determining if the top score candidate meets predetermined criteria; and

if the top score candidate does not meet predetermined criteria, determining if a different top score candidate should be selected.

10. A method, according to claim 9 , wherein a first set of formulas is used to determine the top score candidate and a second, different, set of formulas is used to determine if a different top score candidate should be selected.

11. A method, according to claim 10 , wherein a different top score candidate is not used if using the second set of formulas results in the same top score candidate as using the first set of formulas.

12. A method, according to claim 9 , wherein the predetermined criteria is selected from the group consisting of: whether the top score candidate has less than 25 embedded containers, wherein a container is a Web page element that is associated with tags ‘body’, div′, ‘td’, ‘article/section’, whether the top score candidate has no embedded other candidates and whether the top score candidate has no more than three embedded candidates that have other embedded candidates.

13. A method, according to claim 1 , wherein performing clean up processing includes removing floating elements, link boxes, navigation panels, and videos from unknown sources.

14. A method, according to claim 1 , wherein performing clean up processing includes reformatting the top scoring candidate by rewriting the text thereof as a new HTML page using only feasible HTML, tags and attributes and ignoring stylistic deficiencies and unnecessary elements in the original page format.

15. A method, according to claim 1 , further comprising:

a user providing an indication of portions of a displayed top score candidate; and

clipping portions indicated by the user, wherein the portions are subsequently used by other software.

16. A non-transitory computer-readable medium containing software that extracts principal content from Web pages, the software comprising:

executable code that identifies and classifies items on the Web page;

executable code that builds a list of candidates;

executable code that calculates candidate scores that vary according to a plurality of weights assigned to paragraphs and images of the Web page, wherein a particular paragraph is provided with a first weight based on a number of lines in the particular paragraph and is provided with a second weight based on a number of words in the particular paragraph, the first weight being independent of the second weight and a particular image is provided with a weight based on a size of the particular image;

executable code that selects a top score candidate;

executable code that performs clean up processing for the top score candidate; and

executable code that performs final page processing for the top score candidate.

17. A non-transitory computer-readable medium, according to claim 16 , where a word length of CJK (Chinese-Japanese-Korean) text is determined according to punctuation therein.

18. A non-transitory computer-readable medium, according to claim 16 , wherein executable code that calculates candidate scores determines an Initial Score using the formula:

Initial Score=1000×(1.5×( N 5-line-paragraphs +N 80-word-paragraphs )+ N 3-line-paragraphs +N 50-word-paragraphs +3 ×N large-images −0.5×( N small-images +N skipped-images ))

where N 5-line-paragraphs is a number of five line paragraphs in the candidate, N 80-word-paragraphs is a number of 80 word paragraphs in the candidate, N 3-line-paragraphs is a number of 3 line paragraphs in the candidate, N 50-word-paragraphs is a number of 50 word paragraphs in the candidate, N large-images is a number of images that contain at least 50K pixels or have a width of at least 350 pixels and a height of at least 75 pixels, N skipped-images is a number of images that have a size not exceeding 5×5 or are present on Web page as references to at least one of: blacklisted and quarantined sites, and all other images are count toward N small-images .

19. A non-transitory computer-readable medium, according to claim 16 , wherein candidate scores are modified according to a number of containers and pieces and wherein a container is a Web page element that is associated with tags ‘body’, ‘div’, ‘td’, ‘li’, ‘article/section’ and pieces are candidates that do not include other candidates.

20. A non-transitory computer-readable medium, according to claim 19 , wherein candidate scores are modified using a formula:

Modified Initial Score=Initial Score×(⅓+⅔×⅓×(1 /N pieces +1 /N candidates +1 /N containers )).

21. A non-transitory computer-readable medium, according to claim 19 , wherein candidate scores are modified according to a number of ratios corresponding to text and link density.

22. A non-transitory computer-readable medium, according to claim 21 , wherein the ratios include a ratio of a length of regular text for the candidate to a total length of regular text, a number of words in regular text of the candidate to a total number of words in regular text, a length of regular text above the candidate to a total length of regular text, and a number of words in regular text above the candidate to a total number of words in regular text.

23. A non-transitory computer-readable medium, according to claim 22 , wherein candidate scores are modified using a formula:

Modified Initial Score n+1 =Modified Initial Score n ×(Percentage n ×(ratio degree ) n +(100−Percentage n ))/100

wherein Percentage n and ratio degree are predetermined values that are empirically determined for each of n ratios.

24. A non-transitory computer-readable medium, according to claim 16 , further comprising:

executable code that determines if the top score candidate meets predetermined criteria following selecting the top score candidate; and

executable code that determines if a different top score candidate should be selected if the top score candidate does not meet the predetermined criteria.

25. A non-transitory computer-readable medium, according to claim 24 , wherein a first set of formulas is used to determine the top score candidate and a second, different, set of formulas is used to determine if a different top score candidate should be selected.

26. A non-transitory computer-readable medium, according to claim 25 , wherein a different top score candidate is not used if using the second set of formulas results in the same top score candidate as using the first set of formulas.

27. A non-transitory computer-readable medium, according to claim 24 , wherein the predetermined criteria is selected from the group consisting of: whether the top score candidate has less than 25 embedded containers, wherein a container is a Web page element that is associated with tags ‘body’, ‘div’, ‘td’, ‘li’, ‘article/section’, whether the top score candidate has no embedded other candidates and whether the top score candidate has no more than three embedded candidates that have other embedded candidates.

28. A non-transitory computer-readable medium, according to claim 16 , wherein executable code that performs clean up processing removes floating elements, link boxes, navigation panels, and videos from unknown sources.

29. A non-transitory computer-readable medium, according to claim 16 , wherein executable code that performs clean up processing reformats the top scoring candidate by rewriting the text thereof as a new HTML page using only feasible HTML tags and attributes and ignoring stylistic deficiencies and unnecessary elements in the original page format.

30. A non-transitory computer-readable medium, according to claim 16 , further comprising:

executable code that clips portions indicated by the user, wherein the portions are subsequently used by other software.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2024
From: EVERNOTE CORPORATION
To: BENDING SPOONS S.P.A.
Reel/Frame 066288/0195 →
RELEASE OF SECURITY INTEREST Recorded Mar 17, 2023
From: MUFG BANK, LTD.
To: EVERNOTE CORPORATION
Reel/Frame 063116/0260 →
RELEASE OF SECURITY INTEREST Recorded Oct 8, 2021
From: EAST WEST BANK
To: EVERNOTE CORPORATION
Reel/Frame 057852/0078 →
SECURITY INTEREST Recorded Oct 6, 2021
From: EVERNOTE CORPORATION
To: MUFG UNION BANK, N.A.
Reel/Frame 057722/0876 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT TERMINATION AT R/F 040192/0720 Recorded Oct 22, 2020
From: SILICON VALLEY BANK
To: EVERNOTE CORPORATION
Reel/Frame 054145/0452 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT TERMINATION AT R/F 040240/0945 Recorded Oct 22, 2020
From: HERCULES CAPITAL, INC.
To: EVERNOTE CORPORATION; EVERNOTE GMBH
Reel/Frame 054213/0234 →
SECURITY INTEREST Recorded Oct 19, 2020
From: EVERNOTE CORPORATION
To: EAST WEST BANK
Reel/Frame 054113/0876 →
SECURITY INTEREST Recorded Oct 5, 2016
From: EVERNOTE CORPORATION; EVERNOTE GMBH
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 040240/0945 →
SECURITY AGREEMENT Recorded Sep 30, 2016
From: EVERNOTE CORPORATION
To: SILICON VALLEY BANK
Reel/Frame 040192/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2012
From: BIGNERT, JAKOB; COARNA, GABRIEL ALEXANDRU
To: EVERNOTE CORPORATION
Reel/Frame 028852/0295 →