IP Library Granted Patent US 7,565,350
Granted Patent B2
US 7,565,350 · App. 11/471,403 · Granted Jul 21, 2009

Identifying a web page as belonging to a blog

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,565,350
App. No.
11/471,403
Granted
Jul 21, 2009
Kind
B2
Abstract

A machine learning classifier is used to determine whether a web page belongs to a blog, based on a number of characteristics of web pages (e.g., presence of words such as “permalink”, or being hosted on a known blogging site). The classifier may be initially trained using human-judged examples. After classifying web pages as being blog pages, the blog pages may be further identified or categorized as top level blogs based on their URLs, for example.

Claims (27)

1. A data storage medium for web page classification system, comprising:

a web crawler for crawling a corpus of web pages;

a feature extractor for extracting at least one of the following features from a web page received from the web crawler: a first uniform resource locator (URL) corresponding to the hosting site of the web page, a second URL contained inside the web page that is indicative of a hyperlink to a blog site, at least one substring that is a part of the first URL, and whether the web page contains an ATOM or RSS feed, the feature extractor further configured for extracting the contents of the web page and generating therefrom a set of observed values, wherein each observed value is associated with a feature in the web page that provides an indication that the web page is a blog page, the set of observed values including a first observed value that is generated based on the number of occurrences in the web page of a non-markup word indicative of a blog; and

a machine learning classifier communicatively coupled to the feature extractor for evaluating the extracted features and generating a prediction indicating the probability that the web page is a blog page, the classifier containing an algorithm that is trained to apply i) a heavier classifier weight to the first URL corresponding to the hosting site of the web page than to the second URL contained inside the web page, and ii) a heavier classifier weight to the second URL than the substring that is a part of the first URL.

2. The data storage medium of claim 1 , wherein the extracted feature further comprises at least one phrase contained in the web page in combination with at least one of the other extracted features.

3. The data storage medium of claim 1 , wherein each of the set of observed values generated by the feature extractor is a Boolean value.

4. The data storage medium of claim 1 , wherein the classifier is configured to be initially trained using human-judged examples.

5. The data storage medium of claim 1 , wherein the non-markup word is at least one of the following words: i) “blog”, ii) “permalink”, iii) “comment”, iv) “posted”, and v) “track back”.

6. The data storage medium of claim 1 , wherein the classifier is configured to be language independent, and wherein non-English equivalents of the non-markup word are counted for generating the first observed value.

7. The data storage medium of claim 1 , wherein the feature extractor is configured to generate the set of observed values from a set of features that art selected based on the configuration of the classifier.

8. The data storage medium of claim 1 , wherein the set of observed values further includes a second observed value that is generated based on the number of occurrences in the web page of a phrase indicative of a blog.

9. A web page classification method, comprising:

crawling, via a processor, a corpus of web pages and providing the web page to a feature extractor;

extracting at least one feature, from the received web page using a feature extractor, wherein the at least one feature comprises one or more of the following: a count of the number of occurrences of a blog-related word in the web page, a first uniform resource locator (URL) corresponding to the host of the web page, a second URL contained inside the web page that is indicative of a hyperlink to a blog site, at least one substring that is a part of the first URL, and whether the web page contains an ATOM feed or an RSS feed, the feature extractor further configured for extracting the contents of the web page and generating therefrom a set of observed values, wherein each observed value is associated with a feature in the web rage that provides an indication that the web page is a blog page, the set of observed values including a first observed value that is generated based on the number of occurrences in the web page of a non-markup word indicative of a blog; and

classifying the web page as being a blog page or not based on an evaluation of the at least one extracted feature, the evaluation comprising application of i) a heavier classifier weight to the first URL than to the second URL contained inside the web page, and ii) a heavier classifier weight to the second URL than the substring that is a part of the first URL.

10. The method of claim 9 , further comprising extracting at least one wherein the at least one feature further comprises combining a phrase contained in the web page with at least one of the other extracted features.

11. The method of claim 9 , wherein classifying the web page comprises providing an indication, prediction, or probability that the web page is a blog page or not.

12. The method of claim 9 , further comprising:

forming a set of web pages that are classified as being a blog page; and

identifying a top level blog in the set of web pages.

13. A web page classification method, comprising:

classifying, via a processor, a plurality of web pages, each as being a blog page or not based on at least one extracted feature that comprises a count of the number of occurrences of a blog-related word in the web page, a first uniform resource locator (URL) corresponding to the host of the web page, a second URL contained inside the web page that is indicative of a hyperlink to a blog site, at least one substring that is a part of the first URL, and whether the web page contains an ATOM feed or an RSS feed, the feature extractor further configured for extracting the contents of the web page and generating therefrom a set of observed values, wherein each observed value is associated with a feature in the web page that provides an indication that the web page is a blog page, the set of observed values including a first observed value that is generated based on the number of occurrences in the web page of a non-markup word indicative of a blog, the classifying comprising application of i) a heavier classifier weight to the first URL than to the second URL contained inside the web page, and ii) a heavier classifier weight to the second URL than the substring that is a part of the first URL;

forming a set of web pages that are classified as being a blog page; and

identifying a top level blog in the set of web pages.

14. The method of claim 13 , further comprising lexigraphically sorting the uniform resource locators (URLs) of each of the web pages in the set.

15. The method of claim 14 , wherein identifying the top level blog comprises iterating through the lexigraphically sorted URLs to determine a common prefix of the web pages.

16. The method of claim 13 , wherein the at least one extracted feature further comprises at least one weid-ef phrase contained in the web page in combination with any of the other extracted features.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE INVENTOR'S NAME PREVIOUSLY RECORDED ON REEL 018157 FRAME 0441. ASSIGNOR(S) HEREBY CONFIRMS THE DENNIS CRAIG SHOULD BE DENNIS CRAIG FETTERLY. Recorded Sep 9, 2006
From: FETTERLY, DENNIS CRAIG; CHIEN, STEVE SHAW-TANG
To: MICROSOFT CORPORATION
Reel/Frame 018230/0297 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2006
From: CRAIG, DENNIS; CHIEN, STEVE SHAW-TANG
To: MICROSOFT CORPORATION
Reel/Frame 018157/0441 →