IP Library › Granted Patent US 8,005,782
Granted Patent B2
US 8,005,782 · App. 11/837,476 · Granted Aug 23, 2011

Domain name statistical classification using character-based N-grams

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,005,782
App. No.
11/837,476
Filed
Aug 10, 2007
Granted
Aug 23, 2011
Kind
B2
Examiner
CHANG, LI WU
Art Unit
2129
USPC
706/20
Abstract

Systems and methods of classifying domain names are disclosed. Character-based n-grams are derived from a domain name in order to classify such domain name in one or more categories. In one aspect, a geometrical approach is used. Domain name character-based n-grams are mapped to vector points in a multidimensional space. The relationship between a domain name vector point and vector points of other domain names is used as an indicator of the classification of the domain name vector point. In another aspect, a statistical approach is used. Relative frequencies of one or more character-based n-grams in various classifications are used as indicators. Each character-based n-gram can be associated with a respective probability that indicates a likelihood that the character-based n-gram is found in a domain name of a given classification. Such a probability can serve as an estimator of a classification of a new domain name having such character-based n-gram.

Claims (29)

1. A method of classifying a domain name, comprising:

identifying a dictionary set of character-based n-grams, wherein each character-based n-gram in the dictionary set of character-based n-grams is associated with a pre-established first classification probability value and a pre-established second classification probability value, wherein the pre-established first classification probability value is indicative of whether a corresponding character-based n-gram of the dictionary set is likely to be in a first domain name category, wherein the pre-established second classification probability value is indicative of whether the corresponding character-based n-gram of the dictionary set is likely to be in a second domain name category;

identifying a domain set of character-based n-grams corresponding to a domain name string, the domain name string being associated with a domain name;

classifying the domain name in the first domain name category in response to at least one of the pre-established first classification probability value of a first character-based n-gram in the dictionary set corresponding to a first selected character-based n-gram in the domain set or the pre-established first classification probability value of a second character-based n-gram in the dictionary set corresponding to a second selected character-based n-gram in the domain set being higher than a first classification predetermined threshold; and

classifying the domain name in the second domain name category in response to at least one of the pre-established second classification probability value of the first character-based n-gram in the dictionary set corresponding to the first selected character-based n-gram in the domain set or the pre-established second classification probability value of the second character-based n-gram in the dictionary set corresponding to the second selected character-based n-gram in the domain set being higher than a second classification predetermined threshold.

2. The method of claim 1 , further comprising receiving the domain name string as input for classification.

3. The method of claim 1 , wherein the character-based n-grams are bigrams, trigrams, or four-grams.

4. The method of claim 1 , wherein the first domain name category includes domain names corresponding to non-adult websites, and the second domain name category includes domain names corresponding to adult websites.

5. The method of claim 1 , wherein the pre-established first classification probability value is determined with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the first domain name category.

6. The method of claim 5 , wherein the pre-established second classification probability value is determined with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the second domain name category.

7. A system of classifying a domain name, comprising

a configuration module that identifies a dictionary set of character-based n-grams, wherein each character-based n-gram in the dictionary set of character-based n-grams is associated with a pre-established first classification probability value and a pre-established second classification probability value, wherein the pre-established first classification probability value is indicative of whether a corresponding character-based n-gram of the dictionary set is likely to be in a first domain name category, wherein the pre-established second classification probability value is indicative of whether the corresponding character-based n-gram of the dictionary set is likely to be in a second domain name category;

mapping module that identifies a domain set of character-based n-grams corresponding to a domain name string, the domain name string being associated with a domain name; and

a classification module that classifies the domain name in the first domain name category in response to at least one of the pre-established first classification probability value of a first character-based n-gram in the dictionary set corresponding to a first selected character-based n-gram in the domain set or the pre-established first classification probability value of a second character-based n-gram in the dictionary set corresponding to a second selected character-based n-gram in the domain set being higher than a first classification predetermined threshold and that further classifies the domain name in the second domain name category in response to at least one of the pre-established second classification probability value of the first character-based n-gram in the dictionary set corresponding to the first selected character-based n-gram in the domain set or the pre-established second classification probability value of the second character-based n-gram in the dictionary set corresponding to the second selected character-based n-gram in the domain set being higher than a second classification predetermined threshold.

8. The system of claim 7 , wherein the classification module is configured to receive the domain name string as input for classification.

9. The system of claim 7 , wherein the character-based n-grams are bigrams, trigrams, or four-grams.

10. The system of claim 7 , wherein the first domain name category includes domain names corresponding to non-adult websites, and the second domain name category includes domain names corresponding to adult websites.

11. The system of claim 7 , wherein the configuration module is configured to determine the pre-established first classification probability value with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the first domain name category.

12. The system of claim 11 , wherein the configuration module is configured to determine the pre-established second classification probability value with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the second domain name category.

13. A computer program product comprising a computer storage medium having computer program logic stored thereon for enabling a processor-based system to classify a domain name, the computer program product comprising:

a first program logic module for enabling the processor-based system to identify a dictionary set of character-based n-grams, wherein each character-based n-gram in the dictionary set of character-based n-grams is associated with a pre-established first classification probability value and a pre-established second classification probability value, wherein the pre-established first classification probability value is indicative of whether a corresponding character-based n-gram of the dictionary set is likely to be in a first domain name category, wherein the pre-established second classification probability value is indicative of whether the corresponding character-based n-gram of the dictionary set is likely to be in a second domain name category;

a second program logic module for enabling the processor-based system to identify a domain set of character-based n-grams corresponding to a domain name string, the domain name string being associated with a domain name;

a third program logic module for enabling the processor-based system to classify the domain name in the first domain name category in response to at least one of the pre-established first classification probability value of a first character-based n-gram in the dictionary set corresponding to a first selected character-based n-gram in the domain set or the pre-established first classification probability value of a second character-based n-gram in the dictionary set corresponding to a second selected character-based n-gram in the domain set being higher than a first classification predetermined threshold and for enabling the processor-based system to further classify the domain name in the second domain name category in response to at least one of the pre-established second classification probability value of the first character-based n-gram in the dictionary set corresponding to the first selected character-based n-gram in the domain set or the pre-established second classification probability value of the second character-based n-gram in the dictionary set corresponding to the second selected character-based n-gram in the domain set being higher than a second classification predetermined threshold.

14. The computer program product of claim 13 , wherein the second program logic module is for enabling the processor-based system to identify the domain set of character-based n-grams corresponding to the domain name string which is received as input for classification.

15. The computer program product of claim 13 , wherein the character-based n-grams are bigrams, trigrams, or four-grams.

16. The computer program product of claim 13 , wherein the first domain name category includes domain names corresponding to non-adult websites, and the second domain name category includes domain names corresponding to adult websites.

17. The computer program product of claim 13 , further comprising:

a fourth program logic module for enabling the processor-based system to determine the pre-established first classification probability value with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the first domain name category.

18. The computer program product of claim 17 , wherein the fourth program logic module includes logic for enabling the processor-based system to determine the pre-established second classification probability value with a Bayesian formula that uses a number of occurrences of a character-based n-gram and the number of domain name samples in the second domain name category.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2007
From: REZNIK, ILIA; SIMONSON, ROGER N.
To: MICROSOFT CORPORATION
Reel/Frame 020129/0281 →
Continuity (1)
Related Publication 20090043720A1 · Feb 12, 2009