IP Library Granted Patent US 7,792,846
Granted Patent B1
US 7,792,846 · App. 11/881,770 · Granted Sep 7, 2010

Training procedure for N-gram-based statistical content classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,792,846
App. No.
11/881,770
Granted
Sep 7, 2010
Kind
B1
Abstract

A training procedure for N-gram based statistical document classification has been disclosed. In one embodiment, a set of N-grams is selected out of a second set of N-grams, each of the N-grams having a sequence of N bytes, where N is an integer. Then a statistical content classification model is generated based on occurrences of the N-grams, if any, in a set of training documents and a set of validation documents. The statistical content classification model is provided to content filters to classify content.

Claims (71)

1. A computer-implemented method, comprising:

selecting a plurality of N-grams from a second plurality of N-grams, wherein the second plurality of N-grams are associated with a range of values of N and the plurality of N-grams are associated with a sub-range of the range of values of N, wherein each of the second plurality of N-grams comprises a sequence of N bytes, where N is an integer;

generating a statistical content classification model based on occurrences of the plurality of N-grams, if any, in a set of training documents and a set of validation documents;

providing the statistical content classification model to content filters to classify content into one or more of a plurality of categories;

searching for the plurality of N-grams in the set of training documents;

computing a plurality of scores for each of the plurality of N-grams with respect to the plurality of categories;

searching for the plurality of N-grams in the set of validation documents; and

determining a threshold for each of the plurality of categories.

2. The method of claim 1 , wherein determining the threshold for each of the plurality of categories comprises:

computing the threshold using a frequency of occurrence of each of the plurality of N-grams in a subset of validation documents of the set of validation documents that have been classified into a respective category, a frequency of non-occurrence of each of the plurality of N-grams in the subset of validation documents, the plurality of scores, and a predetermined false positive limit.

3. The method of claim 1 , wherein each of said plurality of N-grams representing at least a portion of a keyword in a non-delimited natural language.

4. The method of claim 3 , wherein the non-delimited natural language is Chinese.

5. The method of claim 1 , wherein the set of training documents includes one or more web pages.

6. The method of claim 1 , wherein the set of training documents includes one or more electronic mail messages.

7. A computer-implemented method, comprising:

selecting a plurality of N-grams from a second plurality of N-grams, wherein the second plurality of N-grams are associated with a range of values of N and the plurality of N-grams are associated with a sub-range of the range of values of N, wherein each of the second plurality of N-grams comprises a sequence of N bytes, where N is an integer;

generating a statistical content classification model based on occurrences of the plurality of N-grams, if any, in a set of training documents and a set of validation documents;

providing the statistical content classification model to content filters to classify content into one or more of a plurality of categories;

determining a utility for each of the second plurality of N-grams using a frequency of occurrence of a respective N-gram in a subset of training documents of the set of training documents that have been classified in a respective category and a frequency of occurrence of the respective N-gram in remaining training documents of the set of training documents; and

selecting the sub-range of values of N based on utilities of the second plurality of N-grams.

8. The method of claim 7 , wherein each of said plurality of N-grams representing at least a portion of a keyword in a non-delimited natural language.

9. The method of claim 8 , wherein the non-delimited natural language is Chinese.

10. The method of claim 7 , wherein the set of training documents includes one or more web pages.

11. The method of claim 7 , wherein the set of training documents includes one or more electronic mail messages.

12. A machine-accessible medium that provides instructions that, if executed by a processor, will cause the processor to perform operations comprising:

selecting a plurality of N-grams from a second plurality of N-grams, wherein the second plurality of N-grams are associated with a range of values of N and the plurality of N-grams are associated with a sub-range of the range of values of N, wherein each of the second plurality of N-grams comprises a sequence of N bytes, where N is an integer;

generating a statistical content classification model based on occurrences of the plurality of N-grams, if any, in a set of training documents and a set of validation documents;

providing the statistical content classification model to content filters to classify content into one or more of a plurality of categories;

searching for the plurality of N-grams in the set of training documents;

computing a plurality of scores for each of the plurality of N-grams with respect to the plurality of categories;

searching for the plurality of N-grams in the set of validation documents; and

determining a threshold for each of the plurality of categories.

13. The machine-accessible medium of claim 12 , wherein determining the threshold for each of the plurality of categories comprises:

computing the threshold using a frequency of occurrence of each of the plurality of N-grams in a subset of validation documents of the set of validation documents that have been classified into a respective category, a frequency of non-occurrence of each of the plurality of N-grams in the subset of validation documents, the plurality of scores, and a predetermined false positive limit.

14. The machine-accessible medium of claim 12 , wherein each of said plurality of N-grams representing at least a portion of a keyword in a non-delimited natural language.

15. The machine-accessible medium of claim 14 , wherein the non-delimited natural language is Chinese.

16. The machine-accessible medium of claim 12 , wherein the set of training documents includes one or more web pages.

17. The machine-accessible medium of claim 12 , wherein the set of training documents includes one or more electronic mail messages.

18. A machine-accessible medium that provides instructions that, if executed by a processor, will cause the processor to perform operations comprising:

selecting a plurality of N-grams from a second plurality of N-grams, wherein the second plurality of N-grams are associated with a range of values of N and the plurality of N-grams are associated with a sub-range of the range of values of N, wherein each of the second plurality of N-grams comprises a sequence of N bytes, where N is an integer;

generating a statistical content classification model based on occurrences of the plurality of N-grams, if any, in a set of training documents and a set of validation documents;

providing the statistical content classification model to content filters to classify content into one or more of a plurality of categories;

determining a utility for each of the second plurality of N-grams using a frequency of occurrence of a respective N-gram in a subset of training documents of the set of training documents that have been classified in a respective category and a frequency of occurrence of the respective N-gram in remaining training documents of the set of training documents; and

selecting the sub-range of values of N based on utilities of the second plurality of N-grams.

19. The machine-accessible medium of claim 18 , wherein each of said plurality of N-grams representing at least a portion of a keyword in a non-delimited natural language.

20. The machine-accessible medium of claim 19 , wherein the non-delimited natural language is Chinese.

21. The machine-accessible medium of claim 18 , wherein the set of training documents includes one or more web pages.

22. The machine-accessible medium of claim 18 , wherein the set of training documents includes one or more electronic mail messages.

23. An apparatus comprising:

a pattern matching engine to search for a plurality of N-grams in a set of training documents and a set of validation documents, each of said plurality of N-grams representing at least a portion of a keyword in a natural language, and the set of training documents and the set of validation documents being written in the natural language, wherein each of said plurality of N-grams comprises a sequence of N bytes, where N is an integer; and

a model generator coupled to the search engine to generate a statistical content classification model based on occurrences of each of the plurality of N-grams in the set of training documents and the set of validation documents,

wherein the search engine is operable to compute a plurality of scores for each of the plurality of N-grams with respect to a plurality of categories;

wherein the model generator is operable to determine a plurality of thresholds for the plurality of categories using the plurality of scores and the set of validation documents, each of the plurality of thresholds being associated with a distinct one of the plurality of categories; wherein the model generator is operable to compute each of the plurality of thresholds using a frequency of occurrences of each of the plurality of N-grams in the set of validation documents, the plurality of scores, and a predetermined false positive limit.

24. The apparatus of claim 23 , further comprising:

a processing module to select the plurality of N-grams from a second plurality of N-grams based on utilities of the plurality of N-grams in content classification with respect to the set of training documents.

25. The apparatus of claim 23 , wherein the model generator is operable to determine a plurality of thresholds for the plurality of categories using the plurality of scores and the set of validation documents, each of the plurality of thresholds being associated with a distinct one of the plurality of categories.

26. The apparatus of claim 25 , wherein the model generator is operable to compute each of the plurality of thresholds using a frequency of occurrences of each of the plurality of N-grams in the set of validation documents, the plurality of scores, and a predetermined false positive limit.

27. A system comprising:

a pattern matching engine to search for a plurality of N-grams in a set of training documents and a set of validation documents, each of said plurality of N-grams representing at least a portion of a keyword in a natural language, and the set of training documents and the set of validation documents being written in the natural language, wherein each of said plurality of N-grams comprises a sequence of N bytes, where N is an integer;

a model generator coupled to the search engine to generate a statistical content classification model based on occurrences of each of the plurality of N-grams in the set of training documents and the set of validation documents;

a repository coupled to the model generator to store the statistical content classification model;

an N-gram-based content rating engine coupled to the repository, to access the statistical content classification model and to rate content of documents in the natural language using the statistical content classification model, wherein the documents are from a network external to the system;

a content filtering module comprising the N-gram-based content rating engine; and

a client machine coupled to the content filtering module, wherein the content filtering module receives a request to access a web page from the client machine and the N-gram-based content rating engine rates content of the requested web page, wherein the content filtering module blocks the requested web page from the client machine if the content of the requested web page is in a prohibited category and the content filtering module passes the requested web page to the client machine if the content of the requested web page is in an allowable category.

28. A system comprising:

a pattern matching engine to search for a plurality of N-grams in a set of training documents and a set of validation documents, each of said plurality of N-grams representing at least a portion of a keyword in a natural language, and the set of training documents and the set of validation documents being written in the natural language, wherein each of said plurality of N-grams comprises a sequence of N bytes, where N is an integer;

a model generator coupled to the search engine to generate a statistical content classification model based on occurrences of each of the plurality of N-grams in the set of training documents and the set of validation documents;

a repository coupled to the model generator to store the statistical content classification model;

an N-gram-based content rating engine coupled to the repository, to access the statistical content classification model and to rate content of documents in the natural language using the statistical content classification model, wherein the documents are from a network external to the system;

a content filtering module comprising the N-gram-based content rating engine; and

a client machine coupled to the content filtering module, wherein the content filtering module receives an incoming electronic mail message and the N-gram-based content rating engine rates content of the electronic mail message, wherein the content filtering module blocks the electronic mail message from the client machine if the content of the electronic mail message is in a prohibited category and the content filtering module passes the electronic mail message to the client machine if the content of the electronic mail message is in an allowable category.

Assignments (23)
RELEASE OF SECOND LIEN SECURITY INTEREST IN PATENTS RECORDED AT RF 046321/0393 Recorded Jun 16, 2025
From: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
To: SONICWALL US HOLDINGS INC.
Reel/Frame 071625/0887 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Jun 7, 2018
From: SONICWALL US HOLDINGS INC.
To: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
Reel/Frame 046321/0393 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Jun 7, 2018
From: SONICWALL US HOLDINGS INC.
To: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
Reel/Frame 046321/0414 →
RELEASE OF FIRST LIEN SECURITY INTEREST IN PATENTS RECORDED AT R/F 040581/0850 Recorded May 22, 2018
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: QUEST SOFTWARE INC. (F/K/A DELL SOFTWARE INC.); AVENTAIL LLC
Reel/Frame 046211/0735 →
CHANGE OF NAME Recorded Apr 30, 2018
From: DELL SOFTWARE INC.
To: QUEST SOFTWARE INC.
Reel/Frame 046040/0277 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE PREVIOUSLY RECORDED AT REEL: 040587 FRAME: 0624. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 28, 2017
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: QUEST SOFTWARE INC. (F/K/A DELL SOFTWARE INC.); AVENTAIL LLC
Reel/Frame 044811/0598 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE NATURE OF CONVEYANCE PREVIOUSLY RECORDED AT REEL: 041073 FRAME: 0001. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY ASSIGNMENT.. Recorded Apr 5, 2017
From: QUEST SOFTWARE INC.
To: SONICWALL US HOLDINGS INC.
Reel/Frame 042168/0114 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jan 23, 2017
From: QUEST SOFTWARE INC.
To: SONICWALL US HOLDINGS, INC.
Reel/Frame 041073/0001 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Nov 10, 2016
From: DELL SOFTWARE INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040587/0624 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Nov 9, 2016
From: DELL SOFTWARE INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040581/0850 →
RELEASE OF SECURITY INTEREST Recorded Oct 31, 2016
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: AVENTAIL LLC; DELL PRODUCTS, L.P.; DELL SOFTWARE INC.
Reel/Frame 040521/0467 →
RELEASE OF SECURITY INTEREST IN CERTAIN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (040039/0642) Recorded Oct 31, 2016
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
To: AVENTAIL LLC; DELL PRODUCTS L.P.; DELL SOFTWARE INC.
Reel/Frame 040521/0016 →
SECURITY AGREEMENT Recorded Sep 14, 2016
From: AVENTAIL LLC; DELL PRODUCTS L.P.; DELL SOFTWARE INC.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 040039/0642 →
SECURITY AGREEMENT Recorded Sep 14, 2016
From: AVENTAIL LLC; DELL PRODUCTS, L.P.; DELL SOFTWARE INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040030/0187 →
MERGER Recorded Dec 15, 2015
From: SONICWALL L.L.C.
To: DELL SOFTWARE INC.
Reel/Frame 037294/0290 →
CONVERSION AND NAME CHANGE Recorded Dec 15, 2015
From: SONICWALL, INC.
To: SONICWALL L.L.C.
Reel/Frame 037301/0765 →
RELEASE OF SECURITY INTEREST IN PATENTS RECORDED ON REEL/FRAME 024776/0337 Recorded May 8, 2012
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: AVENTAIL LLC; SONICWALL, INC.
Reel/Frame 028177/0115 →
RELEASE OF SECURITY INTEREST IN PATENTS RECORDED ON REEL/FRAME 024823/0280 Recorded May 8, 2012
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: AVENTAIL LLC; SONICWALL, INC.
Reel/Frame 028177/0126 →
PATENT SECURITY AGREEMENT (SECOND LIEN) Recorded Aug 3, 2010
From: AVENTAIL LLC; SONICWALL, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 024823/0280 →
SECURITY AGREEMENT Recorded Aug 3, 2010
From: AVENTAIL LLC; SONICWALL, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 024776/0337 →
MERGER Recorded Jul 28, 2010
From: SONICWALL, INC.
To: PSM MERGER SUB (DELAWARE), INC.
Reel/Frame 024755/0083 →
CHANGE OF NAME Recorded Jul 28, 2010
From: PSM MERGER SUB (DELAWARE), INC.
To: SONICWALL, INC.
Reel/Frame 024755/0091 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2007
From: RAFFILL, THOMAS E.; ZHU, SHUNHUI; YANOVSKY, ROMAN; YANOVSKY, BORIS; GMUENDER, JOHN
To: SONICWALL, INC.
Reel/Frame 020152/0521 →