IP Library Granted Patent US 7,996,406
Granted Patent B1
US 7,996,406 · App. 12/241,363 · Granted Aug 9, 2011

Method and apparatus for detecting web-based electronic mail in network traffic

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,996,406
App. No.
12/241,363
Granted
Aug 9, 2011
Kind
B1
Abstract

Method and apparatus for detecting web-based electronic mail in network traffic is described. In some examples, web pages are extracted from the network traffic. Fields in each page of a group of the web pages that share a documents structure are identified. A statistical analysis of the fields of each page in the group of web pages is performed to identify any electronic mail (e-mail) fields. The group of web pages is indicated to include web-based e-mail messages if the fields of each page in the group of web pages include at least one e-mail field.

Claims (59)

1. A method of processing network traffic, comprising:

extracting web pages from network traffic;

identifying fields in each page of a group of the web pages that share a document structure, wherein the step of identifying comprises:

executing a document clustering algorithm using the web pages as parametric input to establish the group of web pages;

executing a structure extraction algorithm configured to:

identify structural portions of the pages in the group of web pages;

determine at least one common structural portion in the structural portions as being shared across all pages in the group of web pages, the at least one common structural portion comprising the document structure; and

determine each of the structural portions other than the at least one structural portion to be the fields;

performing a statistical analysis of the fields of each page in the group of web pages to identify any electronic mail (e-mail) fields; and

indicating the group of web pages to include web-based e-mail messages if the fields of each page in the group of web pages include at least one e-mail field.

2. The method of claim 1 , wherein the step of determining the at least one common structural portion comprises:

selecting two pages of the group of web pages as a set;

iteratively determining common text between all pages in the set and adding at least one additional page of the group of web pages to the set until the common text does not change between two successive iterations; and

identifying the common text between all pages in the set as the at least one common structural portion.

3. The method of claim 1 , wherein the step of performing the statistical analysis comprises:

analyzing each of the fields across all of the pages in the group of web pages against at least one statistical requirement.

4. The method of claim 3 , wherein the step of performing the statistical analysis further comprises:

analyzing at least one of the fields across all of the pages in the group of web pages against at least one regular expression requirement.

5. The method of claim 1 , further comprising:

extracting text content from the fields of each page in the group of web pages if the group of web pages is indicated as including web-based e-mail messages.

6. An apparatus for processing network traffic, comprising:

one or more computer processors communicatively coupled to a network wherein the one or more computer processors are configured to:

extract web pages from the network traffic;

identify fields in each page of a group of the web pages that share a document structure, wherein the step of identifying comprises:

executing a document clustering algorithm using the web pages as parametric input to establish the group of web pages;

executing a structure extraction algorithm configured to:

identify structural portions of the pages in the group of web pages;

determine at least one common structural portion in the structural portions as being shared across all pages in the group of web pages, the at least one common structural portion comprising the document structure; and

determine each of the structural portions other than the at least one structural portion to be the fields;

perform a statistical analysis of the fields of each page in the group of web pages to identify any electronic mail (e-mail) fields; and

indicat the group of web pages to include web-based e-mail messages if the fields of each page in the group of web pages include at least one e-mail field.

7. The apparatus of claim 6 , wherein the determining the at least one common structural portion comprises:

selecting two pages of the group of web pages as a set;

iteratively determining common text between all pages in the set and adding at least one additional page of the group of web pages to the set until the common text does not change between two successive iterations; and

identifying the common text between all pages in the set as the at least one common structural portion.

8. The apparatus of claim 6 , wherein the one or more computer processors configured to perform the statistical analysis comprises:

analyzing each of the fields across all of the pages in the group of web pages against at least one statistical requirement.

9. The apparatus of claim 8 , wherein the one or more computer processors configured to perform the statistical analysis further comprises:

analyzing at least one of the fields across all of the pages in the group of web pages against at least one regular expression requirement.

10. The apparatus of claim 6 , wherein the one or more computer processors are configured to:

extract text content from the fields of each page in the group of web pages if the group of web pages is indicated as including web-based e-mail messages.

11. A computer readable medium having instructions stored thereon that when executed by a processor cause the processor to perform a method of processing network traffic, comprising:

extracting web pages from the network traffic;

identifying fields in each page of a group of the web pages that share a document structure, wherein the step of identifying comprises:

executing a document clustering algorithm using the web pages as parametric input to establish the group of web pages;

executing a structure extraction algorithm configured to:

identify structural portions of the pages in the group of web pages;

determine at least one common structural portion in the structural portions as being shared across all pages in the group of web pages, the at least one common structural portion comprising the document structure; and

determine each of the structural portions other than the at least one structural portion to be the fields;

performing a statistical analysis of the fields of each page in the group of web pages to identify any electronic mail (e-mail) fields; and

indicating the group of web pages to include web-based e-mail messages if the fields of each page in the group of web pages include at least one e-mail field.

12. The computer readable medium of claim 11 , wherein the step of determining the at least one common structural portion comprises:

selecting two pages of the group of web pages as a set;

iteratively determining common text between all pages in the set and adding at least one additional page of the group of web pages to the set until the common text does not change between two successive iterations; and

identifying the common text between all pages in the set as the at least one common structural portion.

13. The computer readable medium of claim 11 , wherein the step of performing the statistical analysis comprises:

analyzing each of the fields across all of the pages in the group of web pages against at least one statistical requirement.

14. The computer readable medium of claim 13 , wherein the step of performing the statistical analysis further comprises:

analyzing at least one of the fields across all of the pages in the group of web pages against at least one regular expression requirement.

Assignments (5)
NOTICE OF SUCCESSION OF AGENCY (REEL 050926 / FRAME 0560) Recorded Sep 13, 2022
From: JPMORGAN CHASE BANK, N.A.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 061422/0371 →
SECURITY AGREEMENT Recorded Sep 13, 2022
From: NORTONLIFELOCK INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062220/0001 →
SECURITY AGREEMENT Recorded Nov 4, 2019
From: SYMANTEC CORPORATION; BLUE COAT LLC; LIFELOCK, INC,; SYMANTEC OPERATING CORPORATION
To: JPMORGAN, N.A.
Reel/Frame 050926/0560 →
CHANGE OF ADDRESS Recorded Jun 28, 2011
From: SYMANTEC CORPORATION
To: SYMANTEC CORPORATION
Reel/Frame 026510/0732 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2008
From: RAJAN, BASANT; DALAL, CHIRAG DEEPAK; KABRA, NAVIN
To: SYMANTEC CORPORATION
Reel/Frame 021877/0241 →