IP Library Granted Patent US 8,973,678
Granted Patent B1
US 8,973,678 · App. 11/562,941 · Granted Mar 10, 2015

Misspelled word analysis for undesirable message classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,973,678
App. No.
11/562,941
Granted
Mar 10, 2015
Kind
B1
Abstract

Misspelled words are identified in incoming email messages. The presence of misspelled words in emails is used to help determine which the emails are spam. Various statistical information concerning the number, prevalence, distribution, etc. of misspelled words in email messages is analyzed to detect spam or other forms of undesirable email, such as phishing emails. In some embodiments, the language in which an email is written is identified in order to aid in the identification of misspelled words. In some embodiments, the analysis of the misspelling information is combined with other techniques used to identify undesirable email.

Claims (58)

1. A computer implemented method for identifying undesirable electronic messages from specific misspellings, the method comprising the steps of:

identifying electronic messages that have been categorized as spam;

identifying specific misspelled words within content of the spam electronic messages;

collecting misspelling statistical metrics for the content of the spam electronic messages based on the identified specific misspelled words; and

for an incoming electronic message, analyzing misspelling metrics to determine if the incoming electronic message is an undesirable electronic message by matching the statistical metrics of spam electronic messages based on the specific misspelled words to misspelled words within the incoming electronic message,

wherein each step of the method is performed by a computer.

2. The method of claim 1 further comprising:

identifying a language in which an identified electronic message is composed, prior to identifying misspelled words in the identified electronic message; and

taking into account the identified language in identifying misspelled words in the identified electronic message.

3. The method of claim 2 further comprising:

using an identified language specific electronic dictionary to identify misspelled words in the identified electronic message.

4. The method of claim 1 wherein analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises:

performing a statistical analysis of electronic message content in order to identify statistical patterns associated with undesirable electronic messages, using misspelling metrics as a factor in the statistical analysis.

5. The method of claim 1 wherein analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises:

performing a heuristic analysis of electronic message content in order to identify undesirable electronic messages, using misspelling metrics as a heuristic in the heuristic analysis.

6. The method of claim 1 wherein analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises:

using data concerning a relationship between at least one misspelling metric and at least one threshold value as evidence in the identification of an electronic message as undesirable.

7. The method of claim 1 further comprising determining that at least one received electronic message comprises a type of undesirable electronic message from a group of undesirable electronic message types consisting of:

a spam email; and

a phishing email.

8. At least one non-transitory computer readable medium containing a computer program product for identifying undesirable electronic messages from specific misspellings, the computer program product comprising program code to:

to identify electronic messages that have been categorized as spam;

to identify specific misspelled words within content of the spam electronic messages;

to collect misspelling statistical metrics for the content of the spam electronic messages based on the identified specific misspelled words; and

for an incoming electronic message, to analyze misspelling metrics to determine if the incoming electronic message is an undesirable electronic message by matching the statistical metrics of spam electronic messages based on the specific misspelled words to misspelled words within the incoming electronic message.

9. The computer program product of claim 8 further comprising program code to:

identify a language in which an identified electronic message is composed, prior to identifying misspelled words in the identified electronic message; and

take into account the identified language in identifying misspelled words in the identified electronic message.

10. The computer program product of claim 9 further comprising program code to:

use an identified language specific electronic dictionary to identify misspelled words in the identified electronic message.

11. The computer program product of claim 8 wherein analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises program code to:

perform a statistical analysis of electronic message content in order to identify statistical patterns associated with undesirable electronic messages, using misspelling metrics as a factor in the statistical analysis.

12. The computer program product of claim 8 wherein the analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises program code to:

perform a heuristic analysis of electronic message content in order to identify undesirable electronic messages, using misspelling metrics as a heuristic in the heuristic analysis.

13. The computer program product of claim 8 wherein analyzing misspelling metrics for identified electronic messages to identify undesirable electronic messages further comprises program code to:

use data concerning a relationship between at least one misspelling metric and at least one threshold value as evidence in the identification of an electronic message as undesirable.

14. The computer program product of claim 8 further comprising program code to:

determine that at least one received electronic message comprises a type of undesirable electronic message from a group of undesirable electronic message types consisting of:

a spam email; and

a phishing email.

15. A computer system for identifying undesirable electronic messages from specific misspellings, the computer system comprising:

a processor;

a computer memory, comprising:

a first module to identify electronic messages that have been categorized as spam;

a second module to identify specific misspelled words within content of the spam electronic messages;

a third module to collect misspelling statistical metrics for the content of the spam electronic messages based on the identified specific misspelled words; and

a fourth module to for an incoming electronic message, analyze misspelling metrics to determine if the incoming electronic message is an undesirable electronic message by matching the statistical metrics of spam electronic messages based on the specific misspelled words to misspelled words within the incoming electronic message.

16. The computer system of claim 15 wherein the memory further comprises:

a fifth module to identify a language in which an identified electronic message is composed, prior to identifying misspelled words in the identified electronic message; and

an executable image stored in the computer memory configured to take into account the identified language in identifying misspelled words in the identified electronic message.

17. The computer system of claim 16 wherein the memory further comprises:

a sixth module to use an identified language specific electronic dictionary to identify misspelled words in the identified electronic message.

18. The computer system of claim 15 wherein the fourth module is further configured to:

perform a statistical analysis of electronic message content in order to identify statistical patterns associated with undesirable electronic messages, using misspelling metrics as a factor in the statistical analysis.

19. The computer system of claim 15 wherein the fourth module is further configured to:

perform a heuristic analysis of electronic message content in order to identify undesirable electronic messages, using misspelling metrics as a heuristic in the heuristic analysis.

20. The computer system of claim 15 wherein the fourth module is further configured to:

use data concerning a relationship between at least one misspelling metric and at least one threshold value as evidence in the identification of an electronic message as undesirable.

Assignments (5)
NOTICE OF SUCCESSION OF AGENCY (REEL 050926 / FRAME 0560) Recorded Sep 13, 2022
From: JPMORGAN CHASE BANK, N.A.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 061422/0371 →
SECURITY AGREEMENT Recorded Sep 13, 2022
From: NORTONLIFELOCK INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062220/0001 →
CHANGE OF NAME Recorded Jun 18, 2020
From: SYMANTEC CORPORATION
To: NORTONLIFELOCK INC.
Reel/Frame 053306/0878 →
SECURITY AGREEMENT Recorded Nov 4, 2019
From: SYMANTEC CORPORATION; BLUE COAT LLC; LIFELOCK, INC,; SYMANTEC OPERATING CORPORATION
To: JPMORGAN, N.A.
Reel/Frame 050926/0560 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2006
From: COOLEY, SHAUN; MCCORKENDALE, BRUCE
To: SYMANTEC CORPORATION
Reel/Frame 018683/0616 →