IP Library Granted Patent US 8,214,437
Granted Patent B1
US 8,214,437 · App. 10/743,015 · Granted Jul 3, 2012

Online adaptive filtering of messages

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,214,437
App. No.
10/743,015
Granted
Jul 3, 2012
Kind
B1
Abstract

In general, a two or more stage spam filtering system is used to filter spam in an e-mail system. One stage includes a global e-mail classifier that classifies e-mail as it enters the e-mail system. The parameters of the global e-mail classifier generally may be determined by the policies of e-mail system owner and generally are set to only classify as spam those e-mails that are likely to be considered spam by a significant number of users of the e-mail system. Another stage includes personal e-mail classifiers at the individual mailboxes of the e-mail system users. The parameters of the personal e-mail classifiers generally are set by the users through retraining, such that the personal e-mail classifiers are refined to track the subjective perceptions of their respective user as to what e-mails are spam e-mails. Retraining data for the personal e-mail classifiers may be aggregated and a subset of the aggregate may be chosen for use in retraining the global e-mail classifier.

Claims (63)

1. A method of handling messages in a messaging system that includes a message gateway and individual message boxes for users of the system, wherein a message addressed to a user is delivered to the user's message box after passing through the message gateway, the method comprising:

knowingly biasing a global, scoring e-mail classifier relative to a personal, scoring e-mail classifier such that the global, scoring e-mail classifier is less stringent than the personal, scoring e-mail classifier as to what is classified as spam, wherein the global, scoring e-mail classifier and the personal, scoring e-mail classifier are probabilistic e-mail classifiers such that, to classify a message, the global, scoring e-mail classifier and the personal, scoring e-mail classifier use respective internal models to determine a probability measure for the message and compare the probability measure to a classification threshold;

receiving messages at the message gateway;

inputting the received messages into the global, scoring e-mail classifier to classify the input messages as spam or non-spam;

handling at least one of the messages input into the global, scoring e-mail classifier based on whether the global, scoring e-mail classifier classified the at least one message as spam or non-spam;

outputting at least one message from the global, scoring e-mail classifier, wherein the outputted message has been classified as non-spam by the global, scoring e-mail classifier;

inputting the outputted message from the global, scoring e-mail classifier into the personal, scoring e-mail classifier to classify the at least one outputted message as spam or non-spam;

handling the at least one outputted message input into the personal, scoring e-mail classifier based on whether the personal, scoring e-mail classifier classified the at least one outputted message as spam or non-spam;

receiving an indication from a user to change the classification of the at least one outputted message;

in response to the indication, changing the classification of the at least one outputted message;

generating retraining data based on the change to the classification of the at least one outputted message; and

retraining the personal, scoring e-mail classifier based on the generated retraining data such that the personal, scoring e-mail classifier's internal model is refined to track the user's subjective perceptions as to what messages constitute spam messages.

2. The method according to claim 1 , further comprising training the global, scoring e-mail classifier using a training set of messages to develop the internal model of the global, scoring e-mail classifier.

3. The method according to claim 2 , wherein the training set of messages comprises messages that are known to be spam messages to a significant number of users of the messaging system.

4. The method according to claim 3 , further comprising collecting the training set of messages through feedback from the users of the messaging system.

5. The method according to claim 1 , wherein to knowingly bias the global, scoring e-mail classifier related to the personal, scoring e-mail classifier, the global, scoring e-mail classifier is trained based on higher misclassification costs than the personal, scoring e-mail classifier.

6. The method according to claim 1 , wherein the messages are at least one of e-mails, instant messages, or SMS messages.

7. The method according to claim 1 , wherein the global, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

8. The method according to claim 1 , wherein the personal, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

9. A non-transitory computer-readable medium including a set of instructions which, when executed by a processor, performs a method of handling messages in a messaging system that includes a message gateway and individual message boxes for users of the system, wherein a message addressed to a user is delivered to the user's message box after passing through the message gateway, the method comprising:

knowingly biasing a global, scoring e-mail classifier relative to a personal, scoring e-mail classifier such that the global, scoring e-mail classifier is less stringent than the personal, scoring e-mail classifier as to what is classified as spam, wherein the global, scoring e-mail classifier and the personal, scoring e-mail classifier are probabilistic e-mail classifiers such that, to classify a message, the global, scoring e-mail classifier and the personal, scoring e-mail classifier use respective internal models to determine a probability measure for the message and compare the probability measure to a classification threshold;

receiving messages at the message gateway;

inputting the received messages into the global, scoring e-mail classifier to classify the input messages as spam or non-spam;

handling at least one of the messages input into the global, scoring e-mail classifier based on whether the global, scoring e-mail classifier classified the at least one message as spam or non-spam;

outputting at least one message from the global, scoring e-mail classifier, wherein the outputted message has been classified as non-spam by the global, scoring e-mail classifier;

inputting the outputted message from the global, scoring e-mail classifier into the personal, scoring e-mail classifier to classify the at least one outputted message as spam or non-spam;

handling the at least one outputted message input into the personal, scoring e-mail classifier based on whether the personal, scoring e-mail classifier classified the at least one outputted message as spam or non-spam;

receiving an indication from a user to change the classification of the at least one outputted message;

in response to the indication, changing the classification of the at least one outputted message;

generating retraining data based on the change to the classification of the at least one outputted message; and

retraining the personal, scoring e-mail classifier based on the generated retraining data such that the personal, scoring e-mail classifier's internal model is refined to track the user's subjective perceptions as to what messages constitute spam messages.

10. The non-transitory computer-readable medium according to claim 9 , further comprising instructions executable by the processor to perform a method including:

training the global, scoring e-mail classifier using a training set of messages to develop the internal model of the global, scoring e-mail classifier.

11. The non-transitory computer-readable medium according to claim 10 , wherein the training set of messages comprises messages that are known to be spam messages to a significant number of users of the messaging system.

12. The non-transitory computer-readable medium according to claim 11 , further comprising instructions executable by the processor to perform a method including:

collecting the training set of messages through feedback from the users of the messaging system.

13. The non-transitory computer-readable medium according to claim 9 , wherein to knowingly bias the global, scoring e-mail classifier related to the personal, scoring e-mail classifier, the global, scoring e-mail classifier is trained based on higher misclassification costs than the personal, scoring e-mail classifier.

14. The non-transitory computer-readable medium according to claim 9 , wherein the messages are at least one of e-mails, instant messages, or SMS messages.

15. The non-transitory computer-readable medium according to claim 9 , wherein the global, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

16. The non-transitory computer-readable medium according to claim 9 , wherein the personal, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

17. A system for handling messages in a messaging system that includes a message gateway and individual message boxes for users of the system, wherein a message addressed to a user is delivered to the user's message box after passing through the message gateway, the system comprising:

a storage medium that stores a set of instructions;

at least one processor that executes the set of instructions to perform a method, the method comprising:

knowingly biasing a global, scoring e-mail classifier relative to a personal, scoring e-mail classifier such that the global, scoring e-mail classifier is less stringent than the personal, scoring e-mail classifier as to what is classified as spam, wherein the global, scoring e-mail classifier and the personal, scoring e-mail classifier are probabilistic e-mail classifiers such that, to classify a message, the global, scoring e-mail classifier and the personal, scoring e-mail classifier use respective internal models to determine a probability measure for the message and compare the probability measure to a classification threshold;

receiving messages at the message gateway;

inputting the received messages into the global, scoring e-mail classifier to classify the input messages as spam or non-spam;

handling at least one of the messages input into the global, scoring e-mail classifier based on whether the global, scoring e-mail classifier classified the at least one message as spam or non-spam;

outputting at least one message from the global, scoring e-mail classifier, wherein the outputted message has been classified as non-spam by the global, scoring e-mail classifier;

inputting the outputted message from the global, scoring e-mail classifier into the personal, scoring e-mail classifier to classify the at least one outputted message as spam or non-spam;

handling the at least one outputted message input into the personal, scoring e-mail classifier based on whether the personal, scoring e-mail classifier classified the at least one outputted message as spam or non-spam;

receiving an indication from a user to change the classification of the at least one outputted message;

in response to the indication, changing the classification of the at least one outputted message;

generating retraining data based on the change to the classification of the at least one outputted message; and

retraining the personal, scoring e-mail classifier based on the generated retraining data such that the personal, scoring e-mail classifier's internal model is refined to track the user's subjective perceptions as to what messages constitute spam messages.

18. The system according to claim 17 , wherein the at least one processor further executes the set of instructions to perform a method including:

training the global, scoring e-mail classifier using a training set of messages to develop the internal model of the global, scoring e-mail classifier.

19. The system according to claim 18 , wherein the training set of messages comprises messages that are known to be spam messages to a significant number of users of the messaging system.

20. The system according to claim 19 , wherein the at least one processor further executes the set of instructions to perform a method including:

collecting the training set of messages through feedback from the users of the messaging system.

21. The system according to claim 17 , wherein to knowingly bias the global, scoring e-mail classifier related to the personal, scoring e-mail classifier, the global, scoring e-mail classifier is trained based on higher misclassification costs than the personal, scoring e-mail classifier.

22. The system according to claim 17 , wherein the messages are at least one of e-mails, instant messages, or SMS messages.

23. The system according to claim 17 , wherein the global, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

24. The system according to claim 17 , wherein the personal, scoring e-mail classifier is configured such that classifying messages as spam or non-spam comprises classifying messages into subcategories of spam or non-spam.

Assignments (12)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
CHANGE OF NAME Recorded Aug 24, 2017
From: AOL INC.
To: OATH INC.
Reel/Frame 043672/0369 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS -RELEASE OF 030936/0011 Recorded Jul 1, 2015
From: JPMORGAN CHASE BANK, N.A.
To: AOL ADVERTISING INC.; AOL INC.; BUYSIGHT, INC.; MAPQUEST, INC.; PICTELA, INC.
Reel/Frame 036042/0053 →
SECURITY AGREEMENT Recorded Aug 2, 2013
From: AOL INC.; AOL ADVERTISING INC.; BUYSIGHT, INC.; MAPQUEST, INC.; PICTELA, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 030936/0011 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Nov 16, 2010
From: BANK OF AMERICA, N A
To: AOL INC; AOL ADVERTISING INC; GOING INC; LIGHTNINGCAST LLC; MAPQUEST, INC; NETSCAPE COMMUNICATIONS CORPORATION; QUIGO TECHNOLOGIES LLC; SPHERE SOURCE, INC; TACODA LLC; TRUVEO, INC; YEDDA, INC
Reel/Frame 025323/0416 →
CHANGE OF NAME Recorded Dec 31, 2009
From: AMERICA ONLINE, INC.
To: AOL LLC
Reel/Frame 023723/0585 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2009
From: AOL LLC
To: AOL INC.
Reel/Frame 023723/0645 →
SECURITY AGREEMENT Recorded Dec 14, 2009
From: AOL INC.; AOL ADVERTISING INC.; BEBO, INC.; ICQ LLC; GOING, INC.; LIGHTNINGCAST LLC; MAPQUEST, INC.; NETSCAPE COMMUNICATIONS CORPORATION; QUIGO TECHNOLOGIES LLC; SPHERE SOURCE, INC.; TACODA LLC; TRUVEO, INC.; YEDDA, INC.
To: BANK OF AMERICAN, N.A. AS COLLATERAL AGENT
Reel/Frame 023649/0061 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2004
From: ALSPECTOR, JOSHUA; KOLCZ, ALEKSANDER
To: AMERICA ONLINE, INC.
Reel/Frame 014616/0714 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2004
From: ALSPECTOR, JOSHUA; KOLCZ, ALEKSANDER
To: AMERICA ONLINE, INC.
Reel/Frame 014616/0476 →