IP Library Granted Patent US 11,425,077
Granted Patent B2
US 11,425,077 · App. 17/477,482 · Granted Aug 23, 2022

Method and system for determining a spam prediction error parameter

Inventor: Dmitry Sergeevich Korotkikh (d Avvakumovo, RU)
Assignee: YANDEX EUROPE AG
H04L51/212G06K9/6276G06N5/022H04L51/224H04L51/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,425,077
App. No.
17/477,482
Granted
Aug 23, 2022
Kind
B2
Abstract

Method and server for determining a spam prediction error parameter for a spam prediction parameter are disclosed. The method includes: receiving a plurality of emails destined to a plurality of users where a given email has a spam prediction parameter and a user-interaction parameter indicative of whether an associated recipient of the plurality of users agrees with the spam prediction parameter, and clustering the plurality of emails into at least two clusters having respective subsets of emails. For a given cluster the method includes determining a ground truth parameter by analyzing its subset of emails and the associated user-interaction parameters, and assigning the ground truth parameter to the given cluster. For the given email, the method includes generating the spam prediction error parameter based on a difference between the spam prediction parameter and the ground truth parameter, and storing the spam prediction error parameter in association with the given email.

Claims (48)

1. A method of determining a spam prediction error parameter for a spam prediction parameter generated by a spam detection algorithm executed by a server, the server associated with an email application and executing the spam detection algorithm, the method executed by the server, the method comprising:

receiving, by the server, an indication of a plurality of emails destined to a plurality of users of the email application, a given one of the plurality of emails having:

a respective spam prediction parameter indicative of the spam detection algorithm determining that the given one of the plurality of emails is one of a spam email and a non-spam email;

a user-interaction parameter indicative of whether an associated recipient of the plurality of users agrees with the respective spam prediction parameter;

clustering, by the server, the plurality of emails into at least two clusters, each one of the at least two clusters having a respective subset of emails;

for a given cluster from the at least two clusters:

determining, by the server, a respective ground truth parameter for the given cluster by analyzing the respective subset of emails and the associated user-interaction parameters,

the respective ground truth parameter being one of the spam email and the non-spam email;

assigning the respective ground truth parameter to the given cluster and each of the respective subset of emails contained therein;

for a given email from the given cluster:

generating, by the server, the spam prediction error parameter based on a difference between the spam prediction parameter and the respective ground truth parameter;

storing, by the server, the spam prediction error parameter in association with the given email from the given cluster.

2. The method of claim 1 , wherein the method further comprises:

determining, by the server, the user-interaction parameter based on at least one user interaction between the associated recipient and a respective email from the plurality of emails,

the at least one user interaction having been collected from an email interface displayed to the associated recipient.

3. The method of claim 2 , wherein the user interaction is at least one of (i) moving the respective email into a folder of the email interface, and (ii) clicking a pre-determined button of the email interface.

4. The method of claim 1 , wherein the clustering the plurality of emails is executed based on email features similarity.

5. The method of claim 4 , wherein the clustering is executed using a K-Nearest Neighbor (KNN) algorithm.

6. The method of claim 1 , wherein the server further executes the email application.

7. The method of claim 1 , wherein the server is configured to connect to a mail server executing the email application.

8. The method of claim 1 , wherein the indication of the plurality of emails comprises the plurality of emails.

9. The method of claim 1 , wherein the indication of the plurality of emails comprises an embedding of each of the plurality of emails, the embedding indicative of a content of the plurality of emails and devoid of any identifiers of associated recipients.

10. The method of claim 1 , wherein the method further comprises:

analyzing, by the server, a total number of emails in a given subset of emails of an other given cluster from the at least two clusters; and

in response to the number being below a pre-determined threshold, excluding, by the server, the other given cluster from further analysis.

11. The method of claim 1 , wherein the method further comprises:

retraining, by the server, the spam detection algorithm by using the spam prediction error parameter.

12. The method of claim 1 , wherein a given one of the at least two clusters comprises at least two sub-clusters.

13. The method of claim 12 , wherein the given one plurality of emails is clustered in both the given one of the at least two clusters and one of the at least two sub-clusters.

14. The method of claim 13 , wherein in response to the given one of the plurality of emails being associated with the ground truth parameter indicative of a wrong categorization in one of the given one of the at least two clusters and one of the at least two sub-clusters, a same value is used for the ground truth parameter for the given one plurality of emails.

15. The method of claim 13 , wherein the ground truth parameter is independently assigned to the given one of the plurality of emails in one of the given one of the at least two clusters and one of the at least two sub-clusters.

16. A server for determining a spam prediction error parameter for a spam prediction parameter generated by a spam detection algorithm executed by the server, the server associated with an email application and executing the spam detection algorithm, the server comprising a processor and being configured to:

receive, by the processor, an indication of a plurality of emails destined to a plurality of users of the email application, a given one of the plurality of emails having:

a respective spam prediction parameter indicative of the spam detection algorithm determining that the given one of the plurality of emails is one of a spam email and a non-spam email;

a user-interaction parameter indicative of an associated recipient of the plurality of users agreeing with the respective spam prediction parameter;

cluster, by the processor, the plurality of emails into at least two clusters, each one of the at least two clusters having a respective subset of emails;

for a given cluster from the at least two clusters:

determine, by the processor, a respective ground truth parameter for the given cluster by analyzing the respective subset of emails and the associated user-interaction parameters, the respective ground truth parameter being one of the spam email and the non-spam email;

assign, by the processor, the respective ground truth parameter to the given cluster and each of the respective subset of emails contained therein;

for a given email from the given cluster:

generate, by the processor, the spam prediction error parameter based on a difference between the spam prediction parameter and the respective ground truth parameter;

store, by the processor, the spam prediction error parameter in association with the given email from the given cluster.

17. The server of claim 16 , wherein the server is further configured to:

determine the user-interaction parameter based on at least one user interaction between the associated recipient and a respective email from the plurality of emails,

the at least one user interaction having been collected from an email interface displayed to the associated recipient.

18. The server of claim 17 , wherein the user interaction is at least one of (i) moving the respective email into a folder of the email interface, and (ii) clicking a pre-determined button of the email interface.

19. The server of claim 16 , wherein the clustering the plurality of emails is executed by the server based on email features similarity.

20. The server of claim 19 , wherein the clustering is executed by the server using a K-Nearest Neighbor (KNN) algorithm.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0687 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: KOROTKIKH, DMITRY SERGEEVICH
To: YANDEX.TECHNOLOGIES LLC
Reel/Frame 057508/0865 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: YANDEX.TECHNOLOGIES LLC
To: YANDEX LLC
Reel/Frame 057508/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 057508/0900 →
Cited By (1)
US 12,566,728