IP Library Granted Patent US 12,388,869
Granted Patent B2
US 12,388,869 · App. 18/732,870 · Granted Aug 12, 2025

Message management platform for performing impersonation analysis and detection

Inventor: Harold Nguyen (San Carlos, CA)
Assignee: Proofpoint, Inc.
H04L63/1483G06N5/04G06N20/00H04L51/212H04L51/224
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,388,869
App. No.
18/732,870
Granted
Aug 12, 2025
Kind
B2
Abstract

Aspects of the disclosure relate to detecting impersonation in email body content using machine learning. Based on email data received from user accounts, a computing platform may generate user identification models that are each specific to one of the user accounts. The computing platform may intercept a message from a first user account to a second user account and may apply a user identification model, specific to the first user account, to the message, so as to calculate feature vectors for the message. The computing platform then may apply impersonation algorithms to the feature vectors and may determine that the message is impersonated. Based on results of the impersonation algorithms, the computing platform may modify delivery of the message.

Claims (100)

1. A computing platform comprising:

at least one processor;

a communication interface communicatively coupled to the at least one processor; and

memory storing computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

generate, based on email data received from a plurality of user accounts, a plurality of user identification models;

intercept a first email message from a first user account of the plurality of user accounts to a second user account of the plurality of user accounts;

apply a first model of the plurality of user identification models to the first email message to calculate a first plurality of feature vectors for the first email message, wherein the first model of the plurality of user identification models is specific to the first user account of the plurality of user accounts;

apply one or more impersonation algorithms to the first plurality of feature vectors to determine results of the one or more impersonation algorithms, wherein applying the one or more impersonation algorithms to the first plurality of feature vectors indicates that the first email message is an impersonated message;

based on the results of the one or more impersonation algorithms, modify delivery of the first email message;

intercept a second email message from a third user account of the plurality of user accounts to the second user account of the plurality of user accounts;

apply a second model of the plurality of user identification models to the second email message to calculate a second plurality of feature vectors for the second email message, wherein the second model of the plurality of user identification models is specific to the third user account of the plurality of user accounts;

apply the one or more impersonation algorithms to the second plurality of feature vectors, wherein applying the one or more impersonation algorithms to the second plurality of feature vectors indicates that the second email message is a legitimate message; and

based on results of the one or more impersonation algorithms, permit delivery of the second email message.

2. The computing platform of claim 1 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

determine, based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determine a deviation value between the confidence score and a predetermined impersonation threshold;

determine that the deviation value does not exceed a first deviation threshold of a plurality of deviation thresholds;

determine, based on the determination that the deviation value does not exceed the first deviation threshold, alert information indicating that the first email message is an impersonated message,

wherein modifying delivery of the first email message comprises:

sending, to a user device associated with the first user account, the alert information, wherein sending the alert information causes the user device associated with the first user account to display an alert indicating that the first email message is an impersonated message; and

sending, to a user device associated with the second user account, the first email message.

3. The computing platform of claim 2 , wherein modifying delivery of the first email message comprises modifying a subject line of the first email message prior to sending the first email message to the user device associated with the second user account.

4. The computing platform of claim 2 , wherein the memory stores additional computer-readable instructions that, when executed by the least one processor, cause the computing platform to:

receive, from the user device associated with the first user account, an indication that the first email message was not impersonated; and

update, based on the indication that the first email message was not impersonated, one or more machine learning datasets to indicate that the first email message was legitimate.

5. The computing platform of claim 2 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

retune the plurality of deviation thresholds based on a target percentage of email messages to be flagged as impersonated, wherein the retuning is based on one or more machine learning datasets comprising indications of identified impersonated messages.

6. The computing platform of claim 1 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

determine, based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determine a deviation value between the confidence score and a predetermined impersonation threshold;

determine that the deviation value exceeds a first deviation threshold but does not exceed a second deviation threshold;

determine, based on the determination that the deviation value exceeds the first deviation threshold but does not exceed the second deviation threshold, that the first email message should be routed to an online mailbox configured to receive messages flagged as impersonated, wherein the online mailbox is accessible by a user device associated with the second user account; and

route, to the online mailbox, the first email message.

7. The computing platform of claim 1 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

determine, based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determine a deviation value between the confidence score and a predetermined impersonation threshold;

determine that the deviation value exceeds a first deviation threshold and a second deviation threshold greater than the first deviation threshold but does not exceed a third deviation threshold greater than the second deviation threshold;

determine, based on the determination that the deviation value exceeds the first deviation threshold and the second deviation threshold but does not exceed the third deviation threshold, that an administrator computing device should be notified that the first email message is an impersonated message; and

send, to the administrator computing device, impersonation alert information, wherein sending the impersonation alert information to the administrator computing device causes the administrator computing device to display an impersonation warning interface.

8. The computing platform of claim 7 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

receive, from the administrator computing device, one or more commands directing the computing platform to delete the first email message.

9. The computing platform of claim 1 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

determine, based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determine a deviation value between the confidence score and a predetermined impersonation threshold;

determine that the deviation value exceeds a deviation threshold of a plurality of deviation thresholds;

determine, based on the determination that the deviation value exceeds the deviation threshold of the plurality of deviation thresholds, that the first email message should be quarantined; and

prevent transmission of the first email message to a user device associated with the second user account.

10. The computing platform of claim 1 , wherein the email data comprises one or more of: a number of blank lines, a total number of lines, an average sentence length, an average word length, a vocabulary richness score, stop word frequency, a number of times one or more distinct words are used a single time, a total number of characters, a total number of alphabetic characters, a total number of upper-case characters, a total number of digits, a total number of white-space characters, a total number of tabs, a total number of punctuation marks, a word length frequency distribution, or a parts of speech frequency distribution.

11. The computing platform of claim 10 , wherein the memory stores additional computer-readable instructions that, when executed by the at least one processor, cause the computing platform to:

determine the vocabulary richness score by determining a number of distinct words and a number of total words; and

dividing the number of distinct words by the number of total words.

12. A method, comprising:

at a computing platform comprising at least one processor, a communication interface, and memory:

generating, by the at least one processor, based on email data received from a plurality of user accounts, a plurality of user identification models;

intercepting, by the at least one processor, a first email message from a first user account of the plurality of user accounts to a second user account of the plurality of user accounts;

applying, by the at least one processor, a first model of the plurality of user identification models to the first email message to calculate a first plurality of feature vectors for the first email message, wherein the first model of the plurality of user identification models is specific to the first user account of the plurality of user accounts;

applying, by the at least one processor, one or more impersonation algorithms to the first plurality of feature vectors to determine results of the one or more impersonation algorithms, wherein applying the one or more impersonation algorithms to the first plurality of feature vectors indicates that the first email message is an impersonated message; and

based on the results of the one or more impersonation algorithms, modifying, by the at least one processor, delivery of the first email message;

intercepting, by the at least one processor, a second email message from a third user account of the plurality of user accounts to the second user account of the plurality of user accounts;

applying, by the at least one processor, a second model of the plurality of user identification models to the second email message to calculate a second plurality of feature vectors for the second email message, wherein the second model of the plurality of user identification models is specific to the third user account of the plurality of user accounts;

applying, by the at least one processor, the one or more impersonation algorithms to the second plurality of feature vectors, wherein applying the one or more impersonation algorithms to the second plurality of feature vectors indicates that the second email message is a legitimate message; and

based on results of the one or more impersonation algorithms, permitting, by the at least one processor, delivery of the second email message.

13. The method of claim 12 , wherein the computing platform is further configured to:

determining, by the at least one processor and based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determining, by the at least one processor, a deviation value between the confidence score and a predetermined impersonation threshold;

determining, by the at least one processor, that the deviation value does not exceed a first deviation threshold of a plurality of deviation thresholds;

determining, by the at least one processor and based on the determination that the deviation value does not exceed the first deviation threshold, alert information indicating that the first email message is an impersonated message,

wherein modifying delivery of the first email message comprises:

sending, by the at least one processor and to a user device associated with the first user account, the alert information, wherein sending the alert information causes the user device associated with the first user account to display an alert indicating that the first email message is an impersonated message; and

sending, by the at least one processor and to a user device associated with the second user account, the first email message.

14. The method of claim 13 , wherein modifying delivery of the first email message comprises modifying a subject line of the first email message prior to sending the first email message to the user device associated with the second user account.

15. The method of claim 13 , wherein the memory stores additional computer-readable instructions that, when executed by the least one processor, cause the computing platform to:

receiving, by the at least one processor and from the user device associated with the first user account, an indication that the first email message was not impersonated; and

updating, by the at least one processor and based on the indication that the first email message was not impersonated, one or more machine learning datasets to indicate that the first email message was legitimate.

16. The method of claim 13 , further comprising:

determining, by the at least one processor and based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determining, by the at least one processor, a deviation value between the confidence score and a predetermined impersonation threshold;

determining, by the at least one processor, that the deviation value exceeds a first deviation threshold and a second deviation threshold greater than the first deviation threshold but does not exceed a third deviation threshold greater than the second deviation threshold;

determining, by the at least one processor and based on the determination that the deviation value exceeds the first deviation threshold and the second deviation threshold but does not exceed the third deviation threshold, that an administrator computing device should be notified that the first email message is an impersonated message; and

sending, by the at least one processor and to the administrator computing device, impersonation alert information, wherein sending the impersonation alert information to the administrator computing device causes the administrator computing device to display an impersonation warning interface.

17. One or more non-transitory computer-readable media storing instructions that, when executed by a computing platform comprising at least one processor, a communication interface, and memory, cause the computing platform to:

generate, based on email data received from a plurality of user accounts, a plurality of user identification models, wherein each of the plurality of user identification models is specific to one of the user accounts;

intercept a first email message from a first user account of the plurality of user accounts to a second user account of the plurality of user accounts;

apply a first model of the plurality of user identification models to the first email message to calculate a first plurality of feature vectors for the first email message, wherein the first model of the plurality of user identification models is specific to the first user account of the plurality of user accounts;

apply one or more impersonation algorithms to the first plurality of feature vectors to determine results of the one or more impersonation algorithms,

wherein applying the one or more impersonation algorithms to the first plurality of feature vectors indicates that the first email message is an impersonated message;

based on the results of the one or more impersonation algorithms, modify delivery of the first email message;

intercept a second email message from a third user account of the plurality of user accounts to the second user account of the plurality of user accounts;

apply a second model of the plurality of user identification models to the second email message to calculate a second plurality of feature vectors for the second email message, wherein the second model of the plurality of user identification models is specific to the third user account of the plurality of user accounts;

apply the one or more impersonation algorithms to the second plurality of feature vectors, wherein applying the one or more impersonation algorithms to the second plurality of feature vectors indicates that the second email message is a legitimate message; and

based on results of the one or more impersonation algorithms, permit delivery of the second email message.

18. The one or more non-transitory computer-readable media of claim 17 , further including instructions that cause the computing platform to:

determine, based on applying the one or more impersonation algorithms, a confidence score indicative of a likelihood that the first email message is an impersonated message;

determine a deviation value between the confidence score and a predetermined impersonation threshold;

determine that the deviation value does not exceed a first deviation threshold of a plurality of deviation thresholds;

determine, based on the determination that the deviation value does not exceed the first deviation threshold, alert information indicating that the first email message is an impersonated message,

wherein modifying delivery of the first email message comprises:

sending, to a user device associated with the first user account, the alert information, wherein sending the alert information causes the user device associated with the first user account to display an alert indicating that the first email message is an impersonated message; and

sending, to a user device associated with the second user account, the first email message.

19. The one or more non-transitory computer-readable media of claim 18 , wherein modifying delivery of the first email message comprises modifying a subject line of the first email message prior to sending the first email message to the user device associated with the second user account.

Assignments (3)
INTELLECTUAL PROPERTY AGREEMENT SUPPLEMENT Recorded Dec 9, 2025
From: PROOFPOINT, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 073910/0027 →
SECOND LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 8, 2025
From: PROOFPOINT, INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 073889/0677 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2025
From: NGUYEN, HAROLD
To: PROOFPOINT, INC.
Reel/Frame 070994/0275 →
Continuity (4)
Continuation 18077479 · Dec 8, 2022
Continuation 16697918 · Nov 27, 2019
Provisional Application 62815088 · Mar 7, 2019
Related Publication 20240422194A1 · Dec 19, 2024
References Cited (38)
US 7363495B2 · Felt et al. · 2008 [cited by applicant]
US 7970705B2 · Patterson · 2011 [cited by applicant]
US 8489635B1 · Phoha et al. · 2013 [cited by applicant]
US 9177293B1 · Gagnon et al. · 2015 [cited by applicant]
US 9268927B1 · Phoha et al. · 2016 [cited by applicant]
US 9418057B2 · de Zeeuw et al. · 2016 [cited by applicant]
US 9667639B1 · Pierson et al. · 2017 [cited by applicant]
US 10320811B1 · Faham et al. · 2019 [cited by applicant]
US 10805311B2 · Greevy · 2020 [cited by applicant]
US 10965696B1 · Amram · 2021 [cited by examiner]
US 11330003B1 · Howell et al. · 2022 [cited by applicant]
US 20020078365A1 · Burnett · 2002 [cited by examiner]
US 20020138583A1 · Takayama · 2002 [cited by examiner]
US 20040221062A1 · Starbuck et al. · 2004 [cited by applicant]
US 20040225647A1 · Connelly et al. · 2004 [cited by applicant]
US 20060168024A1 · Mehr et al. · 2006 [cited by applicant]
US 20070050376A1 · Maida-Smith et al. · 2007 [cited by applicant]
US 20080072294A1 · Chatterjee · 2008 [cited by examiner]
US 20130138427A1 · de Zeeuw et al. · 2013 [cited by applicant]
US 20160014151A1 · Prakash · 2016 [cited by applicant]
US 20170034183A1 · Enqvist et al. · 2017 [cited by applicant]
US 20170078884A1 · Tanabe et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa · 2017 [cited by examiner]
US 20170257395A1 · Goutal · 2017 [cited by examiner]
US 20180152471A1 · Jakobsson · 2018 [cited by applicant]
US 20180375877A1 · Jakobsson et al. · 2018 [cited by applicant]
US 20190132273A1 · Ryan et al. · 2019 [cited by applicant]
US 20190268373A1 · Celik · 2019 [cited by examiner]
US 20190306192A1 · Xie · 2019 [cited by examiner]
US 20200134009A1 · Zhao et al. · 2020 [cited by applicant]
US 20200204572A1 · Jeyakumar et al. · 2020 [cited by applicant]
US 20210058395A1 · Jakobsson · 2021 [cited by examiner]
US 20210126942A1 · Fantham · 2021 [cited by examiner]
US 20210176275A1 · Celik · 2021 [cited by examiner]
Apr. 20, 2020 (EP) Extended European Search Report—App. 19214226.3. [cited by applicant]
Aug. 2, 2012 (EP) First Examination Report—App. 19214226.3. [cited by applicant]
May 27, 2022—Non-Final Office Action—U.S. Appl. No. 16/697,918. [cited by applicant]
Sep. 14, 2022—(US) Notice of Allowance—U.S. Appl. No. 16/697,918. [cited by applicant]