IP Library › Granted Patent US 12,367,229
Granted Patent B2
US 12,367,229 · App. 17/804,055 · Granted Jul 22, 2025

System and method for integrating machine learning in data leakage detection solution through keyword policy prediction

Inventors: Ahmad F. Sirhani (Dammam, SA); Abdullah K. Madani (Dhahran, SA); Abdulrahman M. Alomar (Al Hasa, SA)
Assignee: SAUDI ARABIAN OIL COMPANY
G06F16/35G06F16/93G06F21/554
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,229
App. No.
17/804,055
Granted
Jul 22, 2025
Kind
B2
Abstract

A method which includes receiving a corpus of labelled documents according to a plurality of filters and parsing, by a computer processor, the corpus. The method further includes vectorizing, by the computer processor, the parsed corpus to obtain vectorized documents; and training, by the computer processor, a machine-learned model using, at least a portion, of the vectorized documents. The method further includes extracting word importances from the trained machine-learned model and retaining the words with associated importances that satisfy a criterion, wherein the retained words are suggested keywords. The method further includes incorporating the suggested keywords in a policy of a data leakage prevention system.

Claims (81)

1. A method, comprising:

receiving, by a data leakage prevention (DLP) system comprising a computer processor, a machine learning (ML) system, a memory, and a data fetcher, a corpus of labelled documents using the data fetcher in the DLP system from a SQL-based repository and a plurality of filters comprising an organization filter,

wherein the corpus of labelled documents comprise at least one email, at least one spreadsheet, and at least one binary file,

wherein the data fetcher comprises an interface connected to the SQL-based repository that obtains the corpus of labelled documents based on one or more SQL commands comprising one or more queries based on the plurality of filters, and

wherein the ML system comprises a parser, a vectorizer, and a machine-learned model;

identifying, automatically by the DLP system, a plurality of parsed words in the corpus of labelled documents using the parser in the ML system;

vectorizing, by the DLP system, the corpus of labelled documents using the vectorizer in the ML system, comprising:

generating a matrix, wherein each of the rows of the matrix correspond to a labelled document and each of the columns of the matrix correspond to a parsed word from the plurality of parsed words,

determining, for each labelled document and using a Natural Language Processing (NLP) technique, a numerical value for each parsed word in the plurality of parsed words forming a vectorized document for that labelled document, and

populating the matrix with the with the numerical value of each parsed word in the plurality of parsed words for each labeled document;

training, by the ML system in the DLP system, the machine-learned model comprising:

predicting a document class for at least one vectorized document of the matrix,

comparing the predicted document class to a corresponding label of the at least one vectorized document, and

updating or determining one or more parameters of the machine-learned model based on the comparison,

wherein the trained machine-learned model accepts, as input, a vectorized document and outputs at least one predicted document class from a plurality of document classes comprising a sensitive document class and a non-sensitive document class;

extracting a plurality of word importances from the trained machine-learned model based on the one or more parameters, wherein the plurality of word importances comprises a word importance for each parsed word of the plurality of parsed words;

determining, automatically by the DLP system, a keyword-based policy based on the plurality of word importances and a portion of the plurality of parsed words that satisfy a criterion;

obtaining, by the DLP system, a new document; and

automatically classifying, by the DLP system, the new document as a sensitive document based on the keyword-based policy.

2. The method of claim 1 , wherein the plurality of filters further comprises a date range filter.

3. The method of claim 1 , further comprising evaluating the keyword-based policy by a subject matter expert, wherein the subject matter expert determines which parsed words are incorporated into the keyword-based policy of the data leakage prevention system.

4. The method of claim 1 , further comprising pre-processing the vectorized documents before use with the machine-learned model.

5. The method of claim 1 , further comprising:

selecting a machine-learned model type and an architecture; and

altering the machine-learned model type and/or architecture, or re-training the machine-learned model based, at least in part, on an evaluation of the keyword-based policy by a subject matter expert.

6. The method of claim 1 , wherein the machine-learned model is a logistic regression model.

7. A non-transitory computer readable medium storing instructions executable by a computer processor, the instructions comprising functionality for:

receiving a corpus of labelled documents using a data fetcher from a SQL-based repository and a plurality of filters comprising an organization filter,

wherein the corpus of labelled documents comprise at least one email, at least one spreadsheet, and at least one binary file, and

wherein the data fetcher comprises an interface connected to the SQL-based repository that obtains the corpus of labelled documents based on one or more SQL commands comprising one or more queries based on the plurality of filters;

identifying, automatically, a plurality of parsed words in the corpus of labelled documents using a parser comprised by a machine learning (ML) system,

wherein the ML system further comprises a vectorizer, and a machine-learned model;

vectorizing the corpus of labelled documents using the vectorizer in the ML system, comprising:

generating a matrix, wherein each of the rows of the matrix correspond to a labelled document and each of the columns of the matrix correspond to a parsed word from the plurality of parsed words,

determining, for each labelled document and using a Natural Language Processing (NLP) technique, a numerical value for each parsed word in the plurality of parsed words forming a vectorized document for that labelled document, and

populating the matrix with the with the numerical value of each parsed word in the plurality of parsed words for each labeled document;

training, using the ML system, the machine-learned model, comprising:

predicting a document class for at least one vectorized document of the matrix,

comparing the predicted document class to a corresponding label of the at least one vectorized document, and

updating or determining one or more parameters of the machine-learned model based on the comparison,

wherein the trained machine-learned model accepts, as input, a vectorized document and outputs at least one predicted document class from a plurality of document classes comprising a sensitive document class and a non-sensitive document class;

extracting a plurality of word importances from the trained machine-learned model based on the one or more parameters wherein the plurality of word importances comprises a word importance for each parsed word of the plurality of parsed words;

determining, automatically, a keyword-based policy for a data leakage prevention (DLP) system based on the plurality of word importances and a portion of the plurality of parsed words that satisfy a criterion;

obtaining a new document; and

automatically, using the DLP system, classifying the new document as a sensitive document based on the keyword-based policy.

8. The non-transitory computer readable medium of claim 7 , wherein the plurality of filters further comprises a date range filter.

9. The non-transitory computer readable medium of claim 7 , wherein the keyword-based policy is evaluated by a subject matter expert, wherein the subject matter expert determines which parsed words are incorporated into the keyword-based policy of the data leakage prevention system.

10. The non-transitory computer readable medium of claim 7 , the instructions further comprising functionality for:

pre-processing the vectorized documents before use with the machine-learned model.

11. The non-transitory computer readable medium of claim 7 , the instructions further comprising functionality for:

selecting a machine-learned model type and an architecture; and

altering the machine-learned model type and/or architecture, or re-training the machine-learned model based, at least in part, on an evaluation of the keyword-based policy by a subject matter expert.

12. The non-transitory computer readable medium of claim 7 , wherein the machine-learned model is a logistic regression model.

13. A data leak prevention (DLP) system, comprising:

a computer processor;

a memory;

a data fetcher comprising an interface connected to a SQL-based repository; and

a machine learning (ML) system comprising a parser, a vectorizer, and a machine-learned model,

wherein the DLP system is configured to:

receive, using the data fetcher, a corpus of labelled documents comprising at least one email, at least one spreadsheet, and at least one binary file, wherein the data fetcher obtains the corpus from the SQL-based repository based on one or more SQL commands comprising one or more queries based on a plurality of filters;

identify, automatically, a plurality of parsed words in the corpus of labelled documents using the parser in the ML system;

vectorize the corpus of labelled documents using the vectorizer in the ML system, comprising:

generating a matrix, wherein each of the rows of the matrix correspond to a labelled document and each of the columns of the matrix correspond to a parsed word from the plurality of parsed words,

determining, for each labelled document and using a Natural Language Processing (NLP) technique, a numerical value for each parsed word in the plurality of parsed words forming a vectorized document for that labelled document, and

populating the matrix with the with the numerical value of each parsed word in the plurality of parsed words for each labeled document;

train, by the ML system, the machine-learned model, comprising:

predicting a document class for at least one vectorized document of the matrix,

comparing the predicted document class to a corresponding label of the at least one vectorized document, and

updating or determining one or more parameters of the machine-learned model based on the comparison,

wherein the trained machine-learned model accepts, as input, a vectorized document and outputs at least one predicted document class from a plurality of document classes comprising a sensitive document class and a non-sensitive document class;

extract a plurality of word importances from the trained machine-learned model based on the one or more parameters, wherein the plurality of word importances comprises a word importance for each parsed word of the plurality of parsed words;

determine, automatically, a keyword-based policy based on the plurality of word importances and a portion of the plurality of parsed words that satisfy a criterion;

obtain a new document; and

automatically classify the new document as a sensitive document based on the keyword-based policy.

14. The system of claim 13 , wherein the plurality of filters further comprises a date range filter.

15. The system of claim 13 , wherein the keyword-based policy is evaluated by a subject matter expert, wherein the subject matter expert determines which parsed words are incorporated into the keyword-based policy of the data leakage prevention system.

16. The system of claim 13 , wherein the DLP system is further configured to:

pre-process the vectorized documents before use with the machine-learned model.

17. The system of claim 13 , wherein the DLP system is further configured to:

select a machine-learned model type and an architecture; and

alter the machine-learned model type and/or architecture, or re-train the machine-learned model based, at least in part, on an evaluation of the keyword-based policy by a subject matter expert.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: SIRHANI, AHMAD F.; MADANI, ABDULLAH K.; ALOMAR, ABDULRAHMAN M.
To: SAUDI ARABIAN OIL COMPANY
Reel/Frame 062423/0831 →
Continuity (1)
Related Publication 20230385407A1 · Nov 30, 2023
References Cited (14)
US 8423483B2 · Sadeh-Koniecpol et al. · 2013 [cited by applicant]
US 8671080B1 · Upadhyay et al. · 2014 [cited by applicant]
US 8862522B1 · Jaiswal et al. · 2014 [cited by applicant]
US 9230096B2 · Sarin et al. · 2016 [cited by applicant]
US 9342697B1 · Ren · 2016 [cited by examiner]
US 10521719B1 · Walters · 2019 [cited by examiner]
US 10979458B2 · Narayanaswamy et al. · 2021 [cited by applicant]
US 20090144255A1 · Chow · 2009 [cited by examiner]
US 20140304197A1 · Jaiswal · 2014 [cited by examiner]
US 20160203321A1 · Ayres et al. · 2016 [cited by applicant]
US 20180336278A1 · Agarwal · 2018 [cited by examiner]
US 20200226154A1 · Muffat · 2020 [cited by examiner]
US 20220043850A1 · Mishra · 2022 [cited by examiner]
US 20220337628A1 · Mudgil · 2022 [cited by examiner]