IP Library › Granted Patent US 11,461,371
Granted Patent B2
US 11,461,371 · App. 16/731,356 · Granted Oct 4, 2022

Methods and text summarization systems for data loss prevention and autolabelling

Inventors: Christopher Muffat (Singapore, SG); Tetiana Kodliuk (Singapore, SG)
Assignee: DATHENA SCIENCE PTE LTD.
G06F16/285G06F16/2433G06F16/93G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,371
App. No.
16/731,356
Granted
Oct 4, 2022
Kind
B2
Abstract

Methods and systems for data loss prevention and autolabelling of business categories and confidentiality based on text summarization are provided. The method for data loss prevention includes entering a combination of keywords and/or keyphrases and offline unsupervised mapping of a path of transfer of specific groups of documents. The offline unsupervised mapping includes keyword/keyphrase extraction from the specific groups of documents and normalization of candidates. The method further includes vectorization of the extracted keywords/keyphrases from the specific groups of documents and quantitative performance measurement of the keyword/keyphrase extraction to derive keywords and/or keyphrases suitable for data loss prevention.

Claims (24)

1. A method for data loss prevention comprising:

entering a combination of keywords and/or keyphrases;

offline unsupervised mapping of a path of transfer of specific groups of documents, wherein the offline unsupervised mapping comprises:

keyword/keyphrase extraction from the specific groups of documents; and

normalization of candidates;

vectorization of the extracted keywords/keyphrases from the specific groups of documents; and

quantitative performance measurement of the keyword/keyphrase extraction to derive keywords and/or keyphrases suitable for data loss prevention, wherein the extracted keywords/keyphrases correspond to the specific groups of documents.

2. The method in accordance with claim 1 wherein the offline unsupervised mapping further comprises text summarization from the specific group of documents.

3. The method in accordance with claim 1 further comprising autolabelling categories and confidential statuses of documents of the specific groups of documents in response to the keyword/keyphrase extraction.

4. The method in accordance with claim 1 wherein the combination of keywords and/or key phrases comprises both positive and negative combinations.

5. The method in accordance with claim 1 wherein the keyword/keyphrase extraction comprises keyword extraction by TDF-IDF.

6. The method in accordance with claim 1 wherein the keyword/keyphrase extraction comprises keyphrase extraction by DRAKE.

7. The method in accordance with claim 1 wherein the keyword/keyphrase extraction comprises keyphrase extraction by EmbedDocRank.

8. The method in accordance with claim 1 wherein the vectorization of the extracted keywords/keyphrases from the specific groups of documents comprises valuation based on information gain and cross-validation.

9. A system for autolabelling of documents comprising:

a model comprising a combination of keywords and/or keyphrases;

a feature extraction module for keyword/keyphrase extraction from the documents; and

an autolabelling engine for autolabelling categories and confidential statuses of the documents in response to the keyword/keyphrase extraction, wherein the extracted keywords/keyphrases correspond to the documents.

10. The system in accordance with claim 9 further comprising a data loss prevention module for performing a quantitative performance measurement of the keyword/keyphrase extraction to derive keywords and/or keyphrases suitable for data loss prevention.

11. The system in accordance with claim 9 further comprising a text summarization module for text summarization of the keywords/keyphrases extracted from the documents.

12. The system in accordance with claim 9 wherein the combination of keywords and/or key phrases comprises both positive and negative combinations.

13. The system in accordance with claim 9 wherein the feature extraction module performs keyword extraction by TDF-IDF.

14. The system in accordance with claim 9 wherein the feature extraction module performs keyphrase extraction by DRAKE.

15. The system in accordance with claim 9 wherein the feature extraction module performs keyphrase extraction by EmbedDocRank.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2020
From: MUFFAT, CHRISTOPHER; KODLIUK, TETIANA
To: DATHENA SCIENCE PTE LTD.
Reel/Frame 052355/0558 →
Priority Claims (1)
SG 10201811838T · Dec 31, 2018 · national
Continuity (1)
Related Publication 20200226154A1 · Jul 16, 2020