IP Library Patent Application 18630777
Patent Application
App. No. 18/630,777

IDENTIFICATION OF SENSITIVE INFORMATION IN DATASETS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/630,777
Filed
Apr 9, 2024
Art Unit
2454
USPC
726/26
Abstract

A method, a system, and a computer program product for identifying sensitive data. A plurality of text portions associated with one or more data subjects is identified. A machine learning model is applied to the identified plurality of portions to extract one or more entities representative of one or more data subjects. The entities are grouped into one or more entity groups. Based on one or more entity groups, at least one data subject is identified for replacement or redaction in at least one text portion in the plurality of text portions.

Claims (32)

1 . A computer-implemented method, comprising:

identifying, using at least one processor, a plurality of text portions associated with one or more data subjects;

applying, using the at least one processor, a machine learning model to the identified plurality of text portions to extract one or more entities representative of the one or more data subjects;

grouping, using the at least one processor, the one or more entities into one or more entity groups; and

identifying, using the at least one processor, based on the one or more entity groups, at least one data subject for replacement or redaction in at least one text portion in the plurality of text portions.

2 . The method of claim 1 , wherein the grouping includes grouping the one or more entities using at least one of: a semantic similarity between entities, a relationship between entities, and any combination thereof.

3 . The method of claim 1 , further comprising assigning one or more weights to the one or more entities based on a representation of at least one data subject in the one or more data subjects by each entity in the one or more entities.

4 . The method of claim 3 , wherein the grouping includes grouping the one or more entities using the one or more weights.

5 . The method of claim 1 , wherein the plurality of text portions includes: at least one document, at least one portion of a document, and any combination thereof.

6 . The method of claim 1 , wherein the machine learning model is trained using a plurality of historical data subjects.

7 . The method of claim 1 , wherein the one or more data subjects include at least one of: a sensitive data or information, a commercially sensitive data or information, a trade secret data or information, a secret data or information, a non-public data or information, and any combination thereof.

8 . A system, comprising:

at least one processor; and

at least one non-transitory storage media storing instructions, that when executed by the at least one processor, cause the at least one processor to:

apply a machine learning model to a plurality of text portions to extract one or more entities representative of the one or more data subjects, wherein the plurality of text portions are associated with one or more data subjects;

group the one or more entities into one or more entity groups; and

identify, based on the one or more entity groups, at least one data subject for replacement or redaction in at least one text portion in the plurality of text portions.

9 . The system of claim 8 , wherein the grouping includes grouping the one or more entities using at least one of: a semantic similarity between entities, a relationship between entities, and any combination thereof.

10 . The system of claim 8 , wherein the at least one processor is configured to assign one or more weights to the one or more entities based on a representation of at least one data subject in the one or more data subjects by each entity in the one or more entities.

11 . The system of claim 10 , wherein grouping of the one or more entities includes grouping the one or more entities using the one or more weights.

12 . The system of claim 8 , wherein the plurality of text portions includes: at least one document, at least one portion of a document, and any combination thereof.

13 . The system of claim 8 , wherein the machine learning model is trained using a plurality of historical data subjects.

14 . The system of claim 8 , wherein the one or more data subjects include at least one of: a sensitive data or information, a commercially sensitive data or information, a trade secret data or information, a secret data or information, a non-public data or information, and any combination thereof.

15 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by at least one processor, cause the at least one processor to:

apply a machine learning model to a plurality of text portions to extract one or more entities representative of the one or more data subjects, wherein the plurality of text portions are associated with one or more data subjects;

group the one or more entities into one or more entity groups using at least one of: a semantic similarity between entities, a relationship between entities, and any combination thereof; and

identify, based on the one or more entity groups, at least one data subject for replacement or redaction in at least one text portion in the plurality of text portions.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the at least one processor is configured to assign one or more weights to the one or more entities based on a representation of at least one data subject in the one or more data subjects by each entity in the one or more entities.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein grouping of the one or more entities includes grouping the one or more entities using the one or more weights.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein the plurality of text portions includes: at least one document, at least one portion of a document, and any combination thereof.

19 . The non-transitory computer-readable storage medium of claim 15 , wherein the machine learning model is trained using a plurality of historical data subjects.

20 . The non-transitory computer-readable storage medium of claim 15 , wherein the one or more data subjects include at least one of: a sensitive data or information, a commercially sensitive data or information, a trade secret data or information, a secret data or information, a non-public data or information, and any combination thereof.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 23, 2025
From: DOCUSIGN, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 071337/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: HUANG, YANGCHENG; HASAN, SOULEIMAN; MAHON, SEAN; HE, YAN; MOU, CHENGHAO; PHAM, NGHIA; BELLINI, ALBERTO MARIO; UMER, SHAHEEN; BASI, MOE
To: DOCUSIGN, INC.
Reel/Frame 067189/0776 →