IP Library Patent Application 17163243
Patent Application
App. No. 17/163,243

SELECTING CONDITIONALLY INDEPENDENT INPUT SIGNALS FOR UNSUPERVISED CLASSIFIER TRAINING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/163,243
Abstract

Methods, systems, and computer program products for content management systems. An unlabeled dataset comprising documents that at least potentially comprise personally identifiable information (PII) is used when training a PII content classifier. Such a classifier is trained by (1) determining, based on applying a PII rule to a first portion of a document selected from the unlabeled dataset, a confidence value that the first portion of the document does contain personally identifiable information, (2) selecting a second portion of the document selected from the unlabeled dataset such that the second portion does not include the first portion; and (3) assigning, based on the confidence value, a likelihood value that corresponds to whether characteristics of the second portion are indicative that the document does contain personally identifiable information. Such a PII content classifier is used over selected portions of subject content objects to determine whether the selected portions contain PII.

Claims (48)

1 . A method comprising:

accessing an unlabeled dataset comprising documents that at least potentially comprise personally identifiable information (PII); and

training a content classifier by:

determining, based on applying a PII rule to a first portion of a document selected from the unlabeled dataset, a confidence value that the first portion of the document does contain personally identifiable information;

selecting a second portion of the document selected from the unlabeled dataset, wherein the second portion does not include the first portion; and

associating with the second portion, based on the confidence value, a likelihood value that corresponds to whether characteristics of the second portion are indicative that the document does contain personally identifiable information.

2 . The method of claim 1 , further comprising:

identifying a selected portion of a subject content object and applying the selected portion to the content classifier to determine whether characteristics of the selected portion are indicative that the document does contain PII.

3 . The method of claim 2 , further comprising:

communicating a message to a user device, wherein the message comprises at least a portion of one or more governance restrictions pertaining to communication of personally identifiable information.

4 . The method of claim 1 , wherein application of the PII rule to the first portion of the document is used to identify at least one of, one or more infotype designations, one or more infotype locations, or one or more infotype hotwords.

5 . The method of claim 4 , wherein the second portion of the document selected from the unlabeled dataset does not contain any occurrence of the one or more infotype hotwords.

6 . The method of claim 1 , further comprising:

adjusting a weight of either the likelihood value or the confidence value based on a gradient descent algorithm.

7 . The method of claim 1 , further comprising:

adjusting a weight of either the likelihood value or the confidence value based on an error calculation that compares a vector processor value to a rule processor value.

8 . A non-transitory computer readable medium having stored thereon a sequence of instructions which, when stored in memory and executed by one or more processors causes the one or more processors to perform a set of acts, the set of acts comprising:

accessing an unlabeled dataset comprising documents that at least potentially comprise personally identifiable information (PII); and

training a content classifier by:

determining, based on applying a PII rule to a first portion of a document selected from the unlabeled dataset, a confidence value that the first portion of the document does contain personally identifiable information;

selecting a second portion of the document selected from the unlabeled dataset, wherein the second portion does not include the first portion; and

associating with the second portion, based on the confidence value, a likelihood value that corresponds to whether characteristics of the second portion are indicative that the document does contain personally identifiable information.

9 . The non-transitory computer readable medium of claim 8 , further comprising instructions which, when stored in memory and executed by the one or more processors causes the one or more processors to perform acts of:

identifying a selected portion of a subject content object and applying the selected portion to the content classifier to determine whether characteristics of the selected portion are indicative that the document does contain PII.

10 . The non-transitory computer readable medium of claim 9 , further comprising instructions which, when stored in memory and executed by the one or more processors causes the one or more processors to perform acts of:

communicating a message to a user device, wherein the message comprises at least a portion of one or more governance restrictions pertaining to communication of personally identifiable information.

11 . The non-transitory computer readable medium of claim 8 , wherein application of the PII rule to the first portion of the document is used to identify at least one of, one or more infotype designations, one or more infotype locations, or one or more infotype hotwords.

12 . The non-transitory computer readable medium of claim 11 , wherein the second portion of the document selected from the unlabeled dataset does not contain any occurrence of the one or more infotype hotwords.

13 . The non-transitory computer readable medium of claim 8 , further comprising instructions which, when stored in memory and executed by the one or more processors causes the one or more processors to perform acts of:

adjusting a weight of either the likelihood value or the confidence value based on a gradient descent algorithm.

14 . The non-transitory computer readable medium of claim 8 , further comprising instructions which, when stored in memory and executed by the one or more processors causes the one or more processors to perform acts of:

adjusting a weight of either the likelihood value or the confidence value based on an error calculation that compares a vector processor value to a rule processor value.

15 . A system comprising:

a storage medium having stored thereon a sequence of instructions; and

one or more processors that execute the sequence of instructions to cause the one or more processors to perform a set of acts, the set of acts comprising,

accessing an unlabeled dataset comprising documents that at least potentially comprise personally identifiable information (PII); and

training a content classifier by:

determining, based on applying a PII rule to a first portion of a document selected from the unlabeled dataset, a confidence value that the first portion of the document does contain personally identifiable information;

selecting a second portion of the document selected from the unlabeled dataset, wherein the second portion does not include the first portion; and

associating with the second portion, based on the confidence value, a likelihood value that corresponds to whether characteristics of the second portion are indicative that the document does contain personally identifiable information.

16 . The system of claim 15 , further comprising:

identifying a selected portion of a subject content object and applying the selected portion to the content classifier to determine whether characteristics of the selected portion are indicative that the document does contain PII.

17 . The system of claim 16 , further comprising:

communicating a message to a user device, wherein the message comprises at least a portion of one or more governance restrictions pertaining to communication of personally identifiable information.

18 . The system of claim 15 , wherein application of the PII rule to the first portion of the document is used to identify at least one of, one or more infotype designations, one or more infotype locations, or one or more infotype hotwords.

19 . The system of claim 18 , wherein the second portion of the document selected from the unlabeled dataset does not contain any occurrence of the one or more infotype hotwords.

20 . The system of claim 15 , further comprising:

adjusting a weight of either the likelihood value or the confidence value based on a gradient descent algorithm.

Assignments (2)
SECURITY INTEREST Recorded Jul 26, 2023
From: BOX, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 064389/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: ESHGHI, KAVE; VIKRAMARATNE, VICTOR DE VANSA
To: BOX, INC.
Reel/Frame 055086/0416 →