IP Library › Granted Patent US 11,847,244
Granted Patent B1
US 11,847,244 · App. 16/545,952 · Granted Dec 19, 2023

Private information detector for data loss prevention

Inventors: Isaac Abhay Madan (San Francisco, CA); Rohan Shrikant Sathe (San Francisco, CA); Trung Hoai Nguyen (Alameda, CA); Yiang Zheng (San Francisco, CA)
Assignee: Shoreline Labs, Inc.
G06F21/6245G06F9/54G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,847,244
App. No.
16/545,952
Granted
Dec 19, 2023
Kind
B1
Abstract

A private information detector for data loss prevention is described. In one embodiment, a method includes training a first, machine learning model on a set of known application programming interface keys to detect application programming interface keys, training a second, machine learning model on code including known application programming interface keys to detect adjacent characters to application programming interface keys, scanning a repository with the first, machine learning model to select a proper subset of the repository that includes possible application programming interface keys and adjacent characters, scanning the proper subset of the repository with the second, machine learning model to detect and remove potential false positives of the possible application programming interface keys based on adjacent characters of the possible application programming interface keys to generate a list of probable application programming interface keys, and sending an indication for the list of probable application programming interface keys.

Claims (50)

1. A method comprising:

training a first machine learning model on a first set of known application programming interface keys to detect application programming interface keys;

training a second machine learning model on code including a second set of known application programming interface keys to detect adjacent characters to application programming interface keys;

scanning a repository with the first machine learning model to select a subset of the repository that includes possible application programming interface keys and adjacent characters;

scanning the subset of the repository with the second machine learning model to detect and remove potential false positives of the possible application programming interface keys, based on generating an inference by the second machine learning model for adjacent characters of the possible application programming interface keys, to generate a list of probable application programming interface keys, wherein the generating the inference by the second machine learning model comprises:

converting a set of the adjacent characters of the subset of the repository according to a dictionary into a vector of numbers,

embedding the vector of numbers into an embedded matrix comprising a row for each character of the set of the adjacent characters and a plurality of weights in each column for a particular row,

generating a state matrix from the embedded matrix,

generating a context vector from the state matrix, and

generating an output from the context vector, and

sending an indication for the list of probable application programming interface keys.

2. The method of claim 1 , wherein the scanning with the first machine learning model is in response to receipt of a uniform resource locator value and a corresponding key for the repository.

3. The method of claim 1 , wherein the possible application programming interface keys includes at least one application programming interface key that is not one of the first set of known application programming interface keys or the second set of known application programming interface keys.

4. The method of claim 1 , wherein the indication comprises a link to each of the probable application programming interface keys within the repository.

5. The method of claim 1 , wherein the indication does not include a probable application programming interface key itself of the probable application programming interface keys.

6. The method of claim 1 , wherein the indication includes a subset of characters of each of the probable application programming interface keys.

7. The method of claim 1 , wherein the scanning with the first machine learning model and the scanning with the second machine learning model does not retain a copy of the repository.

8. A method comprising:

training a first machine learning model on a first set of known private information to detect private information;

training a second machine learning model on code including a second set of known private information to detect adjacent characters to known private information;

scanning a repository with the first machine learning model to select a subset of the repository that includes possible private information and adjacent characters;

scanning the subset of the repository with the second machine learning model to detect and remove potential false positives of the possible private information, based on generating an inference by the second machine learning model for adjacent characters of the possible private information, to generate a list of probable private information, wherein the generating the inference by the second machine learning model comprises:

converting a set of the adjacent characters of the subset of the repository according to a dictionary into a vector of numbers,

embedding the vector of numbers into an embedded matrix comprising a row for each character of the set of the adjacent characters and a plurality of weights in each column for a particular row,

generating a state matrix from the embedded matrix,

generating a context vector from the state matrix, and

generating an output from the context vector, and

sending an indication for the list of probable private information.

9. The method of claim 8 , wherein the scanning with the first machine learning model is in response to receipt of a uniform resource locator value and a corresponding key for the repository.

10. The method of claim 8 , wherein the possible private information is not one of the first set of known private information or the second set of known private information.

11. The method of claim 8 , wherein the indication comprises a link to the probable private information within the repository.

12. The method of claim 8 , wherein the indication does not include the probable private information itself.

13. The method of claim 8 , wherein the indication only includes a subset of characters of the probable private information.

14. The method of claim 8 , wherein the scanning with the first machine learning model and the scanning with the second machine learning model does not retain a copy of the repository.

15. A non-transitory machine readable medium that stores program code that when executed by a machine causes the machine to perform a method comprising:

training a first machine learning model on a first set of known private information to detect private information;

training a second machine learning model on code including a second set of known private information to detect adjacent characters to known private information;

scanning a repository with the first machine learning model to select a subset of the repository that includes possible private information and adjacent characters;

scanning the subset of the repository with the second machine learning model to detect and remove potential false positives of the possible private information, based on generating an inference by the second machine learning model for adjacent characters of the possible private information, to generate a list of probable private information, wherein the generating the inference by the second machine learning model comprises:

converting a set of the adjacent characters of the subset of the repository according to a dictionary into a vector of numbers,

embedding the vector of numbers into an embedded matrix comprising a row for each character of the set of the adjacent characters and a plurality of weights in each column for a particular row,

generating a state matrix from the embedded matrix,

generating a context vector from the state matrix, and

generating an output from the context vector, and

sending an indication for the list of probable private information.

16. The non-transitory machine readable medium of claim 15 , wherein the scanning with the first machine learning model is in response to receipt of a uniform resource locator value and a corresponding key for the repository.

17. The non-transitory machine readable medium of claim 15 , wherein the possible private information is not one of the first set of known private information or the second set of known private information.

18. The non-transitory machine readable medium of claim 15 , wherein the indication comprises a link to the probable private information within the repository.

19. The non-transitory machine readable medium of claim 15 , wherein the indication does not include the probable private information itself.

20. The non-transitory machine readable medium of claim 15 , wherein the indication only includes a subset of characters of the probable private information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2019
From: MADAN, ISAAC ABHAY; SATHE, ROHAN SHRIKANT; NGUYEN, TRUNG HOAI; ZHENG, YIANG
To: SHORELINE LABS, INC.
Reel/Frame 050107/0864 →
Cited By (5)
US 12,301,551 US 12,432,063 US 12,445,427 US 12,530,491 US 12,632,594