Systems, methods, and apparatuses for detecting sensitive terms in data
Methods, systems, and apparatuses are provided for predicting if an item of text data includes one or more sensitive terms. An item of text data may be received by a computing device from a second computing device associated with a user. The computing device may determine one or more words within the item of text data. An arrangement of the one or more words may be determined. The arrangement of the one or more words may be based on an assignment bearing vectorization of the item of text data and/or a position bearing vectorization of the item of text data. Based on the arrangement of the one or more words within the item of text data, the computing device may determine a potential sensitive term within the item of text data. The computing device may cause an output of an indication of the potential sensitive term.
1 . A method comprising:
receiving, by a computing device, text data;
determining, based on the text data, one or more words in the text data;
determining, by a trained machine-learning model and based on an arrangement of a first word of the one or more words with reference to a second word of the one or more words, a potential sensitive term in the text data; and
causing output of an indication of the potential sensitive term.
2 . The method of claim 1 , further comprising:
determining, based on the text data, a position for the one or more words within the text data,
wherein determining the potential sensitive term is further based on the position for the one or more words within the text data.
3 . The method of claim 1 , wherein the text data comprises a line of computer code.
4 . The method of claim 1 , further comprising:
receiving an indication that the potential sensitive term is not a sensitive term; and
updating, based on the potential sensitive term not being a sensitive term, an error variable.
5 . The method of claim 4 , further comprising:
determining that the error variable satisfies an error threshold; and
retraining the trained machine-learning model for detecting potential sensitive terms.
6 . The method of claim 1 , further comprising:
receiving user data identifying a user associated with the text data,
wherein determining the potential sensitive term is further based on the user associated with the text data.
7 . The method of claim 1 , wherein the potential sensitive term is one of a secret term, a confidential term, or a password.
8 . The method of claim 1 , further comprising:
receiving modified text data, wherein the modified text data does not comprise the potential sensitive term; and
determining, based on the modified text data not comprising the potential sensitive term, that the potential sensitive term is a sensitive term.
9 . The method of claim 1 , further comprising:
determining, the potential sensitive term is a sensitive term;
generating a representation of the sensitive term; and
comparing the representation to additional terms to identify additional sensitive terms.
10 . A method comprising:
receiving, by a computing device, text data;
determining, based on the text data, one or more words in the text data;
determining, by a trained machine-learning model and based on a position of a first word of the one or more words with reference to a second word of the one or more words in the text data, a potential sensitive term in the text data; and
causing output of an indication of the potential sensitive term.
11 . The method of claim 10 , further comprising:
determining an arrangement of the one or more words in the text data,
wherein determining the potential sensitive term is further based on the arrangement of the one or more words in the text data.
12 . The method of claim 10 , wherein the text data comprises the one or more words and one or more non-word elements, wherein the method further comprises determining, based on the text data, a position for each of the one or more words and the one or more non-word elements in the text data.
13 . The method of claim 10 , further comprising:
receiving an indication that the potential sensitive term is not a sensitive term; and
updating, based on the potential sensitive term not being the sensitive term, an error variable.
14 . The method of claim 13 , further comprising:
determining that the error variable satisfies an error threshold; and
retraining the trained machine-learning model for detecting potential sensitive terms.
15 . The method of claim 10 , further comprising:
receiving user data identifying a user associated with the text data,
wherein determining the potential sensitive term is further based on the user associated with the text data.
16 . The method of claim 10 , further comprising:
receiving modified text data, wherein the modified text data does not comprise the potential sensitive term; and
determining, based on the modified text data not comprising the potential sensitive term, that the potential sensitive term is a sensitive term.
17 . The method of claim 10 , wherein the potential sensitive term is one of a secret term, a confidential term, or a password.
18 . The method of claim 10 , further comprising:
determining, the potential sensitive term is a sensitive term;
generating a representation of the sensitive term; and
comparing the representation to additional words in the text data to identify additional sensitive terms in the text data.
19 . A method comprising:
receiving, by a computing device, text data;
determining, based on the text data, one or more words in the text data;
determining, by a trained machine-learning model and based on a position of a word of the one or more words, a potential sensitive term in the text data; and
causing output of an indication of the potential sensitive term.
20 . The method of claim 19 , further comprising:
determining an arrangement of the one or more words in the text data,
wherein determining the potential sensitive term is further based on the arrangement of the one or more words in the text data.
21 . The method of claim 19 , further comprising:
determining, based on the text data, a position for the one or more words within the text data,
wherein determining the potential sensitive term is further based on the position for the one or more words within the text data.