Personalized natural language processing system
A personalized natural language processing system tokenizes a plurality of sets of raw text data to generate a plurality of sets of tokenized text data for the plurality of users, respectively. The tokenized text data includes a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text. The system appends predetermined user-specific tokens to the sets of tokenized text data from the users, respectively. Each predetermined user-specific token corresponds to one of the users. The system processes the sets of tokenized text data using the NLP model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for the sets of tokenized text data from each of the users, and outputs the personalized classifications of the tokenized text data for each of the users.
1. A personalized natural language processing system comprising:
at least one processor, communicatively coupled to non-volatile memory storing a natural language processing (NLP) model personalized for use by multiple users and instructions that, when executed by the processor, cause the processor to:
receive or retrieve a plurality of sets of raw text data from a plurality of users, respectively;
tokenize the plurality of sets of raw text data to generate a plurality of sets of tokenized text data for the plurality of users, respectively, each set of tokenized text data including a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text data;
append predetermined user-specific tokens to the plurality of sets of tokenized text data from the plurality of users, respectively, to generate a plurality of user-specific token sets, each predetermined user-specific token corresponding to one user of the plurality of users;
process the plurality of user-specific token sets using the NLP model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for each of the plurality of user-specific token sets; and
output the personalized classifications of the plurality of user-specific token sets, wherein the NLP model is trained using a training data set comprising multiple tuples of training data, each tuple including user identifier data that identifies a particular user, raw text data from the particular user, and ground truth classification data, the trained NLP model being configured to perform personalized text classification and/or text prediction tasks for each user of the multiple users, and during processing of the plurality of user-specific token sets, embeddings are produced for each token in each set of the plurality of user-specific token sets, and attention weights are computed between each token in each user-specific token set using the embeddings for each token.
2. The personalized natural language processing system of claim 1 , wherein the NLP model is a text classification model; and the personalized classifications are personalized text classifications for each of the plurality of users.
3. The personalized natural language processing system of claim 1 , wherein the NLP model is a text prediction model; and the personalized classifications are personalized text predictions for each of the plurality of users.
4. The personalized natural language processing system of claim 1 , wherein the predetermined user-specific tokens include at least one of consecutive numbers, usernames, random sequences of digits, random sequences of tokens with non-alphanumeric characters, or random sequences of all available tokens in a tokenizer vocabulary.
5. The personalized natural language processing system of claim 1 , wherein the training of the NLP model includes minimizing cross-entropy loss for classification.
6. The personalized natural language processing system of claim 1 , wherein the predetermined user-specific tokens are appended to the beginning and the end of each set of tokenized text data.
7. The personalized natural language processing system of claim 1 , wherein lengths of the predetermined user-specific tokens do not exceed a predetermined number of tokens.
8. The personalized natural language processing system of claim 1 , wherein the NLP model is a transformer sequence classifier, transformer sequence-to-sequence model, or long short-term memory (LSTM) recurrent neural network (RNN) classifier.
9. The personalized natural language processing system of claim 1 , wherein the NLP model is a transformer sequence classifier, and user embedding parameters of the predetermined user-specific tokens are tied to embedding parameters of the transformer sequence classifier.
10. A personalized natural language processing method, comprising:
receiving or retrieve a plurality of sets of raw text data from a plurality of users, respectively;
tokenizing the plurality of sets of raw text data to generate a plurality of sets of tokenized text data for the plurality of users, respectively, each set of tokenized text data including a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text data;
appending predetermined user-specific tokens to the plurality of sets of tokenized text data from the plurality of users, respectively, to generate a plurality of user-specific token sets, each predetermined user-specific token corresponding to one user of the plurality of users;
processing the plurality of user-specific token sets sing a natural language processing (NLP) model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for each of the plurality of user-specific token sets; and
outputting the personalized classifications of the plurality of user-specific token sets, wherein the NLP model is trained using a training data set comprising multiple tuples of training data, each tuple including user identifier data that identifies a particular user, raw text data from the particular user, and ground truth classification data, the trained NLP model being configured to perform personalized text classification and/or text prediction tasks for each user of the multiple users, and during processing of the plurality of user-specific token sets, embeddings are produced for each token in each set of the plurality of user-specific token sets, and attention weights are computed between each token in each user-specific token set using the embeddings for each token.
11. The personalized natural language processing method of claim 10 , wherein the NLP model is a text classification model; and the personalized classifications are personalized text classifications for each of the plurality of users.
12. The personalized natural language processing method of claim 10 , wherein the NLP model is a text prediction model; and the personalized classifications are personalized text predictions for each of the plurality of users.
13. The personalized natural language processing method of claim 10 , wherein the predetermined user-specific tokens comprise one of consecutive numbers, usernames, random sequences of digits, random sequences of tokens with non-alphanumeric characters, or random sequences of all available tokens in a tokenizer vocabulary.
14. The personalized natural language processing method of claim 10 , wherein the training of the NLP model includes minimizing cross-entropy loss for classification.
15. The personalized natural language processing method of claim 10 , wherein the predetermined user-specific tokens are appended to a beginning and an end of each set of tokenized text data.
16. The personalized natural language processing method of claim 10 , wherein the NLP model is a transformer sequence classifier, transformer sequence-to-sequence model, or long short-term memory (LSTM) recurrent neural network (RNN) classifier.
17. The personalized natural language processing method of claim 10 , wherein the NLP model is a transformer sequence classifier, and user embedding parameters of the predetermined user-specific tokens are tied to embedding parameters of the transformer sequence classifier.
18. A personalized natural language processing system comprising:
at least one processor, communicatively coupled to non-volatile memory storing a sentiment analysis model and instructions that, when executed by the processor, cause the processor to:
receive or retrieve a plurality of sets of utterances from a plurality of users, respectively;
tokenize the plurality of sets of utterances to generate a plurality of sets of tokenized text data for the plurality of users, respectively, each set of tokenized text data including a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text data;
append predetermined user-specific tokens to the plurality of sets of tokenized text data from the plurality of users, respectively, to generate a plurality of user-specific token sets, each predetermined user-specific token corresponding to one user of the plurality of users;
process the plurality of user-specific token sets using the sentiment analysis model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for each of the plurality of user-specific token sets; and
output the personalized classifications of the plurality of user-specific token sets, wherein the sentiment analysis model is trained using a training data set comprising multiple tuples of training data, each tuple including user identifier data that identifies a particular user, raw text data from the particular user, and ground truth classification data, the trained NLP model being configured to perform personalized text classification and/or text prediction tasks for each user of the multiple users, during processing of the plurality of user-specific token sets, embeddings are produced for each token in each set of the plurality of user-specific token sets, and attention weights are computed between each token in each user-specific token set using the embeddings for each token, and the personalized classifications include a plurality of sentiment labels including at least a positive sentiment, a neutral sentiment, and a negative sentiment.