Machine learning dataset generation using a natural language processing technique
A server can receive a plurality of records at a databases such that each record is associated with a phone call and includes at least one request generated based on a transcript of the phone call. The server can generate a training dataset based on the plurality of records. The server can further train a binary classification model using the training dataset. Next, the server can receive a live transcript of a phone call in progress. The server can generate at least one live request based on the live transcript using a natural language processing module of the server. The server can provide the at least one live request to the binary classification model as input to generate a prediction. Lastly, the server can transmit the prediction to an entity receiving the phone call in progress. The prediction can cause a transfer of the call to a chatbot.
1. A method comprising:
receiving, at a server, a plurality of call records, wherein:
each call record includes a call recording, a phone number, a time stamp and a fraud designation; and
the fraud designation is fraudulent or non-fraudulent;
generating, using a processor of the server, a call transcript for each call recording;
creating, using the processor, a training dataset including a plurality of data points, each data point including the call transcript, the phone number, the time stamp and the fraud designation;
training, using the processor, a classification model using the training dataset;
receiving, at the server, a new call record including a new call recording, a new phone number and a new time stamp;
labeling, using the processor, the new call record with a new fraud designation based on a classification by the classification model, wherein the new fraud designation is fraudulent or non-fraudulent; and
transmitting, using the processor, a transfer signal to a device when the new fraud designation is fraudulent, wherein the transfer signal is configured to cause a transfer of a phone call to a chatbot.
2. The method of claim 1 , further comprising multiplying the training dataset using a sampling technique.
3. The method of claim 2 , wherein the sampling technique is undersampling the training dataset or oversampling the training dataset.
4. The method of claim 2 , wherein the sampling technique is one or a combination of: Synthetic Minority Over-sampling Technique; Modified synthetic minority oversampling technique; Random Under-Sampling; or Random Over-Sampling.
5. The method of claim 1 , further comprising generating, using the processor, a background noise for each call record.
6. The method of claim 5 , wherein each data point includes the background noise.
7. The method of claim 6 , wherein the new call record includes the background noise.
8. The method of claim 1 , further comprising transmitting, using the processor, a transfer signal to a device when the new fraud designation is fraudulent, wherein the transfer signal is configured to cause a transfer of a phone call to an Interactive Voice Response (IVR) phone loop.
9. The method of claim 1 , further comprising generating, using the processor, a voice profile for each call record.
10. The method of claim 9 , wherein each data point includes the voice profile.
11. The method of claim 10 , wherein the new call record includes the voice profile.
12. The method of claim 1 , further comprising assigning, using the processor, an intent to the call recording.
13. The method of claim 12 , wherein the intent is determined using a natural language processing module stored on the server.
14. The method of claim 12 , wherein the intent is determined based on keywords used in the call transcript.
15. The method of claim 12 , wherein the intent is determined using an intent recognition module using a machine learning model.
16. The method of claim 15 , wherein the intent recognition module includes preprocessing modules to convert text into character, word, or sentence embeddings that can be fed into the machine learning model.
17. The method of claim 16 , wherein the preprocessing modules are configured for stemming or lemmatization, sentence or word tokenization, and/or stopword removal.
18. The method of claim 1 , further comprising generating, using the processor, an accent for each call record.
19. The method of claim 18 , wherein each data point includes the accent.
20. The method of claim 19 , wherein the new call record includes the accent.