System and method for generating user specific interactive voice responses based on user speech and voice characteristics
A system includes a memory configured to store user profiles associated with a plurality of users and an interactive voice response (IVR) system configured to service calls. The system includes processors configured to receive a call from a first user, generate a first voice interaction configured to prompt the first user to perform an utterance of a second voice interaction, and detect the utterance of the second voice interaction. The processors are configured to execute a first machine-learning model trained to identify speech and voice characteristics of the first user and to generate a third voice interaction based on the identified speech and voice characteristics. In response to identifying an intent and one or more named entities of the request, the processors are configured to initiate the execution of one or more interactions with the first user profile in accordance with the identified intent and one or more named entities.
1 . A system, comprising:
a memory configured to store a plurality of user profiles associated with a plurality of users and an interactive voice response (IVR) system configured to service calls with respect to the plurality of user profiles; and
one or more processors operably coupled to the memory and configured to:
receive a call from a first user of the plurality of users, wherein the call comprises a potential request to initiate an execution of one or more interactions with a first user profile associated with the first user, and, in response:
generate, based at least in part on the call from the first user, a first voice interaction configured to prompt the first user to perform an utterance of a second voice interaction;
detect, based at least in part on the first voice interaction, the utterance of the second voice interaction performed by the first user;
in response to detecting the utterance of the second voice interaction, execute a first machine-learning model trained to identify speech characteristics and voice characteristics of the first user and to generate, based at least in part on the identified speech characteristics and the identified one or more voice characteristics, a third voice interaction reflective of the identified speech characteristics and the identified voice characteristics, wherein the third voice interaction comprises a personalized generated voice interaction including a speech and a voice personalized to the first user;
provide the third voice interaction to the first user;
execute a second machine-learning model trained to identify an intent and one or more named entities of a request of the first user based at least in part on the second voice interaction and the identified speech characteristics and the identified voice characteristics; and
in response to identifying the intent and the one or more named entities of the request of the first user, initiate the execution of the one or more interactions with the first user profile in accordance with the identified intent and the one or more named entities of the request.
2 . The system of claim 1 , wherein the first machine-learning model comprises a first natural language processing (NLP) model trained or fine-tuned based on the identified speech characteristics and the identified voice characteristics.
3 . The system of claim 2 , wherein the first natural language processing (NLP) model comprises one or more of a bidirectional and auto-regressive transformer (BART) model, a bidirectional encoder representations for transformer (BERT) model, a knowledge enhanced bidirectional encoder representations for transformer (KnowBERT) model, a robustly optimized bidirectional encoder representations for transformer pretraining approach (ROBERTa) model, or a generative pre-trained transformer (GPT) model.
4 . The system of claim 1 , wherein the second machine-learning model comprises a second natural language processing (NLP) model pretrained to identify intent and one or more named entities from a plurality of different utterances of voice interactions performed by the plurality of users.
5 . The system of claim 1 , wherein the identified speech characteristics comprises one or more of a language, an accent, a dialect, a speech context, a speech complexity, a pause rate, a word length, a word frequency, a syntactic depth, a use of particles, a use of nouns, or a use of pronouns.
6 . The system of claim 1 , wherein the identified voice characteristics comprises one or more of a tone, a pitch, a volume, a tempo, a timbre, a rate, a voice type, or a voice register.
7 . The system of claim 1 , wherein the personalized generated voice interaction includes the speech, the voice, and a speech rate pattern personalized to the first user.
8 . The system of claim 1 , wherein the one or more processors are further configured to initiate the execution of the one or more interactions with the first user profile to execute a predetermined action.
9 . A method, comprising:
receiving a call from a first user of a plurality of users, wherein the call comprises a potential request to initiate an execution of one or more interactions with a first user profile of a plurality of user profiles associated with a plurality of users, wherein the first user profile is associated with a first user, and wherein the call is received by an interactive voice response (IVR) system configured to service calls with respect to the plurality of user profiles, and, in response:
generating, based at least in part on the call from the first user, a first voice interaction configured to prompt the first user to perform an utterance of a second voice interaction;
detecting, based at least in part on the first voice interaction, the utterance of the second voice interaction performed by the first user;
in response to detecting the utterance of the second voice interaction, executing a first machine-learning model trained to identify speech characteristics and voice characteristics of the first user and to generate, based at least in part on the identified speech characteristics and the identified voice characteristics, a third voice interaction reflective of the identified speech characteristics and the identified voice characteristics, wherein the third voice interaction comprises a personalized generated voice interaction including a speech and a voice personalized to the first user;
providing the third voice interaction to the first user;
executing a second machine-learning model trained to identify an intent and one or more named entities of a request of the first user based at least in part on the second voice interaction and the identified speech characteristics and the identified voice characteristics; and
in response to identifying the intent and the one or more named entities of the request of the first user, initiating the execution of the one or more interactions with the first user profile in accordance with the identified intent and the one or more named entities of the request.
10 . The method of claim 9 , wherein the first machine-learning model comprises a first natural language processing (NLP) model trained or fine-tuned based on the identified speech characteristics and the identified voice characteristics.
11 . The method of claim 10 , wherein the first natural language processing (NLP) model comprises one or more of a bidirectional and auto-regressive transformer (BART) model, a bidirectional encoder representations for transformer (BERT) model, a knowledge enhanced bidirectional encoder representations for transformer (KnowBERT) model, a robustly optimized bidirectional encoder representations for transformer pretraining approach (ROBERTa) model, or a generative pre-trained transformer (GPT) model.
12 . The method of claim 9 , wherein the second machine-learning model comprises a second natural language processing (NLP) model pretrained to identify intent and one or more named entities from a plurality of different utterances of voice interactions performed by the plurality of users.
13 . The method of claim 9 , wherein the identified speech characteristics comprises one or more of a language, an accent, a dialect, a speech context, a speech complexity, a pause rate, a word length, a word frequency, a syntactic depth, a use of particles, a use of nouns, or a use of pronouns.
14 . The method of claim 9 , wherein the identified voice characteristics comprises one or more of a tone, a pitch, a volume, a tempo, a timbre, a rate, a voice type, or a voice register.
15 . The method of claim 9 , wherein personalized generated voice interaction includes the speech, the voice, and a speech rate pattern personalized to the first user.
16 . The method of claim 9 , wherein initiating the execution of the one or more interactions with the first user profile comprises initiating the execution of the one or more interactions with the first user profile to execute a predetermined action.
17 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:
receive a call from a first user of a plurality of users, wherein the call comprises a potential request to initiate an execution of one or more interactions with a first user profile of a plurality of user profiles associated with a plurality of users, wherein the first user profile is associated with a first user, and wherein the call is received by an interactive voice response (IVR) system configured to service calls with respect to the plurality of user profiles, and, in response:
generate, based at least in part on the call from the first user, a first voice interaction configured to prompt the first user to perform an utterance of a second voice interaction;
detect, based at least in part on the first voice interaction, the utterance of the second voice interaction performed by the first user;
in response to detecting the utterance of the second voice interaction, execute a first machine-learning model trained to identify speech characteristics and voice characteristics of the first user and to generate, based at least in part on the identified speech characteristics and the identified voice characteristics, a third voice interaction reflective of the identified speech characteristics and the identified voice characteristics, wherein the third voice interaction comprises a personalized generated voice interaction including a speech and a voice personalized to the first user;
provide the third voice interaction to the first user;
execute a second machine-learning model trained to identify an intent and one or more named entities of a request of the first user based at least in part on the second voice interaction and the identified speech characteristics and the identified voice characteristics; and
in response to identifying the intent and the one or more named entities of the request of the first user, initiate the execution of the one or more interactions with the first user profile in accordance with the identified intent and the one or more named entities of the request.
18 . The non-transitory computer-readable medium of claim 17 , wherein the first machine-learning model comprises a first natural language processing (NLP) model trained or fine-tuned based on the identified speech characteristics and the identified voice characteristics.
19 . The non-transitory computer-readable medium of claim 18 , wherein the first natural language processing (NLP) model comprises one or more of a bidirectional and auto-regressive transformer (BART) model, a bidirectional encoder representations for transformer (BERT) model, a knowledge enhanced bidirectional encoder representations for transformer (KnowBERT) model, a robustly optimized bidirectional encoder representations for transformer pretraining approach (ROBERTa) model, or a generative pre-trained transformer (GPT) model.
20 . The non-transitory computer-readable medium of claim 17 , wherein the second machine-learning model comprises a second natural language processing (NLP) model pretrained to identify intent and one or more named entities from a plurality of different utterances of voice interactions performed by the plurality of users.