Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method
View Patent ↗Provided are a conversational artificial intelligence (AI) system and method based on real-time multimodal emotion recognition. The system includes a model server configured to provide a machine learning-based conversational model, a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation, and a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal.
1 . A conversational artificial intelligence (AI) system based on real-time multimodal emotion recognition, the conversational AI system comprising:
a model server configured to provide a machine learning-based conversational model;
a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation; and
a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal,
wherein the terminal comprises a voice-text conversion unit configured to convert a voice of the user into text, and
wherein the multimodal empathetic conversation-generation system comprises:
an image emotion recognition unit configured to analyze the facial image from the terminal to recognize an emotion of the user in the facial image;
a text emotion recognition unit configured to assess an emotion of the user inherent in the text;
a composite emotion recognition unit configured to recognize a composite emotion by integrating an emotion obtained from the text and an emotion obtained from the facial image; and
an empathetic expression generation unit configured to control or manipulate, by applying the composite emotion to a rule for converting an emotion recognition result value into a variable for manipulating the expression of the virtual human, the expression of the virtual human displayed on the terminal so as to display an empathetic emotion corresponding to the emotion of the user.
2 . The conversational AI system of claim 1 , wherein the model server comprises reinforcement learning from human feedback (RLHF)-based large language models (LLMs).
3 . The conversational AI system of claim 1 , wherein the terminal comprises:
a capturing unit including a camera photographing a face of the user; and
a recording unit including a microphone generating an electrical voice signal of the user.
4 . The conversational AI system of claim 3 , wherein the terminal further comprises an input window configured to display text returned from the multimodal empathetic conversation-generation system to be correctable.
5 . The conversational AI system of claim 1 , wherein the model server, the terminal, and the multimodal empathetic conversation-generation system are connected to one another through a communication network, and
the communication network is accessed by a database storing information related to a conversation between the user and the virtual human.
6 . A conversation generation method based on real-time multimodal emotion recognition, the conversation generation method comprising:
providing, via a model server, a conversational model according to claim 1 ;
displaying a virtual human through a display and obtaining a facial image of a user and recording a voice of the user via a terminal used by the user to perform a voice conversation with the conversational model;
via a multimodal empathetic conversation-generation system, accessing the model server and receiving a response to a question of the user from the terminal and assessing a composite emotion of the user based on an emotion inherent in the question of the user and the facial image, and based on the assessed composite emotion, controlling an expression of the virtual human displayed on the terminal,
wherein the multimodal empathetic conversation-generation system analyzes the facial image received from the terminal to recognize an emotion of the user in the image, recognizes an emotion inherent in text converted from the voice of the user, recognizes the composite emotion by integrating the emotion inherent in the voice with the emotion obtained from the facial image, and controls or manipulates the expression of the virtual human displayed on the terminal by applying the composite emotion to a rule for converting an emotion recognition result value into a variable for manipulating the expression of the virtual human so as to display an empathetic emotion corresponding to the emotion of the user.
7 . The conversation generation method of claim 6 , wherein the model server comprises reinforcement learning from human feedback (RLHF)-based large language models (LLMs).
8 . The conversation generation method of claim 7 , wherein the terminal is further configured to provide an input window configured to display text returned from the multimodal empathetic conversation-generation system to be correctable.