Virtual meeting coaching with dynamically extracted content
Methods and systems provide for virtual meeting coaching with dynamically extracted content. In one embodiment, the system receives transcripts for a plurality of selected meetings from within a communication platform, the meetings being selected for question extraction; extracts, via a question extraction model, questions from the transcripts for the selected meetings; extracts, from the questions, a number of expected answers, each expected answer associated with one of the questions; connects to a coaching session with one or more participants and a virtual coaching agent; transmits, by the virtual coaching agent, at least a subset of the extracted questions to the participants; receives answers to the subset of the extracted questions by participants; and generates one or more evaluation scores for the participants corresponding to each of the answers, the evaluation scores being based at least in part on the extracted expected answers associated with the subset of the extracted questions.
1 . A computer-implemented method, comprising:
receiving, by a processing engine integrated with a communication platform, transcripts for a plurality of selected meetings from within a communication platform, the meetings being selected for question extraction;
extracting, via a question extraction model, a plurality of questions from the transcripts for the selected meetings;
extracting, from the plurality of questions, a plurality of expected answers, each expected answer associated with one of the questions, wherein extracting the plurality of expected answers comprises applying a machine learning model to the transcripts to generate, for each question, a set of key points and corresponding conversation sentences representing an expected answer;
connecting to a coaching session comprising one or more participants using one or more client devices and a virtual coaching agent;
transmitting at least a subset of the extracted questions to at least one of the client devices, the extracted questions being transmitted as uttered by the virtual coaching agent;
receiving, from the at least one client device, answers to the subset of the extracted questions by at least one participant of the one or more participants, wherein the answers to the questions by the one or more participants each comprises audio and video of the answering participant, and wherein the virtual coaching agent is represented in another video by a digital avatar;
generating one or more evaluation scores for the at least one participant of the one or more participants corresponding to each of the answers, the evaluation scores being based at least in part on the extracted expected answers associated with the subset of the extracted questions, wherein generating the one or more evaluation scores comprises, for each answer, automatically computing a content coverage score based on a text embedding similarity between a participant's utterances and the extracted expected answer for the corresponding question using the machine learning model, and wherein generating the one or more evaluation scores for each of the answers is further based at least in part on evaluating a style of the audio and video of the answering participant; and
transmitting the evaluation score in real time to a corresponding client device.
2 . The method of claim 1 , wherein the plurality of extracted questions and the plurality of expected answers both comprise one or more scenarios, where the extracted questions and answers within each scenario all relate to a common context.
3 . The method of claim 2 , wherein the one or more scenarios and the common context for each scenario are automatically generated based on the content of the questions and expected answers.
4 . The method of claim 1 , wherein at least one of the participants of the coaching session was assigned to the coaching session by a user of the communication platform authorized to assign coaching sessions to the participant.
5 . The method of claim 1 , wherein at least one of the participants of the coaching session was automatically assigned to the coaching session automatically based on a performance by the participant within one or more communication sessions previously attended by the participant.
6 . The method of claim 1 , further comprising:
for each received answer from a participant, receiving a transcript of utterances spoken by the participant during the answer,
wherein generating the one or more evaluation scores for each of the answers is further based at least in part on comparing the extracted expected answers associated with the subset of the extracted questions to the utterances for the received answer.
7 . The method of claim 1 , further comprising:
determining an overall evaluation score for the coaching session based on the generated evaluation scores for the answers; and
transmitting, to at least the client device, the overall evaluation score for the coaching session.
8 . The method of claim 1 , wherein the evaluation scores for each of the answers are transmitted in real time to at least a corresponding client device of the answering participant upon receiving the answer.
9 . The method of claim 1 , wherein at least one of the evaluation scores for each question represents a percentage of content coverage for the question, and further comprising:
transmitting, to at least the client device in real time upon receiving the answer, the percentage of content coverage for the question.
10 . The method of claim 1 , wherein each of the extracted expected answers comprises one or more key points, each of the key points comprising a headline and a conversation sentence.
11 . The method of claim 1 , wherein extracting the plurality of questions and extracting the plurality of expected answers both are performed via machine learning (ML) techniques.
12 . The method of claim 1 , further comprising:
extracting, via a question extraction model, a plurality of expected questions from the transcripts for the selected meetings, the expected questions representing questions the one or more participants may ask the virtual coaching agent.
13 . The method of claim 12 further comprising:
extracting, from video of the selected meetings, one or more expected visual expressions for answers.
14 . The method of claim 13 , further comprising:
receiving, from the client device, a question from one of the participants to the virtual coaching agent;
determining a similarity match of the question from the participant to an expected question from the plurality of expected questions; and
transmitting, to the client device, the answer associated with the expected question, the answer being transmitted as uttered by the virtual coaching agent.
15 . A communication system comprising one or more processors configured to:
receive, by a processing engine integrated with a communication platform, transcripts for a plurality of selected meetings from within a communication platform, the meetings being selected for question extraction;
extract, via a question extraction model, a plurality of questions from the transcripts for the selected meetings;
extract, from the plurality of questions, a plurality of expected answers, each expected answer associated with one of the questions, wherein extracting the plurality of expected answers comprises applying a machine learning model to the transcripts to generate, for each question, a set of key points and corresponding conversation sentences representing an expected answer;
connect to a coaching session comprising one or more participants using one or more client devices and a virtual coaching agent;
transmit at least a subset of the extracted questions to at least one of the client devices, the extracted questions being transmitted as uttered by the virtual coaching agent;
receive, from the at least one client device, answers to the subset of the extracted questions by at least one participant of the one or more participants, wherein the answers to the questions by the one or more participants each comprises audio and video of the answering participant, and wherein the virtual coaching agent is represented in another video by a digital avatar;
generate one or more evaluation scores for the at least one participant of the one or more participants corresponding to each of the answers, the evaluation scores being based at least in part on the extracted expected answers associated with the subset of the extracted questions, wherein generating the one or more evaluation scores comprises, for each answer, automatically computing a content coverage score based on a text embedding similarity between a participant's utterances and the extracted expected answer for the corresponding question using the machine learning model, and wherein generating the one or more evaluation scores for each of the answers is further based at least in part on evaluating a style of the audio and video of the answering participant; and
transmit the evaluation score in real time to a corresponding client device.
16 . The communication system of claim 15 , wherein the one or more processors are further configured to:
prior to transmitting a next question to the client device, determine, via one or more ML techniques, that the answer has terminated.
17 . The communication system of claim 15 , wherein the one or more processors are further configured to:
extract, from audio of the selected meetings, one or more expected tones for one or more answers, each answer comprising audio of one of the participants, and generate the one or more evaluation scores for each of the answers being further based on comparing one or more tones from the audio of the participant in the answer to the one or more expected tones for answers.
18 . The communication system of claim 15 , wherein the one or more processors are further configured to:
extract, from video of the selected meetings, one or more expected visual expressions for answers,
each answer comprising video of one of the participants, and generate the one or more evaluation scores for each of the answers being further based on comparing one or more visual expressions from the video of the participant in the answer to the one or more expected visual expressions for answers.
19 . A non-transitory computer-readable medium containing instructions, that when executed by one or more processors, cause the one or more processors to perform operations comprising:
receiving, by a processing engine integrated with a communication platform, transcripts for a plurality of selected meetings from within a communication platform, the meetings being selected for question extraction;
extracting, via a question extraction model, a plurality of questions from the transcripts for the selected meetings;
extracting, from the plurality of questions, a plurality of expected answers, each expected answer associated with one of the questions, wherein extracting the plurality of expected answers comprises applying a machine learning model to the transcripts to generate, for each question, a set of key points and corresponding conversation sentences representing an expected answer;
connecting to a coaching session comprising one or more participants using one or more client devices and a virtual coaching agent;
transmitting at least a subset of the extracted questions to at least one of the client devices, the extracted questions being transmitted as uttered by the virtual coaching agent;
receiving, from the at least one client device, answers to the subset of the extracted questions by at least one participant of the one or more participants, wherein the answers to the questions by the one or more participants each comprises audio and video of the answering participant, and wherein the virtual coaching agent is represented in another video by a digital avatar;
generating one or more evaluation scores for the at least one participant of the one or more participants corresponding to each of the answers, the evaluation scores being based at least in part on the extracted expected answers associated with the subset of the extracted questions, wherein generating the one or more evaluation scores comprises, for each answer, automatically computing a content coverage score based on a text embedding similarity between a participant's utterances and the extracted expected answer for the corresponding question using the machine learning model, and wherein generating the one or more evaluation scores for each of the answers is further based at least in part on evaluating a style of the audio and video of the answering participant; and
transmitting the evaluation score in real time to a corresponding client device.
20 . The non-transitory computer-readable medium of claim 19 , wherein the plurality of extracted questions and the plurality of expected answers both comprise one or more scenarios, where the extracted questions and answers within each scenario all relate to a common context.