Detection of interaction events in recorded audio streams
Detection of interaction events in recorded audio streams is disclosed, including: detecting an interaction event within a recorded audio stream; analyzing text before and after the interaction event in the recorded audio stream to determine a causer of the interaction event; and determining an action to be performed in response to the interaction event based at least in part on the causer of the interaction event.
1 . A system, comprising:
a processor configured to:
detect an interaction event within a recorded audio stream, wherein to detect the interaction event within the recorded audio stream comprises to:
divide the recorded audio stream into a plurality of audio frames;
determine an audio segment based at least in part on audio frames that are included in a sliding window;
determine that a subset of audio frames within the audio segment includes the interaction event; and
determine a time segment within the recorded audio stream corresponding to a duration during which the interaction event is detected within the recorded audio stream, wherein the time segment comprises a start timestamp that denotes a start associated with the interaction event and an end timestamp that denotes an end associated with the interaction event;
obtain a text transcription of the recorded audio stream, wherein the text transcription is time aligned with speech in the recorded audio stream and includes annotations of participants for respective speaking turns;
analyze contextual text comprising a first portion of text in the text transcription that precedes the time segment associated with the interaction event and a second portion of text in the text transcription that follows the time segment associated with the interaction event in the recorded audio stream to determine a causer of the interaction event, wherein the first portion of text in the text transcription comprises a first predetermined number of words in the text transcription that precedes the start timestamp associated with the interaction event, and wherein the second portion of text in the text transcription comprises a second predetermined number of words in the text transcription that follows the end timestamp associated with the interaction event, wherein to determine the causer of the interaction event comprises to input the contextual text comprising the first portion of text and the second portion of text into a text-based interaction event classification model, wherein the text-based interaction event classification model is configured to output information indicating the causer of the interaction event, wherein the text-based interaction event classification model has been trained on training data comprising labeled text snippets associated with recorded interactions, wherein a labeled text snippet includes: contextual text surrounding a detected interaction event within the labeled text snippet, a set of first labels indicating which speaking turn in the contextual text belongs to which of a first participant role and a second participant role in a corresponding recorded interaction, a second label indicating which of the first participant role and the second participant role had caused the detected interaction event within the labeled text snippet, and a third label indicating whether the detected interaction event within the labeled text snippet was expected or unexpected; and
determine an action to be performed in response to the interaction event based at least in part on the causer of the interaction event; and
a memory coupled to the processor and configured to provide the processor with instructions.
2 . The system of claim 1 , wherein the processor is further configured to determine whether the interaction event is associated with a predetermined type.
3 . The system of claim 2 , wherein the predetermined type is silence.
4 . The system of claim 1 , wherein to determine whether the audio segment includes the interaction event comprises to input the audio segment into an audio-based interaction event detection machine learning model.
5 . The system of claim 1 , wherein the processor is configured to start a new clock in response to the detection of the interaction event.
6 . The system of claim 5 , wherein the duration comprises a first duration, and wherein the processor is configured to determine that a second duration indicated by the new clock is greater than a threshold duration and send a prompt to a client device.
7 . The system of claim 1 , wherein the first portion of text and the second portion of text comprises the contextual text associated with the interaction event.
8 . The system of claim 1 , wherein the text-based interaction event classification model is further configured to output whether the interaction event was a first predefined classification or a second predefined classification.
9 . The system of claim 1 , wherein the processor is further configured to determine an indicated duration of the interaction event within text before the interaction event.
10 . The system of claim 9 , wherein the processor is further configured to:
determine an actual duration associated with the interaction event;
compare the actual duration to the indicated duration; and
in response to a determination that the actual duration is greater than the indicated duration, send a prompt to a client device associated with a participant in the recorded audio stream.
11 . The system of claim 9 , wherein the processor is further configured to:
determine an actual duration associated with the interaction event;
compare the actual duration to the indicated duration; and
in response to a determination that the actual duration is greater than the indicated duration, store data corresponding to the interaction event and wherein the stored data is configured to be used for report generation, analytics, or downstream processing.
12 . The system of claim 9 , wherein the processor is further configured to:
determine an actual duration associated with the interaction event;
compare the actual duration to the indicated duration; and
in response to a determination that the actual duration is greater than the indicated duration, send a prompt to a supervisor of a participant in the recorded audio stream.
13 . The system of claim 12 , wherein the prompt indicates that the interaction event comprises a hold violation.
14 . The system of claim 1 , wherein the action to be performed comprises to send a prompt to a supervisor of a first speaking role in the recorded audio stream to indicate one or more of the following: that a second speaking role in the recorded audio stream was exhibiting a negative sentiment or that the second speaking role was exhibiting a lack of comprehension over speech by the first speaking role.
15 . The system of claim 1 , wherein the recorded audio stream comprises a live recording of audio.
16 . The system of claim 1 , wherein the recorded audio stream is associated with a videoconference-based meeting.
17 . The system of claim 1 , wherein the first predetermined number of words is different from the second predetermined number of words.
18 . The system of claim 1 , wherein the labeled text snippet further includes a fourth label comprising an explanation for why whichever of the first participant role and the second participant role had caused the detected interaction event within the labeled text snippet.
19 . A method, comprising:
detecting an interaction event within a recorded audio stream, wherein detecting the interaction event within the recorded audio stream comprises:
dividing the recorded audio stream into a plurality of audio frames;
determining an audio segment based at least in part on audio frames that are included in a sliding window;
determining that a subset of audio frames within the audio segment includes the interaction event; and
determining a time segment within the recorded audio stream corresponding to a duration during which the interaction event is detected within the recorded audio stream, wherein the time segment comprises a start timestamp that denotes a start associated with the interaction event and an end timestamp that denotes an end associated with the interaction event;
obtaining a text transcription of the recorded audio stream, wherein the text transcription is time aligned with speech in the recorded audio stream and includes annotations of participants for respective speaking turns;
analyzing contextual text comprising a first portion of text in the text transcription that precedes the time segment associated with the interaction event and a second portion of text in the text transcription that follows the time segment associated with the interaction event in the recorded audio stream to determine a causer of the interaction event, wherein the first portion of text in the text transcription comprises a first predetermined number of words in the text transcription that precedes the start timestamp associated with the interaction event, and wherein the second portion of text in the text transcription comprises a second predetermined number of words in the text transcription that follows the end timestamp associated with the interaction event, wherein to determine the causer of the interaction event comprises to input the contextual text comprising the first portion of text and the second portion of text into a text-based interaction event classification model, wherein the text-based interaction event classification model is configured to output information indicating the causer of the interaction event, wherein the text-based interaction event classification model has been trained on training data comprising labeled text snippets associated with recorded interactions, wherein a labeled text snippet includes: contextual text surrounding a detected interaction event within the labeled text snippet, a set of first labels indicating which speaking turn in the contextual text belongs to which of a first participant role and a second participant role in a corresponding recorded interaction, a second label indicating which of the first participant role and the second participant role had caused the detected interaction event within the labeled text snippet, and a third label indicating whether the detected interaction event within the labeled text snippet was expected or unexpected; and
determining an action to be performed in response to the interaction event based at least in part on the causer of the interaction event.
20 . A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:
detecting an interaction event within a recorded audio stream, wherein detecting the interaction event within the recorded audio stream comprises:
dividing the recorded audio stream into a plurality of audio frames;
determining an audio segment based at least in part on audio frames that are included in a sliding window;
determining that a subset of audio frames within the audio segment includes the interaction event; and
determining a time segment within the recorded audio stream corresponding to a duration during which the interaction event is detected within the recorded audio stream, wherein the time segment comprises a start timestamp that denotes a start associated with the interaction event and an end timestamp that denotes an end associated with the interaction event;
obtaining a text transcription of the recorded audio stream, wherein the text transcription is time aligned with speech in the recorded audio stream and includes annotations of participants for respective speaking turns;
analyzing contextual text comprising a first portion of text in the text transcription that precedes the time segment associated with the interaction event and a second portion of text in the text transcription that follows the time segment associated with the interaction event in the recorded audio stream to determine a causer of the interaction event, wherein the first portion of text in the text transcription comprises a first predetermined number of words in the text transcription that precedes the start timestamp associated with the interaction event, and wherein the second portion of text in the text transcription comprises a second predetermined number of words in the text transcription that follows the end timestamp associated with the interaction event, wherein to determine the causer of the interaction event comprises to input the contextual text comprising the first portion of text and the second portion of text into a text-based interaction event classification model, wherein the text-based interaction event classification model is configured to output information indicating the causer of the interaction event, wherein the text-based interaction event classification model has been trained on training data comprising labeled text snippets associated with recorded interactions, wherein a labeled text snippet includes: contextual text surrounding a detected interaction event within the labeled text snippet, a set of first labels indicating which speaking turn in the contextual text belongs to which of a first participant role and a second participant role in a corresponding recorded interaction, a second label indicating which of the first participant role and the second participant role had caused the detected interaction event within the labeled text snippet, and a third label indicating whether the detected interaction event within the labeled text snippet was expected or unexpected; and
determining an action to be performed in response to the interaction event based at least in part on the causer of the interaction event.