Transcript question search for text-based video editing
Embodiments of the present invention provide systems, methods, and computer storage media for a question search for meaningful questions that appear in a video. In an example embodiment, an audio track from a video is transcribed, and the transcript is parsed to identify sentences that end with a question mark. Depending on the embodiment, one or more types of questions are filtered out, such as short questions less than a designated length or duration, logistical questions, and/or rhetorical questions. As such, in response to a command to perform a question search, the questions are identified, and search result tiles representing video segments of the questions are presented. Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.
1 . One or more computer storage media storing computer-useable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:
responsive to receiving a command to navigate a video by searchable questions through a video editing interface, identifying the searchable questions asked in the video based on triggering:
identifying questions asked in the video by parsing a diarized transcript of the video;
identifying a subset of the questions asked and answered by a common diarized speaker from the diarized transcript;
encoding each question of the questions asked in the video into a corresponding vector representation of each question;
filtering out the subset from the questions and logistical questions from the questions to output the searchable questions, wherein filtering out the logistical questions is based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and
combining, based on determining a group of consecutive questions of the searchable questions to be within a threshold cosine similarity, the group of consecutive questions into a single question of the searchable questions; and
causing the video editing interface to present a plurality of tiles representing and configured to navigate the video to corresponding video segments during which each of the searchable questions was asked in the video in a search interface of the video editing interface.
2 . The one or more computer storage media of claim 1 , the parsing of the diarized transcript of the video comprising identifying the questions based on the sentences ending with a question mark.
3 . The one or more computer storage media of claim 1 , the searchable questions asked in the video further identified based on triggering:
further identifying the subset based on the questions asked by a particular diarized speaker and not answered by a different diarized speaker in the diarized transcript within a designated amount of time.
4 . The one or more computer storage media of claim 1 , the searchable questions asked in the video further identified based on filtering out short questions that are shorter than a designated duration of time.
5 . The one or more computer storage media of claim 1 , the operations further comprising causing the search interface to present within each of the plurality of tiles: a corresponding video thumbnail comprising a visual representation of a video frame of one of the corresponding video segments during which one of the searchable questions was asked in the video, a corresponding speaker thumbnail representing a diarized speaker detected from the one of the corresponding video segments, and corresponding transcript text of the one of the searchable questions asked in the one of the corresponding video segments from the diarized transcript of the video.
6 . The one or more computer storage media of claim 1 , wherein the search interface of the video editing interface is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of the diarized transcript of the video.
7 . A method comprising:
receiving, via a video editing interface, a command to navigate a loaded video by searchable questions;
responsive to receiving the command to navigate the loaded video by the searchable questions, identifying the searchable questions asked in the loaded video based on triggering:
identifying questions asked in the loaded video by parsing a diarized transcript of the loaded video;
identifying a subset of the questions asked by a particular diarized speaker and not answered by a different diarized speaker in the diarized transcript within a designated amount of time;
encoding each question of the questions asked in the loaded video into a corresponding vector representation of each question; and
filtering out the subset from the questions and logistical questions from the questions to output the searchable questions, wherein filtering out the logistical questions is based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and
causing the video editing interface to present a plurality of tiles representing and configured to navigate the loaded video to corresponding video segments during which each of the searchable questions was asked in the loaded video in a search interface of the video editing interface.
8 . The method of claim 7 , the parsing the diarized transcript of the loaded video comprising identifying the questions based on the sentences ending with a question mark.
9 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on triggering combining a group of consecutive questions of the searchable questions into a single question of the searchable questions based on a determination that the group of consecutive questions are within a threshold cosine similarity.
10 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on triggering:
further identifying the subset based on the questions asked and answered by a common diarized speaker from the diarized transcript.
11 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on filtering out short questions that are shorter than a designated duration of time.
12 . The method of claim 7 , further comprising causing the search interface to present within each of the plurality of tiles: a corresponding video thumbnail comprising a visual representation of a video frame of one of the corresponding video segments during which one of the searchable questions was asked in the loaded video, a corresponding speaker thumbnail representing a diarized speaker detected from the one of the corresponding video segments, and corresponding transcript text of the one of the searchable questions asked in the one of the corresponding video segments from the diarized transcript of the loaded video.
13 . The method of claim 7 , wherein the search interface of the video editing interface is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of the diarized transcript of the loaded video.
14 . A computer system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors, the computer program instructions comprising:
question identifier component configured to receive a command to navigate a loaded video by questions by identifying a plurality of questions asked in the loaded video based on triggering:
identifying an initial set of questions asked in the loaded video by parsing a transcript of the loaded video;
encoding each question of the initial set of questions asked in the loaded video into a corresponding vector representation of each question;
filtering out logistical questions from the initial set of questions asked in the loaded video based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and
combining a group of consecutive questions of the filtered initial set of questions into a single question of the plurality of questions based on a determination that the group of consecutive questions are within a threshold cosine similarity; and
a question navigator component configured to present a plurality of tiles representing and configured to navigate the loaded video to corresponding video segments during which each of the identified plurality of questions was asked in the loaded video.
15 . The computer system of claim 14 , the plurality of questions asked in the loaded video further identified based on triggering:
identifying a subset of questions of the initial set of questions asked and answered by a common diarized speaker from a diarized transcript; and
filtering out the subset of questions from the initial set of questions.
16 . The computer system of claim 14 , the plurality of questions asked in the loaded video further identified based on triggering:
identifying a subset of questions of the initial set of questions asked by a particular diarized speaker and not answered by a different diarized speaker in a diarized transcript within a designated amount of time; and
filtering out the subset of questions from the initial set of questions.
17 . The computer system of claim 14 , wherein the question navigator component is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of a diarized transcript of the loaded video.