IP Library › Granted Patent US 12,748,565
Granted Patent B2
US 12,748,565 · App. 17/967,714 · Granted Sep 29, 2026

Transcript question search for text-based video editing

Inventors: Lubomira Assenova Dontcheva (Seattle, WA); Anh Lan Truong (San Carlos, CA); Hanieh Deilamsalehy (Seattle, WA); Kim Pascal Pimmel (Seattle, WA); Aseem Omprakash Agarwala (Seattle, WA); Dingzeyu Li (Seattle, WA); Joel Richard Brandt (Venice, CA); Joy Oakyung Kim (Sunnyvale, CA)
Assignee: ADOBE INC.
G06F3/167G06F3/0482G06F3/0484G06F16/735G06F16/738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,565
App. No.
17/967,714
Granted
Sep 29, 2026
Kind
B2
Abstract

Embodiments of the present invention provide systems, methods, and computer storage media for a question search for meaningful questions that appear in a video. In an example embodiment, an audio track from a video is transcribed, and the transcript is parsed to identify sentences that end with a question mark. Depending on the embodiment, one or more types of questions are filtered out, such as short questions less than a designated length or duration, logistical questions, and/or rhetorical questions. As such, in response to a command to perform a question search, the questions are identified, and search result tiles representing video segments of the questions are presented. Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.

Claims (43)

1 . One or more computer storage media storing computer-useable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

responsive to receiving a command to navigate a video by searchable questions through a video editing interface, identifying the searchable questions asked in the video based on triggering:

identifying questions asked in the video by parsing a diarized transcript of the video;

identifying a subset of the questions asked and answered by a common diarized speaker from the diarized transcript;

encoding each question of the questions asked in the video into a corresponding vector representation of each question;

filtering out the subset from the questions and logistical questions from the questions to output the searchable questions, wherein filtering out the logistical questions is based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and

combining, based on determining a group of consecutive questions of the searchable questions to be within a threshold cosine similarity, the group of consecutive questions into a single question of the searchable questions; and

causing the video editing interface to present a plurality of tiles representing and configured to navigate the video to corresponding video segments during which each of the searchable questions was asked in the video in a search interface of the video editing interface.

2 . The one or more computer storage media of claim 1 , the parsing of the diarized transcript of the video comprising identifying the questions based on the sentences ending with a question mark.

3 . The one or more computer storage media of claim 1 , the searchable questions asked in the video further identified based on triggering:

further identifying the subset based on the questions asked by a particular diarized speaker and not answered by a different diarized speaker in the diarized transcript within a designated amount of time.

4 . The one or more computer storage media of claim 1 , the searchable questions asked in the video further identified based on filtering out short questions that are shorter than a designated duration of time.

5 . The one or more computer storage media of claim 1 , the operations further comprising causing the search interface to present within each of the plurality of tiles: a corresponding video thumbnail comprising a visual representation of a video frame of one of the corresponding video segments during which one of the searchable questions was asked in the video, a corresponding speaker thumbnail representing a diarized speaker detected from the one of the corresponding video segments, and corresponding transcript text of the one of the searchable questions asked in the one of the corresponding video segments from the diarized transcript of the video.

6 . The one or more computer storage media of claim 1 , wherein the search interface of the video editing interface is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of the diarized transcript of the video.

7 . A method comprising:

receiving, via a video editing interface, a command to navigate a loaded video by searchable questions;

responsive to receiving the command to navigate the loaded video by the searchable questions, identifying the searchable questions asked in the loaded video based on triggering:

identifying questions asked in the loaded video by parsing a diarized transcript of the loaded video;

identifying a subset of the questions asked by a particular diarized speaker and not answered by a different diarized speaker in the diarized transcript within a designated amount of time;

encoding each question of the questions asked in the loaded video into a corresponding vector representation of each question; and

filtering out the subset from the questions and logistical questions from the questions to output the searchable questions, wherein filtering out the logistical questions is based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and

causing the video editing interface to present a plurality of tiles representing and configured to navigate the loaded video to corresponding video segments during which each of the searchable questions was asked in the loaded video in a search interface of the video editing interface.

8 . The method of claim 7 , the parsing the diarized transcript of the loaded video comprising identifying the questions based on the sentences ending with a question mark.

9 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on triggering combining a group of consecutive questions of the searchable questions into a single question of the searchable questions based on a determination that the group of consecutive questions are within a threshold cosine similarity.

10 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on triggering:

further identifying the subset based on the questions asked and answered by a common diarized speaker from the diarized transcript.

11 . The method of claim 7 , the searchable questions asked in the loaded video further identified based on filtering out short questions that are shorter than a designated duration of time.

12 . The method of claim 7 , further comprising causing the search interface to present within each of the plurality of tiles: a corresponding video thumbnail comprising a visual representation of a video frame of one of the corresponding video segments during which one of the searchable questions was asked in the loaded video, a corresponding speaker thumbnail representing a diarized speaker detected from the one of the corresponding video segments, and corresponding transcript text of the one of the searchable questions asked in the one of the corresponding video segments from the diarized transcript of the loaded video.

13 . The method of claim 7 , wherein the search interface of the video editing interface is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of the diarized transcript of the loaded video.

14 . A computer system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors, the computer program instructions comprising:

question identifier component configured to receive a command to navigate a loaded video by questions by identifying a plurality of questions asked in the loaded video based on triggering:

identifying an initial set of questions asked in the loaded video by parsing a transcript of the loaded video;

encoding each question of the initial set of questions asked in the loaded video into a corresponding vector representation of each question;

filtering out logistical questions from the initial set of questions asked in the loaded video based on comparing the corresponding vector representation of each question to a sentence embedding generated by combining an encoded representation of example logistical questions into a composite representation of the example logistical questions; and

combining a group of consecutive questions of the filtered initial set of questions into a single question of the plurality of questions based on a determination that the group of consecutive questions are within a threshold cosine similarity; and

a question navigator component configured to present a plurality of tiles representing and configured to navigate the loaded video to corresponding video segments during which each of the identified plurality of questions was asked in the loaded video.

15 . The computer system of claim 14 , the plurality of questions asked in the loaded video further identified based on triggering:

identifying a subset of questions of the initial set of questions asked and answered by a common diarized speaker from a diarized transcript; and

filtering out the subset of questions from the initial set of questions.

16 . The computer system of claim 14 , the plurality of questions asked in the loaded video further identified based on triggering:

identifying a subset of questions of the initial set of questions asked by a particular diarized speaker and not answered by a different diarized speaker in a diarized transcript within a designated amount of time; and

filtering out the subset of questions from the initial set of questions.

17 . The computer system of claim 14 , wherein the question navigator component is configured to navigate, responsive to selection of one of the plurality of tiles, to a corresponding portion of a diarized transcript of the loaded video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2022
From: DONTCHEVA, LUBOMIRA ASSENOVA; TRUONG, ANH LAN; DEILAMSALEHY, HANIEH; PIMMEL, KIM PASCAL; AGARWALA, ASEEM OMPRAKASH; LI, DINGZEYU; BRANDT, JOEL RICHARD; KIM, JOY OAKYUNG
To: ADOBE INC.
Reel/Frame 061718/0738 →
Continuity (1)
Related Publication 20240134597A1 · Apr 25, 2024
References Cited (36)
US 10540906B1 · Fieldman · 2020 [cited by examiner]
US 11410038B2 · Lin et al. · 2022 [cited by applicant]
US 11907315B1 · Cui · 2024 [cited by examiner]
US 20080126191A1 · Schiavi · 2008 [cited by examiner]
US 20100260466A1 · Shima · 2010 [cited by examiner]
US 20130018909A1 · Dicker · 2013 [cited by examiner]
US 20140161416A1 · Chou · 2014 [cited by examiner]
US 20160171003A1 · Ju · 2016 [cited by examiner]
US 20170011774A1 · Ju · 2017 [cited by examiner]
US 20180359530A1 · Marlow · 2018 [cited by examiner]
US 20190079918A1 · Thörn · 2019 [cited by examiner]
US 20200167524A1 · Hunter · 2020 [cited by examiner]
US 20200311164A1 · Golan · 2020 [cited by examiner]
US 20200357009A1 · Podgorny · 2020 [cited by examiner]
US 20210051120A1 · Pottier · 2021 [cited by examiner]
US 20220075513A1 · Walker et al. · 2022 [cited by applicant]
US 20220075820A1 · Walker et al. · 2022 [cited by applicant]
US 20220076023A1 · Shin et al. · 2022 [cited by applicant]
US 20220076024A1 · Walker et al. · 2022 [cited by applicant]
US 20220076025A1 · Shin et al. · 2022 [cited by applicant]
US 20220076026A1 · Walker et al. · 2022 [cited by applicant]
US 20220076424A1 · Shin et al. · 2022 [cited by applicant]
US 20220076705A1 · Walker et al. · 2022 [cited by applicant]
US 20220076706A1 · Walker et al. · 2022 [cited by applicant]
US 20220076707A1 · Walker et al. · 2022 [cited by applicant]
US 20220237637A1 · Susel · 2022 [cited by examiner]
Bhattasali, S., Cytryn, J., Feldman, E., & Park, J. (2015). Automatic Identification of Rhetorical Questions (pp. 743-749). Association for Computational Linguistics. https://aclanthology.org/P15-2122.pdf (Year: 2015). [cited by examiner]
“Customizable Framework to Extract Moments of Interest”, U.S. Appl. No. 17/452,626, filed Oct. 28, 2021. [cited by applicant]
Xiao, X., Kanda, N., Chen, Z., Zhou, T., Yoshioka, T., Chen, S., . . . & Gong, Y. (Jun. 2021). Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020. In ICASSP 2021-2021 IEEE Internati… [cited by applicant]
Liu, Y. C., Han, E., Lee, C., & Stolcke, A. (2021). End-to-end neural diarization: From transformer to conformer. arXiv preprint arXiv:2106.07167.(pp. 1-5). [cited by applicant]
“Speechmatics”. https://docs.speechmatics.com/en/batch-container/speech-api/v6.3.0/#diarization, (30 pages). [cited by applicant]
Alcäzar, J. L., Caba, F., Mai, L., Perazzi, F., Lee, J. Y., Arbeläez, P., & Ghanem, B. (2020). Active speakers in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12465-… [cited by applicant]
Descript, “Descript Storyboard”, Retrieved From Internet on Sep. 21, 2022 from URL: < https://www.descript.com/video-editing>, 8 Pages. [cited by applicant]
Descript Storyboard, “Descript Storyboard—The Future of Video Editing”, Retrieved From Internet on Sep. 22, 2022 from URL: <https://www.descript.com/storyboard>, 9 Pages. [cited by applicant]
TypeStudio, “Online Video Editor”, Retrieved From lntemet on Sep. 22, 2022 from URL: <https://www.typestudio.co/tool/online-video-editor>, 9 pages. [cited by applicant]
Reduct Video, “Where your team and video work together”, Retrieved From Internet on Sep. 21, 2022 from URL: <https://reduct.video/>, 10 Pages. [cited by applicant]