IP Library Granted Patent US 12,299,401
Granted Patent B2
US 12,299,401 · App. 17/967,562 · Granted May 13, 2025

Transcript paragraph segmentation and visualization of transcript paragraphs

Inventors: Hanieh Deilamsalehy (Seattle, WA); Aseem Omprakash Agarwala (Seattle, WA); Haoran Cai (Mercer Island, WA); Hijung Shin (Arlington, MA); Joel Richard Brandt (Venice, CA); Lubomira Assenova Dontcheva (Seattle, WA)
Assignee: ADOBE INC.
G06F40/30G06F40/205H04N5/9305G06V20/49G10L15/04G10L25/78G10L25/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,401
App. No.
17/967,562
Granted
May 13, 2025
Kind
B2
Abstract

Embodiments of the present invention provide systems, methods, and computer storage media for segmenting a transcript into paragraphs. In an example embodiment, a transcript is segmented to start a new paragraph whenever there is a change in speaker and/or a long pause in speech. If any remaining paragraphs are longer than a designated length or duration (e.g., 50 or 100 words), each of those paragraphs is segmented using dynamic programming to minimize a cost function that penalizes candidate paragraphs based on divergence from a target paragraph length and/or that rewards candidate paragraphs that group semantically similar sentences. As such, the transcript is visualized, segmented at the identified paragraphs.

Claims (26)

1. One or more computer storage media storing computer-useable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

causing generation of a representation of a paragraph segmentation of a transcript that segments one or more transcript paragraphs into smaller transcript paragraphs by using dynamic programming to minimize a cost function based on a target paragraph length and semantic coherency of text within one or more paragraphs of the paragraph segmentation; and

causing a user interface to present a visualization of the paragraph segmentation of the transcript.

2. The one or more computer storage media of claim 1 , wherein the transcript is of an audio track, and the generation of the representation of the paragraph segmentation is based on speech pauses detected in the audio track using an audio classifier.

3. The one or more computer storage media of claim 1 , wherein the generation of the representation of the paragraph segmentation is based on speech pauses detected using differences in word or sentence timings extracted from the transcript.

4. The one or more computer storage media of claim 1 , wherein the generation of the representation of the paragraph segmentation comprises identifying the one or more transcript paragraphs that are longer than a designated length or duration and, for each identified transcript paragraph, segmenting the identified transcript paragraph using dynamic programming to minimize the cost function.

5. The one or more computer storage media of claim 1 , wherein the generation of the representation of the paragraph segmentation comprises using dynamic programming to minimize the cost function that penalizes candidate segmentations that include candidate paragraphs with detected speech pauses longer than a target length or duration.

6. The one or more computer storage media of claim 1 , wherein the generation of the representation of the paragraph segmentation comprises using dynamic programming to minimize the cost function that rewards candidate paragraphs that group semantically similar sentences.

7. The one or more computer storage media of claim 1 , wherein the generation of the representation of the paragraph segmentation comprises quantifying the semantic coherency by generating a measure of similarity of each pair of sentences in a candidate paragraph and combining the measure of similarity for each pair to generate a measure of paragraph similarity for the candidate paragraph.

8. The one or more computer storage media of claim 1 , wherein the paragraph segmentation segments the transcript at a subset of sentence boundaries identified from the transcript.

9. The one or more computer storage media of claim 1 , wherein the transcript is of an audio track of a video, the user interface is a transcript interface of a video editing application, and the transcript interface is configured to interpret input selecting transcript text, via an interaction with the visualization of the paragraph segmentation of the transcript, as an instruction to select a corresponding video segment of the video.

10. A method comprising:

triggering generation of a representation of a paragraph segmentation of a transcript that segments one or more transcript paragraphs into smaller transcript paragraphs by using dynamic programming to minimize a cost function based on a target paragraph length and semantic coherency; and

causing a user interface to present transcript text of the transcript using the paragraph segmentation.

11. The method of claim 10 , wherein the transcript is of an audio track, and the generation of the representation of the paragraph segmentation is based on speech pauses detected in the audio track using an audio classifier.

12. The method of claim 10 , wherein the generation of the representation of the paragraph segmentation is based on speech pauses detected using differences in word or sentence timings extracted from the transcript.

13. The method of claim 10 , wherein the generation of the representation of the paragraph segmentation comprises identifying the one or more transcript paragraphs that are longer than a designated length or duration and, for each identified transcript paragraph, segmenting the identified transcript paragraph using dynamic programming to minimize the cost function.

14. The method of claim 10 , wherein the generation of the representation of the paragraph segmentation comprises using dynamic programming to minimize the cost function that penalizes candidate segmentations that include candidate paragraphs with detected speech pauses longer than a designated length or duration.

15. The method of claim 10 , wherein the generation of the representation of the paragraph segmentation comprises using dynamic programming to minimize the cost function that rewards candidate paragraphs that group semantically similar sentences.

16. The method of claim 10 , wherein the generation of the representation of the paragraph segmentation comprises quantifying the semantic coherency by generating a measure of similarity of each pair of sentences in a candidate paragraph and combining the measure of similarity for each pair to generate a measure of paragraph similarity for the candidate paragraph.

17. The method of claim 10 , wherein the paragraph segmentation segments the transcript at a subset of sentence boundaries identified from the transcript.

18. The method of claim 10 , wherein the transcript is of an audio track of a video, the user interface is a transcript interface of a video editing application, and the transcript interface is configured to interpret input selecting a portion of the transcript text as an instruction to select a corresponding video segment of the video.

19. A computer system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors, the computer program instructions comprising:

a video interaction engine configured to trigger generation of a representation of a paragraph segmentation of a transcript of a video that segments one or more transcript paragraphs into smaller transcript paragraphs by using dynamic programming to minimize a cost function based on a target paragraph length and semantic coherency; and

a transcript tool configured to cause presentation of transcript text of the transcript using the paragraph segmentation.

20. The computer system of claim 19 , wherein the video interaction engine and the transcript tool are part of a video editing application, and the transcript tool is configured to interpret input selecting a portion of the transcript text as an instruction to select a corresponding video segment of the video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: DEILAMSALEHY, HANIEH; AGARWALA, ASEEM OMPRAKASH; CAI, HAORAN; SHIN, HIJUNG; BRANDT, JOEL RICHARD; DONTCHEVA, LUBOMIRA ASSENOVA
To: ADOBE INC.
Reel/Frame 061780/0537 →
Continuity (1)
Related Publication 20240126994A1 · Apr 18, 2024
References Cited (23)
US 11410038B2 · Lin et al. · 2022 [cited by applicant]
US 20070055695A1 · Dorai · 2007 [cited by examiner]
US 20100223051A1 · Burstein · 2010 [cited by examiner]
US 20150206544A1 · Carter · 2015 [cited by examiner]
US 20220075513A1 · Walker et al. · 2022 [cited by applicant]
US 20220075820A1 · Walker et al. · 2022 [cited by applicant]
US 20220076023A1 · Shin · 2022 [cited by examiner]
US 20220076024A1 · Walker et al. · 2022 [cited by applicant]
US 20220076025A1 · Shin et al. · 2022 [cited by applicant]
US 20220076026A1 · Walker et al. · 2022 [cited by applicant]
US 20220076424A1 · Shin et al. · 2022 [cited by applicant]
US 20220076705A1 · Walker et al. · 2022 [cited by applicant]
US 20220076706A1 · Walker et al. · 2022 [cited by applicant]
US 20220076707A1 · Walker et al. · 2022 [cited by applicant]
“Customizable Framework To Extract Moments Of Interest”, U.S. Appl. No. 17/452,626, filed Oct. 28, 2021. [cited by applicant]
Xiao, X., Kanda, N., Chen, Z., Zhou, T., Yoshioka, T., Chen, S., . . . & Gong, Y. (Jun. 2021). Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020. In ICASSP 2021-2021 IEEE Internati… [cited by applicant]
Liu, Y. C., Han, E., Lee, C., & Stolcke, A. (2021). End-to-end neural diarization: From transformer to conformer. arXiv preprint arXiv:2106.07167.(pp. 1-5). [cited by applicant]
“Speechmatics”. https://docs.speechmatics.com/en/batch-container/speech-api/v6.3.0/#diarization, (30 pages). [cited by applicant]
Alcázar, J. L., Caba, F., Mai, L., Perazzi, F., Lee, J. Y., Arbeláez, P., & Ghanem, B. (2020). Active speakers in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12465-… [cited by applicant]
Descript, “Descript Storyboard”, Retrieved From Internet on Sep. 21, 2022 from URL: <https://www.descript.com/video-editing>, 8 Pages. [cited by applicant]
Descript Storyboard, “Descript Storyboard—The Future of Video Editing”, Retrieved From Internet on Sep. 22, 2022 from URL: <https://www.descript.com/storyboard>, 9 Pages. [cited by applicant]
TypeStudio, “Online Video Editor”, Retrieved From Internet on Sep. 22, 2022 from URL: <https://www.typestudio.co/tool/online-video-editor>, 9 pages. [cited by applicant]
Reduct Video, “Where your team and video work together”, Retrieved From Internet on Sep. 21, 2022 from URL: <https://reduct.video/>, 10 Pages. [cited by applicant]