IP Library Granted Patent US 12,223,962
Granted Patent B2
US 12,223,962 · App. 17/967,502 · Granted Feb 11, 2025

Music-aware speaker diarization for transcripts and text-based video editing

Inventors: Justin Jonathan Salamon (San Francisco, CA); Fabian David Caba Heilbron (Campbell, CA); Xue Bai (Bellevue, WA); Aseem Omprakash Agarwala (Seattle, WA); Hijung Shin (Arlington, MA); Lubomira Assenova Dontcheva (Seattle, WA)
Assignee: ADOBE INC.
G10L15/26G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,962
App. No.
17/967,502
Granted
Feb 11, 2025
Kind
B2
Abstract

Embodiments of the present invention provide systems, methods, and computer storage media for music-aware speaker diarization. In an example embodiment, one or more audio classifiers detect speech and music independently of each other, which facilitates detecting regions in an audio track that contain music but do not contain speech. These music-only regions are compared to the transcript, and any transcription and speakers that overlap in time with the music-only regions are removed from the transcript. In some embodiments, rather than having the transcript display the text from this detected music, a visual representation of the audio waveform is included in the corresponding regions of the transcript.

Claims (34)

1. One or more computer storage media storing computer-useable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

detecting, using one or more audio classifiers, (i) speech regions of an audio track that contain detected speech and (ii) music regions of the audio track that contain detected music;

detecting music-only regions based at least on comparing times of the music regions to times of the speech regions;

identifying transcribed singing in a transcript of the audio track based on corresponding sentences overlapping with the music-only regions; and

causing presentation of a visualization of the transcript that omits the transcribed singing from transcript text of the transcript.

2. The one or more computer storage media of claim 1 , wherein detecting the speech regions comprises generating a detection curve representing likelihood over time that speech is present in the audio track, and detecting start and stop times for the speech regions by applying smoothing and thresholding to the detection curve.

3. The one or more computer storage media of claim 1 , the operations further comprising merging adjacent speech regions of the speech regions, or merging adjacent music regions of the music regions, within a threshold temporal spacing of each other.

4. The one or more computer storage media of claim 1 , the operations further comprising removing a set of the music-only regions that are shorter than a threshold duration.

5. The one or more computer storage media of claim 1 ,

wherein the visualization of the transcript includes a visual representation of an audio waveform of the transcribed singing in the transcript.

6. The one or more computer storage media of claim 1 ,

wherein the visualization of the transcript includes a soundbar of the transcribed singing in a row of the transcript without any of the transcript text.

7. The one or more computer storage media of claim 1 ,

wherein the visualization of the transcript includes a spatially condensed visual representation of an audio waveform of the transcribed singing.

8. The one or more computer storage media of claim 1 ,

wherein the visualization of the transcript includes a visual representation of an audio waveform of the transcribed singing annotated with a label identifying a classification of the audio waveform.

9. A method comprising:

generating a representation of speech events identifying regions of an audio track that contain detected speech, and music events identifying regions of the audio track that contain detected music;

detecting music-only events identifying regions of the audio track where the music events do not overlap the speech events;

identifying transcribed singing in a transcript of the audio track based on corresponding sentences overlapping with the music-only events; and

causing presentation of a visualization of the transcript that omits the transcribed singing from transcript text of the transcript.

10. The method of claim 9 , wherein detecting the speech events comprises generating a detection curve representing likelihood over time that speech is present in the audio track, and detecting start and stop times for the speech events by applying smoothing and thresholding to the detection curve.

11. The method of claim 9 , further comprising merging adjacent speech events of the speech events, or merging adjacent music events of the music events, within a threshold temporal spacing of each other.

12. The method of claim 9 , further comprising removing a set of the music-only events that are shorter than a threshold duration.

13. The method of claim 9 , wherein the visualization of the transcript includes a visual representation of an audio waveform of the transcribed singing in the transcript.

14. The method of claim 9 , wherein the visualization of the transcript includes a soundbar of the transcribed singing in a row of the transcript without any of the transcript text.

15. The method of claim 9 , wherein the visualization of the transcript includes a spatially condensed visual representation of an audio waveform of the transcribed singing.

16. The method of claim 9 , wherein the visualization of the transcript includes a visual representation of an audio waveform of the transcribed singing annotated with a label identifying a classification of the audio waveform.

17. A computer system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors, the computer program instructions comprising:

a video interaction engine configured to trigger detecting music-only events that identify regions of an audio track of a video where detected music events do not overlap detected speech events; and

a transcript tool configured to cause presentation of a visualization of a transcript that omits, from transcript text of the transcript, transcribed sentences that overlap with the music-only events.

18. The computer system of claim 17 , wherein the visualization of the transcript includes a visual representation of an audio waveform of the transcribed sentences in the transcript.

19. The computer system of claim 17 , wherein the visualization of the transcript includes a soundbar of the transcribed sentences in a row of the transcript without any of the transcript text.

20. The computer system of claim 17 , wherein the visualization of the transcript includes a spatially condensed visual representation of an audio waveform of the transcribed sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: SALAMON, JUSTIN JONATHAN; CABA HEILBRON, FABIAN DAVID; BAI, XUE; AGARWALA, ASEEM OMPRAKASH; SHIN, HIJUNG; DONTCHEVA, LUBOMIRA ASSENOVA
To: ADOBE INC.
Reel/Frame 061780/0692 →
Continuity (1)
Related Publication 20240127820A1 · Apr 18, 2024
References Cited (22)
US 11410038B2 · Lin et al. · 2022 [cited by applicant]
US 12020708B2 · Bradley · 2024 [cited by examiner]
US 20160267074A1 · Nozue · 2016 [cited by examiner]
US 20220075513A1 · Walker et al. · 2022 [cited by applicant]
US 20220075820A1 · Walker et al. · 2022 [cited by applicant]
US 20220076023A1 · Shin et al. · 2022 [cited by applicant]
US 20220076024A1 · Walker et al. · 2022 [cited by applicant]
US 20220076025A1 · Shin et al. · 2022 [cited by applicant]
US 20220076026A1 · Walker et al. · 2022 [cited by applicant]
US 20220076424A1 · Shin et al. · 2022 [cited by applicant]
US 20220076705A1 · Walker et al. · 2022 [cited by applicant]
US 20220076706A1 · Walker et al. · 2022 [cited by applicant]
US 20220076707A1 · Walker et al. · 2022 [cited by applicant]
US 20240087547A1 · Ivers · 2024 [cited by examiner]
“Customizable Framework To Extract Moments Of Interest”, U.S. Appl. No. 17/452,626, filed Oct. 28, 2021. [cited by applicant]
Xiao, X., Kanda, N., Chen, Z., Zhou, T., Yoshioka, T., Chen, S., . . . & Gong, Y. (Jun. 2021). Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020. In ICASSP 2021-2021 IEEE Internati… [cited by applicant]
Liu, Y. C., Han, E., Lee, C., & Stolcke, A. (2021). End-to-end neural diarization: From transformer to conformer. arXiv preprint arXiv:2106.07167.(pp. 1-5). [cited by applicant]
Alcázar, J. L., Caba, F., Mai, L., Perazzi, F., Lee, J. Y., Arbeláez, P., & Ghanem, B. (2020). Active speakers in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12465-… [cited by applicant]
Descript, “Descript Storyboard”, Retrieved From Internet on Sep. 21, 2022 from URL: <https://www.descript.com/video-editing>, 8 Pages. [cited by applicant]
Descript Storyboard, “Descript Storyboard—The Future of Video Editing”, Retrieved From Internet on Sep. 22, 2022 from URL: <https://www.descript.com/storyboard>, 9 Pages. [cited by applicant]
TypeStudio, “Online Video Editor”, Retrieved From Internet on Sep. 22, 2022 from URL: <https://www.typestudio.co/tool/online-video-editor>, 9 pages. [cited by applicant]
Reduct Video, “Where your team and video work together”, Retrieved From Internet on Sep. 21, 2022 from URL: <https://reduct.video/>, 10 Pages. [cited by applicant]