IP Library Granted Patent US 11,875,781
Granted Patent B2
US 11,875,781 · App. 17/008,427 · Granted Jan 16, 2024

Audio-based media edit point selection

Inventors: Amol Jindal (San Jose, CA); Somya Jain (Uttar Pradesh, IN); Ajay Bedi (Himchal Pradesh, IN)
Assignee: Adobe Inc.
G10L15/04G06F40/253G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,875,781
App. No.
17/008,427
Granted
Jan 16, 2024
Kind
B2
Abstract

A media edit point selection process can include a media editing software application programmatically converting speech to text and storing a timestamp-to-text map. The map correlates text corresponding to speech extracted from an audio track for the media clip to timestamps for the media clip. The timestamps correspond to words and some gaps in the speech from the audio track. The probability of identified gaps corresponding to a grammatical pause by the speaker is determined using the timestamp-to-text map and a semantic model. Potential edit points corresponding to grammatical pauses in the speech are stored for display or for additional use by the media editing software application. Text can optionally be displayed to a user during media editing.

Claims (41)

1. A method comprising:

accessing, by a processor, a timestamp-to-text map that maps text corresponding to speech represented in an audio track for a video clip to a plurality of timestamps for the video clip;

determining, by the processor using a deep-learning neural network trained to output word occurrence probabilities, a word-based probability of at least one identified gap in the text corresponding to a grammatical pause in the speech represented in the audio track, the word-based probability based at least on part on word occurrence probabilities in relation to prior word;

combining, by the processor, the word-based probability and a time-based probability corresponding to the timestamp-to-text map to provide a comparative probability of the at least one identified gap in the text corresponding to the grammatical pause;

identifying, by the processor based on the comparative probability relative to an input-adjustable grammatical threshold that determines a number of displayed edit points, a potential edit point for the video clip that corresponds to the grammatical pause in the speech represented in the audio track of the video clip;

displaying, by the processor, on a presentation device, a marker indicating the potential edit point in the video clip;

receiving, by the processor, user input directed to an editing action at the marker; and

performing, by the processor, the editing action at the marker in the video clip in response to the user input.

2. The method of claim 1 further comprising:

kernel-additive modeling the audio track of the video clip to identify audio sources to isolate the speech from among the audio sources; and

producing the timestamp-to-text map based on the speech, wherein the text mapped by timestamp-to-text map includes both words and gaps in the speech.

3. The method of claim 1 further comprising causing the presentation device to display at least a portion of the text corresponding to the potential edit point.

4. The method of claim 3 further comprising causing the presentation device to display a visual attribute corresponding to each of at least two speakers in the video clip.

5. The method of claim 1 wherein the deep-learning neural network trained to output word occurrence probabilities comprises at least one of an N-gram language model, a neural network language model, or a recurrent neural network language model.

6. A non-transitory computer-readable medium storing program code executable by a processor to perform operations, the operations comprising:

accessing a timestamp-to-text map that maps text corresponding to speech represented in an audio track for a video clip to a plurality of timestamps for the video clip;

producing, using the timestamp-to-text map, based on an average-length of gaps between words in the speech, an indexed list containing high, time-based probability gaps in the speech;

determining, based on the indexed list and a word-based probability produced using a deep-learning neural network trained to output word occurrence probabilities, a comparative probability of at least one identified gap from the high, time-based probability gaps corresponding to a grammatical pause in the speech represented in the audio track, the word-based probability based at least on part on word occurrence probabilities in relation to prior word;

identifying, based on the comparative probability relative to an input-adjustable grammatical threshold configured to determine a number of displayed edit points, a potential edit point for the video clip that corresponds to the grammatical pause in the speech represented in the audio track of the video clip;

displaying, on a presentation device, a marker indicating the potential edit point in the video clip;

receiving user input directed to an editing action at the marker; and

performing the editing action at the marker in the video clip in response to the user input.

7. The non-transitory computer-readable medium of claim 6 wherein the operations further comprise:

kernel-additive modeling the audio track of the video clip to identify audio sources to isolate the speech from among the audio sources; and

producing the timestamp-to-text map based on the speech, wherein the text mapped by timestamp-to-text map includes both words and gaps in the speech.

8. The non-transitory computer-readable medium of claim 6 wherein the operations further comprise causing the presentation device to display at least a portion of the text corresponding to the potential edit point.

9. The non-transitory computer-readable medium of claim 8 wherein the operations further comprise causing the presentation device to display a visual attribute corresponding to a speaker.

10. The non-transitory computer-readable medium of claim 6 wherein the deep-learning neural network trained to output word occurrence probabilities comprises at least one of an N-gram language model, a neural network language model, or a recurrent neural network language model.

11. A method for producing a potential edit point for a video clip, the method comprising:

accessing an audio track for a video clip;

a step for determining, with a timestamp-to-text map and a deep-learning neural network trained to output word occurrence probabilities, a time-based probability and a word-based probability of at least one identified gap in text corresponding to a grammatical pause in speech represented in the audio track, the word-based probability based at least on part on word occurrence probabilities in relation to prior word;

combining the word-based probability and the time-based probability to provide a comparative probability of the at least one identified gap in the text corresponding to the grammatical pause;

identifying, based on the comparative probability relative to an input-adjustable grammatical threshold configured to determine a number of displayed edit points, a potential edit point for the video clip that corresponds to the grammatical pause in the speech represented in the audio track of the video clip;

displaying, on a presentation device, a marker indicating the potential edit point in the video clip;

receiving user input directed to an editing action at the marker; and

performing the editing action at the marker in the video clip in response to the user input.

12. The method of claim 11 further comprising kernel-additive modeling the audio track of the video clip to identify audio sources to isolate the speech from among the audio sources.

13. The method of claim 11 further comprising producing an indexed list containing high, time-based probability gaps in the speech based on an average-length of gaps between words in the speech as determined at least in part from the timestamp-to-text map.

14. The method of claim 11 further comprising a step for causing the presentation device to display at least a portion of the text corresponding to the potential edit point.

15. The method of claim 14 further comprising a step for causing the presentation device to display a visual attribute corresponding to each of at least two speakers in the video clip.

16. The method of claim 11 wherein the deep-learning neural network trained to output word occurrence probabilities comprises at least one of an N-gram language model, a neural network language model, or a recurrent neural network language model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: JINDAL, AMOL; JAIN, SOMYA; BEDI, AJAY
To: ADOBE INC.
Reel/Frame 053649/0130 →
Continuity (1)
Related Publication 20220068258A1 · Mar 3, 2022