IP Library Granted Patent US 12682931
Granted Patent B2
US 12682931 · App. 18/292,223 · Granted Jul 14, 2026

Generating audiovisual content based on video clips

Inventors: Nicholas James Clark (London, GB); Glen Murphy (Palo Alto, CA); Jason Briggs Cornwell (Lafayette, CA); Conor Patrick O'Sullivan (San Francisco, CA); Philip Loyd Burk (Larkspur, CA); Eunyoung Park (Sunnyvale, CA); Philip Francis Rowe (Portland, OR); Donald Peter Turner (Fetcham, GB); Karl David Möllerstedt (Stockholm, SE); Finn Åke Axel Ericson (Stockholm, SE); Svante Sten Johan Stadler (Stockholm, SE); Johan Philip Claesson (Stockholm, SE); Ola Fredrik Josefsson (Stockholm, SE)
Assignee: Google LLC
G11B27/031G06F3/0483G06F3/04847G10H1/368
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682931
App. No.
18/292,223
Filed
Jan 25, 2024
Granted
Jul 14, 2026
Kind
B2
Art Unit
2484
USPC
386/282
Abstract

A method includes capturing, by a content generation component of a computing device, initial content comprising video, and audio associated with the video; identifying one or more audio clips in the audio associated with the video based on one or more transient points in the audio; extracting, for each audio clip, a corresponding video clip from the video of the initial content; providing a control interface to enable a user-generated sequence of audio clips, wherein each audio clip in the sequence of audio clips is selected from the one or more identified audio clips; generating new audiovisual content comprising a sequence of video clips to correspond to the user-generated sequence of audio clips, wherein each video clip in the sequence of video clips is the extracted corresponding video clip for each audio clip in the user-generated sequence of audio clips; and providing, by the control interface, the new audiovisual content.

Claims (54)

1 . A computing device, comprising:

a graphical user interface configured to enable generation of audiovisual content;

one or more processors; and

data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions comprising:

capturing, by a content generation component of the computing device, initial content comprising co-captured video, and co-captured audio associated with the video;

identifying one or more audio clips in the audio associated with the video based on one or more transient points in the audio;

extracting, for each audio clip of the one or more identified audio clips, a corresponding video clip from the video of the initial content;

providing, via the graphical user interface, a control interface to enable a user-generated sequence of audio clips, wherein each audio clip in the sequence of audio clips is selected from the one or more identified audio clips, and wherein the control interface comprises a plurality of selectable tabs corresponding to a plurality of audio channels, wherein user selection of a tab of the plurality of selectable tabs enables user access to one or more channel interfaces to interact with one or more of an audio clip or a video clip in the audio channel corresponding to the user selected tab;

generating new audiovisual content comprising a sequence of video clips to correspond to the user-generated sequence of audio clips, wherein each video clip in the sequence of video clips is the extracted corresponding video clip for each audio clip in the user-generated sequence of audio clips; and

providing, by the control interface, the new audiovisual content.

2 . The computing device of claim 1 , wherein the one or more identified audio clips comprise a plurality of percussive sounds comprising an initial rhythm, and wherein the providing of the control interface further comprises:

generating a plurality of modified versions of the plurality of percussive sounds, wherein the plurality of modified versions is associated with a modified rhythm different from the initial rhythm; and

providing, via the control interface, the plurality of modified versions of the plurality of percussive sounds,

wherein the user-generated sequence of audio clips is based on the plurality of modified versions of the plurality of percussive sounds.

3 . The computing device of claim 1 , wherein the plurality of audio channels comprise audio clips corresponding to one or more of a melodic note, a percussive sound, a musical composition, an instrumental sound, a silence, or a vocal phrase.

4 . The computing device of claim 1 , wherein each audio channel of the plurality of audio channels is associated with a given audiovisual content different from the initial content.

5 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface with one or more icons corresponding to the one or more identified audio clips, and wherein the user-generated sequence of audio clips is based on user indication of selecting at least one icon of the one or more icons to generate the sequence.

6 . The computing device of claim 1 , wherein an audio clip of the one or more identified audio clips comprises a musical note, and wherein the providing of the control interface further comprises:

generating a plurality of repitched versions of the musical note,

wherein the one or more channel interfaces comprises an interface with one or more icons corresponding to the plurality of repitched versions of the musical note, and

wherein the user-generated sequence of audio clips is based on user indication of selecting at least one icon of the one or more icons to generate the sequence.

7 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a plurality of user-generated sequences, each sequence of the plurality of user-generated sequences corresponding to the plurality of audio channels, and further comprising a selectable option enabling a user to chain the one or more sequences to generate a new sequence.

8 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a plurality of user-generated sequences, each sequence of the plurality of user-generated sequences corresponding to the plurality of audio channels, and further comprising a selectable option enabling a user to mix the one or more sequences to generate a new audio track.

9 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a pair of coordinate axes, wherein a horizontal axis corresponds to a plurality of pitch adjustments for the user-generated sequence, and a vertical axis corresponds to a plurality of simultaneously adjustable audio filter adjustments for the user-generated sequence.

10 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a plurality of respective volume controls for the plurality of audio channels, wherein the plurality of respective volume controls enable a user to simultaneously control volume settings of each of the plurality of audio channels.

11 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a first tool to adjust a tempo, a second tool to adjust a swing, and a third tool to adjust a root musical note, for an audio clip in the sequence of audio clips.

12 . The computing device of claim 1 , wherein the one or more channel interfaces comprises an interface displaying a plurality of video edit icons, and wherein user selection of a video edit icon of the plurality of video edit icons enables application of a video edit feature to a video clip of the sequence of video clips.

13 . The computing device of claim 1 , the functions further comprising:

identifying one or more second audio clips in a second initial content based on one or more second transient points in the second initial content; and

enabling, via the control interface, a second user-generated sequence of second audio clips, wherein each second audio clip in the sequence of second audio clips is selected from the one or more identified second audio clips,

wherein the generating of the new audiovisual content comprises generating a second sequence of video clips to correspond to the user-generated sequence of audio clips and the user-generated sequence of second audio clips.

14 . The computing device of claim 1 , wherein the transient points in the initial content comprise one or more of a transient location, a pause, or a cut.

15 . The computing device of claim 1 , wherein the identifying of the one or more audio clips in the initial content comprises identifying, in a soundtrack of the initial content, one or more of a melodic note, a percussive sound, a musical composition, an instrumental sound, a change in audio intensity, a silence, or a vocal phrase.

16 . The computing device of claim 1 , wherein the identifying of the one or more audio clips in the initial content is performed by a trained machine learning model.

17 . The computing device of claim 1 , wherein the identifying of the one or more audio clips further comprises:

identifying, by a trained machine learning model, a classification for an audio clip of the one or more audio clips;

generating, based on the classification, a visual label associated with the audio clip; and

displaying, via the control interface, the visual label on a selectable icon corresponding to the audio clip.

18 . The computing device of claim 1 , further comprising:

providing user selectable virtual tabs to enable automatic upload of the new audiovisual content to a social networking site.

19 . A computer-implemented method comprising:

capturing, by a content generation component of the computing device, initial content comprising co-captured video, and co-captured audio associated with the video;

identifying one or more audio clips in the audio associated with the video based on one or more transient points in the audio;

extracting, for each audio clip of the one or more identified audio clips, a corresponding video clip from the video of the initial content;

providing, via a graphical user interface of the computing device, a control interface to enable a user-generated sequence of audio clips, wherein each audio clip in the sequence of audio clips is selected from the one or more identified audio clips, and wherein the control interface comprises a plurality of selectable tabs corresponding to a plurality of audio channels, wherein user selection of a tab of the plurality of selectable tabs enables user access to one or more channel interfaces to interact with one or more of an audio clip or a video clip in the audio channel corresponding to the user selected tab;

generating new audiovisual content comprising a sequence of video clips to correspond to the user-generated sequence of audio clips, wherein each video clip in the sequence of video clips is the extracted corresponding video clip for each audio clip in the user-generated sequence of audio clips; and

providing, by the control interface, the new audiovisual content.

20 . An article of manufacture including a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by one or more processors of a computing device, cause the computing device to carry out operations comprising:

capturing, by a content generation component of the computing device, initial content comprising co-captured video, and co-captured audio associated with the video;

identifying one or more audio clips in the audio associated with the video based on one or more transient points in the audio;

extracting, for each audio clip of the one or more identified audio clips, a corresponding video clip from the video of the initial content;

providing, via a graphical user interface of the computing device, a control interface to enable a user-generated sequence of audio clips, wherein each audio clip in the sequence of audio clips is selected from the one or more identified audio clips, and wherein the control interface comprises a plurality of selectable tabs corresponding to a plurality of audio channels, wherein user selection of a tab of the plurality of selectable tabs enables user access to one or more channel interfaces to interact with one or more of an audio clip or a video clip in the audio channel corresponding to the user selected tab;

generating new audiovisual content comprising a sequence of video clips to correspond to the user-generated sequence of audio clips, wherein each video clip in the sequence of video clips is the extracted corresponding video clip for each audio clip in the user-generated sequence of audio clips; and

providing, by the control interface, the new audiovisual content.