IP Library Granted Patent US 11,264,058
Granted Patent B2
US 11,264,058 · App. 16/834,775 · Granted Mar 1, 2022

Audiovisual capture and sharing framework with coordinated, user-selectable audio and video effects filters

Inventors: Parag P. Chordia (Los Altos, CA); Perry R. Cook (Jacksonville, OR); Mark T. Godfrey (Atlanta, GA); Prerna Gupta (Los Altos Hills, CA); Nicholas M. Kruge (San Francisco, CA); Randal J. Leistikow (Palo Alto, CA); Alexander M. D. Rae (Atlanta, GA); Ian S. Simon (San Francisco, CA)
Assignee: Smule, Inc.
G11B27/031G06F3/0482G06F3/04842G10H1/0025G10H1/383G10L21/003G10L21/055G10H2210/576G10L21/013H04N21/41407H04N21/854
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,264,058
App. No.
16/834,775
Granted
Mar 1, 2022
Kind
B2
Abstract

Coordinated audio and video filter pairs are applied to enhance artistic and emotional content of audiovisual performances. Such filter pairs, when applied in audio and video processing pipelines of an audiovisual application hosted on a portable computing device (such as a mobile phone or media player, a computing pad or tablet, a game controller or a personal digital assistant or book reader) can allow user selection of effects that enhance both audio and video coordinated therewith. Coordinated audio and video are captured, filtered and rendered at the portable computing device using camera and microphone interfaces, using digital signal processing software executable on a processor and using storage, speaker and display devices of, or interoperable with, the device. By providing audiovisual capture and personalization on an intimate handheld device, social interactions and postings of a type made popular by modern social networking platforms can now be extended to audiovisual content.

Claims (65)

1. An audiovisual processing method comprising:

processing corresponding audio and video streams in coordinated audio and video pipelines, wherein the processing the audio and video streams includes:

in the audio pipeline, segmenting the audio stream into a plurality of segments and mapping the plurality of segments to respective subphrase portions of a phrase template for a target song; and

in the video pipeline, segmenting the video stream and mapping segments thereof in correspondence with the audio segmentation and mapping, and applying a video effect filter to the mapped video segments;

automatically generating, in the audio pipeline, a musical accompaniment for vocals in the captured audio stream as specified by an audio filter in the audio pipeline,

wherein the audio and video effect filters are coordinated;

audiovisually rendering the processed audio and video streams, wherein the audiovisual rendering includes the automatically generated musical accompaniment; and

storing, transmitting, or posting the rendered audiovisual content.

2. The method of claim 1 , wherein the video effect filter is applied to the mapped video segments based on audio features extracted in the audio pipeline.

3. The method of claim 2 , wherein the audio features include

melody pitches, and wherein the musical accompaniment is generated based on a

selection of chords that are harmonies of the melody pitches.

4. The method of claim 1 ,

wherein the captured audio stream includes vocals temporally synchronized with the video stream, and

wherein the segments are delimited in the audio pipeline based on onsets detected in the vocals.

5. The method of claim 1 , further comprising:

in the audio pipeline, temporally aligning successive ones of the segments with respective pulses of a rhythmic skeleton for the target song, and temporally adjusting at least some of the temporally aligned segments,

in the video pipeline, temporally aligning and adjusting respective segments thereof in correspondence with the audio segmentation aligning and adjusting.

6. The method of claim 1 , wherein the processing the audio and video streams further includes:

using, in the audio pipeline, one or more temporally localizable features extracted in the video pipeline.

7. The method of claim 1 , wherein the processing corresponding audio and video streams in the coordinated audio and video pipelines further includes:

applying artistically consistent effects to the audio and video streams.

8. A non-transitory machine-readable medium comprising a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method comprising:

processing corresponding audio and video streams in coordinated audio and video pipelines, wherein the processing the audio and video streams includes:

in the audio pipeline, segmenting the audio stream into a plurality of segments and mapping the plurality of segments to respective subphrase portions of a phrase template for a target song; and

in the video pipeline, segmenting the video stream and mapping segments thereof in correspondence with the audio segmentation and mapping, and applying a video effect filter to the mapped video segments;

automatically generating, in the audio pipeline, a musical accompaniment for vocals in the captured audio stream as specified by an audio filter in the audio pipeline,

wherein the audio and video effect filters are coordinated;

audiovisually rendering the processed audio and video streams, wherein the audiovisual rendering includes the automatically generated musical accompaniment; and

storing, transmitting, or posting the rendered audiovisual content.

9. The non-transitory machine-readable medium of claim 8 , wherein the video effect filter is applied to the mapped video segments based on audio features extracted in the audio pipeline.

10. The non-transitory machine-readable medium of claim 9 ,

wherein the audio features include melody pitches, and wherein the musical accompaniment is generated based on a selection of chords that are harmonies of the melody pitches.

11. The non-transitory machine-readable medium of claim 8 ,

wherein the captured audio stream includes vocals temporally synchronized with the video stream, and

wherein the segments are delimited in the audio pipeline based on onsets detected in the vocals.

12. The non-transitory machine-readable medium of claim 8 , wherein the method further comprises:

in the audio pipeline, temporally aligning successive ones of the segments with respective pulses of a rhythmic skeleton for the target song, and temporally adjusting at least some of the temporally aligned segments,

in the video pipeline, temporally aligning and adjusting respective segments thereof in correspondence with the audio segmentation aligning and adjusting.

13. The non-transitory machine-readable medium of claim 8 , wherein the processing the audio and video streams further includes:

using, in the audio pipeline, one or more temporally localizable features extracted in the video pipeline.

14. The non-transitory machine-readable medium of claim 8 , wherein the processing corresponding audio and video streams in coordinated audio and video pipelines further includes:

applying artistically consistent effects to the audio and video streams.

15. A system, comprising:

a non-transitory memory; and

one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform a method comprising:

processing corresponding audio and video streams in coordinated audio and video pipelines, wherein the processing the audio and video streams includes:

in the audio pipeline, segmenting the audio stream into a plurality of segments and mapping the plurality of segments to respective subphrase portions of a phrase template for a target song; and

in the video pipeline, segmenting the video stream and mapping segments thereof in correspondence with the audio segmentation and mapping, and applying a video effect filter to the mapped video segments;

automatically generating, in the audio pipeline, a musical accompaniment for vocals in the captured audio stream as specified by an audio filter in the audio pipeline,

wherein the audio and video effect filters are coordinated;

audiovisually rendering the processed audio and video streams, wherein the audiovisual rendering includes the automatically generated musical accompaniment; and

storing, transmitting, or posting the rendered audiovisual content.

16. The system of claim 15 , wherein the video effect filter is applied to the mapped video segments based on audio features extracted in the audio pipeline.

17. The system of claim 16 , wherein the audio features include

melody pitches, and wherein the musical accompaniment is generated based on a

selection of chords that are harmonies of the melody pitches.

18. The system of claim 15 ,

wherein the captured audio stream includes vocals temporally synchronized with the video stream, and

wherein the segments are delimited in the audio pipeline based on onsets detected in the vocals.

19. The system of claim 15 , wherein the method further comprises:

in the audio pipeline, temporally aligning successive ones of the segments with respective pulses of a rhythmic skeleton for the target song, and temporally adjusting at least some of the temporally aligned segments,

in the video pipeline, temporally aligning and adjusting respective segments thereof in correspondence with the audio segmentation aligning and adjusting.

20. The system of claim 15 , wherein the processing the audio and video streams further includes:

using, in the audio pipeline, one or more temporally localizable features extracted in the video pipeline.

Assignments (3)
SECURITY INTEREST Recorded Dec 30, 2024
From: SMULE, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 069703/0571 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: CHORDIA, PARAG P.; COOK, PERRY R.; GODFREY, MARK T.; GUPTA, PRERNA; KRUGE, NICHOLAS M.; LEISTIKOW, RANDAL J.; RAE, ALEXANDER M.D.; SIMON, IAN S.
To: SMULE, INC.
Reel/Frame 058724/0591 →
SECURITY INTEREST Recorded Apr 15, 2021
From: SMULE, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 055937/0207 →
Continuity (5)
Continuation 15284229 · Oct 3, 2016
Continuation 14104618 · Dec 12, 2013
Continuation In Part 13853759 · Mar 29, 2013
Provisional Application 61736503 · Dec 12, 2012
Related Publication 20200294550A1 · Sep 17, 2020