IP Library Granted Patent US 11,394,855
Granted Patent B2
US 11,394,855 · App. 16/814,132 · Granted Jul 19, 2022

Coordinating and mixing audiovisual content captured from geographically distributed performers

Inventors: Mark T. Godfrey (San Francisco, CA); Perry R. Cook (Jacksonville, OR)
Assignee: Smule, Inc.
H04N5/04G10H1/366G10L13/0335G10L21/013G10H2210/066G10H2210/331G10H2240/251Y10S84/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,394,855
App. No.
16/814,132
Granted
Jul 19, 2022
Kind
B2
Abstract

Audiovisual performances, including vocal music, are captured and coordinated with those of other users in ways that create compelling user experiences. In some cases, the vocal performances of individual users are captured (together with performance synchronized video) on mobile devices, television-type display and/or set-top box equipment in the context of karaoke-style presentations of lyrics in correspondence with audible renderings of a backing track. Contributions of multiple vocalists are coordinated and mixed in a manner that selects for visually prominent presentation performance synchronized video of one or more of the contributors. Prominence of particular performance synchronized video may be based, at least in part, on computationally-defined audio features extracted from (or computed over) captured vocal audio. Over the course of a coordinated audiovisual performance timeline, these computationally-defined audio features are selective for performance synchronized video of one or more of the contributing vocalists.

Claims (39)

1. A method of preparing a coordinated audiovisual performance from distributed performer contributions, the method comprising:

receiving via a communication network, a first audiovisual encoding of a first performer, the first audiovisual encoding comprising at least first performer audio captured at the first remote device and first performer performance synchronized video captured by a camera of the first remote device;

mixing the first performer audio with a first performer-selected backing track, wherein the mixing results in a first mixed audiovisual performance;

receiving via the communication network a selection of the first mixed audiovisual performance;

supplying via the communication network to a second remote device the first mixed audiovisual performance;

receiving via the communication network, a second audiovisual encoding of a second performer, the second audiovisual encoding comprising at least second performer audio captured at the second remote device against a local audio rendering, at the second remote device, of the first audiovisual encoding of the first performer and second performer performance synchronized video captured by a camera of the second remote device; and

generating a combined audiovisual performance mix of the first mixed audiovisual performance and the second audiovisual encoding of the second performer.

2. The method of claim 1 , further comprising:

determining, from the first performer audio, a first time-varying, computationally-defined audio feature;

determining, from the second performer audio, a second time-varying, computationally-defined audio feature; and

based on comparison of the first time-varying computationally-defined audio feature and the second time-varying computationally-defined audio feature dynamically varying relative visual prominence within the combined audiovisual performance mix of the first performer performance synchronized video and the second performer performance synchronized video.

3. The method of claim 2 , wherein the first time-varying computationally-defined audio feature includes an audio power measure, and wherein the second time-varying computationally-defined audio feature includes an audio power measure.

4. The method of claim 2 , wherein the combined audiovisual performance mix includes visual presentation of both first and second performer performance synchronized video with differing visual prominence for at least some values of the time-varying computationally-defined audio feature.

5. The method of claim 4 , wherein the combined audiovisual performance mix includes visual presentation of first performer video or second performer performance synchronized video, but not both, for at least some values of the computationally-defined audio feature.

6. The method of claim 2 , wherein the dynamic varying of relative visual prominence includes transitioning between prominent visual presentation of first performer performance synchronized video and prominent visual presentation of second performer performance synchronized video.

7. The method of claim 6 , wherein the transitioning includes switching, wiping, or crossfading of respective performer performance synchronized video.

8. The method of claim 6 , wherein the transitioning is subject to duration filtering or a hysteresis function.

9. The method of claim 1 , wherein the combined audiovisual performance mix includes visual presentation of both first and second performer performance synchronized video with equal levels of visual prominence.

10. The method of claim 1 , further comprising: inviting via electronic message or social network posting at least the second performer to provide the second audiovisual encoding.

11. The method of claim 1 , further comprising:

supplying the first and second remote devices with corresponding, but differing versions of the combined audiovisual performance mix,

wherein the combined audiovisual performance mix supplied to the first remote device features the first performer performance synchronized video more prominently than the second performer performance synchronized video, and

wherein the combined audiovisual performance mix supplied to the second remote device features the second performer performance synchronized video more prominently than the first performer performance synchronized video.

12. The method of claim 1 , further comprising:

pitch correcting at least one of the received first and second performer audio in accord with a vocal score that encodes (i) a sequence of notes for a vocal melody and (ii) at least a first set of harmony notes for at least some portions of the vocal melody.

13. The method of claim 1 , further comprising distributing the first mixed audiovisual performance to social media contacts of the first performer.

14. The method of claim 1 , wherein the combined audiovisual performance mix renders the first performer performance synchronized video on a left side of the combined audiovisual performance and renders the second performer performance synchronized video on a right side of the combined audiovisual performance.

15. The method of claim 1 , wherein the combined audiovisual performance mix includes visual ornamentation on at least one of the first performer performance synchronized video and second performer performance synchronized video.

16. The method of claim 1 , further comprising supplying via the communication network to the first and second remote devices the combined audiovisual performance mix.

17. The method of claim 1 , further comprising displaying, on a display of the second remote device, lyrics in correspondence with the local audio rendering of the first audiovisual encoding of the first performer.

18. A service platform, comprising:

one or more computing devices; and

machine-readable code embodied in a non-transitory medium and executable on at least one of the one or more computing devices to receive via a communication network, a first audiovisual encoding of a first performer, the first audiovisual encoding comprising at least first performer audio captured at the first remote device and first performer performance synchronized video captured by a camera of the first remote device;

the machine-readable code further executable to mix the first performer audio with a first performer-selected backing track, wherein the mixing results in a first mixed audiovisual performance;

the machine-readable code further executable to receive via the communication network a selection of the first mixed audiovisual performance;

the machine-readable code further executable to supply via the communication network to a second remote device the first mixed audiovisual performance;

the machine-readable code further executable to receive via the communication network, a second audiovisual encoding of a second performer, the second audiovisual encoding comprising at least second performer audio captured at the second remote device against a local audio rendering, at the second remote device, of the first audiovisual encoding of the first performer and second performer performance synchronized video captured by a camera of the second remote device;

the machine-readable code further executable to generate a combined audiovisual performance mix of the first mixed audiovisual performance and the second audiovisual encoding of the second performer; and

the machine-readable code further executable to supply via the communication network to the first and second remote devices the combined audiovisual performance mix.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: COOK, PERRY R.
To: SMULE, INC.
Reel/Frame 059077/0598 →
MERGER AND CHANGE OF NAME Recorded Feb 23, 2022
From: KHUSH, INC.; SMULE, INC.
To: SMULE, INC.
Reel/Frame 059078/0689 →
CONFIDENTIAL INFORMATION AGREEMENT Recorded Feb 23, 2022
From: GODFREY, MARK
To: KHUSH INC., A SUBSIDIARY OF SMULE, INC.
Reel/Frame 059239/0709 →
SECURITY INTEREST Recorded Apr 15, 2021
From: SMULE, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 055937/0207 →
Continuity (6)
Continuation 15864819 · Jan 8, 2018
Continuation 14928727 · Oct 30, 2015
Continuation In Part 14656344 · Mar 12, 2015
Division 13085414 · Apr 12, 2011
Provisional Application 62072558 · Oct 30, 2014
Related Publication 20210037166A1 · Feb 4, 2021