IP Library Granted Patent US 12,308,005
Granted Patent B2
US 12,308,005 · App. 18/051,391 · Granted May 20, 2025

Audio-visual effects system for augmentation of captured performance based on content thereof

Inventors: David Steinwedel (San Francisco, CA); Perry R. Cook (Jacksonville, OR); Paul T. Chi (San Jose, CA); Wei Zhou (San Francisco, CA); Jon Moldover (San Francisco, CA); Anton Holmberg (San Francisco, CA); Jingxi Li (San Francisco, CA)
Assignee: SMULE, INC.
G10H1/368G06T11/00G10H1/366G11B27/02G10H2210/331G10H2220/005G10H2220/011G10H2220/355G10H2230/015G10H2240/175G10H2240/251
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,005
App. No.
18/051,391
Granted
May 20, 2025
Kind
B2
Abstract

Visual effects schedules are applied to audiovisual performances with differing visual effects applied in correspondence with differing elements of musical structure. Segmentation techniques applied to one or more audio tracks (e.g., vocal or backing tracks) are used to compute some of the components of the musical structure. In some cases, applied visual effects schedules are mood-denominated and may be selected by a performer as a component of his or her visual expression or determined from an audiovisual performance using machine learning techniques.

Claims (68)

1. A system including an audiovisual processing pipeline comprising:

a video effects planner that, based on a selected video style recipe, maps computationally extracted audio features corresponding to a machine-readable encoding of an audiovisual performance to particular sets of visual effects; and

a video effects renderer that applies differing visual effects of the particular sets of visual effects to the audiovisual performance encoding in correspondence with differing audio features of the computationally extracted audio features.

2. The system of claim 1 , further comprising:

a segmentation engine that computationally extracts the audio features from one or more tracks corresponding to the machine-readable encoding of the audiovisual performance.

3. The system of claim 2 ,

wherein at least one of the tracks includes captured vocal audio; and

wherein the segmentation engine operates on one or more of:

the captured vocal audio track;

a backing track against which the vocal audio track was captured; and

a MIDI file corresponding to the captured vocal audio track and/or the backing track,

to extract the audio features including at least musical section boundaries coded for temporal alignment with the captured vocal audio track.

4. The system of claim 3 ,

wherein the applied visual effects encode differing visual effects for differing musical structure elements of the audiovisual performance encoding and provide visual effect transitions in temporal alignment with at least some of the coded musical section boundaries.

5. The system of claim 4 , wherein the visual effects encoded for respective musical structure elements of the audiovisual performance encode one or more of:

a particle-based effect or lens flare;

transitions between, or layouts of, distinct source videos;

animations or motion of a frame within a source video

vector graphics or images of patterns or textures; and

color, saturation or contrast.

6. The system of claim 1 ,

wherein the selected video style recipe is selected from amongst a plurality of video style recipes based on a computationally-determined mood for the audiovisual performance.

7. The system of claim 1 ,

wherein the selected video style recipe is selected based on a user interface selection by the vocal audio performer prior to, or coincident with, capture of the vocal audio.

8. The system of claim 1 ,

wherein the selected video style recipe is selected based on a user interface selection by the vocal audio performer after an initial, post vocal audio capture, rendering of the audiovisual performance.

9. The system of claim 1 , further comprising:

a communications interface configured to stream the audiovisual performance to an audience at one or more remote client devices.

10. The system of claim 9 ,

wherein the streamed audiovisual performance is mixed with an encoding of a backing track against which the vocal audio was captured.

11. The system of claim 9 ,

wherein the streamed audiovisual performance includes the applied set of visual effects.

12. The system of claim 9 ,

wherein the streamed audiovisual performance is supplied with an identification of the applied set of visual effects for video effect rendering at one or more of the remote client devices.

13. The system of claim 1 ,

wherein at least some of the video style recipes map to mood-denominated sets of visual effects, and

wherein for a particular mood-denominated set of visual effects applied by the renderer, mood values are parameterized as a two-dimensional quantity, wherein a first dimension of the mood parameterization codes an emotion and wherein second dimension of the mood parameterization codes intensity.

14. The system of claim 13 ,

wherein the intensity dimension of the mood parameterization based on one or more of (i) a time-varying audio signal strength or vocal energy density measure computationally determined from a vocal audio track and (ii) beats, tempo, signal strength or energy density of a backing audio track against which the vocal audio track was captured.

15. The system of claim 1 ,

wherein operation of the segmentation engine is based at least in part on a computational determination of vocal intensity of a vocal audio track with at least some segmentation boundaries constrained to temporally align with beats or tempo computationally extracted from a corresponding audio backing track against which the vocal audio track was captured.

16. The system of claim 1 ,

embodied, at least in part, as a content server or service platform to which geographically-distributed, network-connected, vocal capture devices are communicatively coupled.

17. The system of claim 1 ,

embodied, at least in part, as a network-connected, vocal capture device communicatively coupled to stream the audiovisual performance to one or more client devices.

18. A method comprising: processing an audiovisual performance in an audiovisual processing pipeline to augment a machine-readable encoding of the audiovisual performance with temporally aligned visual effects,

the processing including executing video effects planning code that, based on a selected video style recipe, maps computationally extracted audio features corresponding to a machine-readable encoding of the audiovisual performance to particular sets of visual effects and executing video effects rendering code that applies differing visual effects of the particular sets of visual effects to the audiovisual performance encoding in correspondence with differing audio features of the computationally extracted audio features.

19. The method of claim 18 , further comprising:

segmenting the machine-readable encoding of the audiovisual performance to computationally extract the corresponding audio features.

20. The method of claim 19 ,

wherein the machine-readable encoding of the audiovisual performance includes one or more tracks including at least one captured vocal audio track; and

wherein the segmenting operates on one or more of:

the captured vocal audio track;

a backing track against which the vocal audio track was captured; and

a MIDI file corresponding to the captured vocal audio track and/or the backing track,

to extract the audio features including at least musical section boundaries coded for temporal alignment with the captured vocal audio track.

21. The method of claim 18 ,

wherein the applied visual effects encode differing visual effects for differing musical structure elements of the audiovisual performance encoding and provide visual effect transitions in temporal alignment with at least some of the coded musical section boundaries.

22. The method of claim 18 , further comprising:

after an audiovisual rendering of the audiovisual performance with a first set of visual effects applied, selecting a second video style recipe from amongst a plurality of video style recipes, the video style recipe mapping to a second set of visual effects differing first set of visual effects; and

in the video effects rendering code of the audiovisual processing pipeline, applying the second set visual effects to at least a portion of the audiovisual performance encoding.

23. The method of claim 22 ,

wherein at least some of the plurality of video style recipes map to mood-denominated sets of visual effects, and

wherein the applied second set visual effects is a mood-denominated set of visual effects.

24. The method of claim 18 , further comprising:

capturing the audiovisual performance at a network-connected vocal capture device communicatively coupled to a content server or service platform from which the musical structure encoding is supplied.

25. The method of claim 18 ,

wherein the audiovisual performance capture is performed at the network-connected vocal capture device in accordance with a Karaoke-style operational mechanic in which lyrics are visually presented in correspondence with audible rendering of a backing track.

Assignments (4)
SECURITY INTEREST Recorded Dec 30, 2024
From: SMULE, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 069703/0571 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2024
From: STEINWEDEL, DAVID; COOK, PERRY R.; CHI, PAUL T.; ZHOU, WEI; MOLDOVER, JON; HOLMBERG, ANTON
To: SMULE, INC.
Reel/Frame 066491/0577 →
CONFIDENTIAL INFORMATION AGREEMENT Recorded Feb 19, 2024
From: LI, JINGXI
To: SMULE, INC.
Reel/Frame 066625/0521 →
SECURITY INTEREST Recorded Apr 19, 2023
From: SMULE, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 063376/0889 →
Continuity (8)
Continuation 16107351 · Aug 21, 2018
Continuation 15910646 · Mar 2, 2018
Continuation 15173041 · Jun 3, 2016
Continuation In Part 15944537 · Apr 3, 2018
Provisional Application 62548122 · Aug 21, 2017
Provisional Application 62170255 · Jun 3, 2015
Provisional Application 62480610 · Apr 3, 2017
Related Publication 20230335094A1 · Oct 19, 2023
References Cited (147)
US 4688464A · Gibson · 1987 [cited by applicant]
US 5231671A · Gibson · 1993 [cited by applicant]
US 5301259A · Gibson · 1994 [cited by applicant]
US 5477003A · Muraki · 1995 [cited by applicant]
US 5719346A · Yoshida · 1998 [cited by applicant]
US 5811708A · Matsumoto · 1998 [cited by applicant]
US 5889223A · Matsumoto · 1999 [cited by applicant]
US 5902950A · Kato · 1999 [cited by applicant]
US 5939654A · Anada · 1999 [cited by applicant]
US 5966687A · Ojard · 1999 [cited by applicant]
US 6121531A · Kato · 2000 [cited by applicant]
US 6307140B1 · Iwamoto · 2001 [cited by applicant]
US 6336092B1 · Gibson · 2002 [cited by applicant]
US 6353174B1 · Schmidt · 2002 [cited by applicant]
US 6369311B1 · Iwamoto · 2002 [cited by applicant]
US 6971882B1 · Kumar · 2005 [cited by applicant]
US 7003496B2 · Ishii · 2006 [cited by applicant]
US 7068596B1 · Mou · 2006 [cited by applicant]
US 7096080B2 · Asada · 2006 [cited by applicant]
US 7297858B2 · Paepcke · 2007 [cited by applicant]
US 7853342B2 · Redmann · 2010 [cited by applicant]
US 8682653B2 · Salazar · 2014 [cited by applicant]
US 8868411B2 · Cook · 2014 [cited by applicant]
US 8983829B2 · Cook · 2015 [cited by applicant]
US 8996364B2 · Cook · 2015 [cited by applicant]
US 9058797B2 · Salazar · 2015 [cited by applicant]
US 9147385B2 · Salazar · 2015 [cited by applicant]
US 9324330B2 · Chordia et al. · 2016 [cited by applicant]
US 9601127B2 · Yang · 2017 [cited by applicant]
US 9866731B2 · Godfrey et al. · 2018 [cited by applicant]
US 9911403B2 · Sung · 2018 [cited by applicant]
US 20020004191A1 · Tice · 2002 [cited by applicant]
US 20020032728A1 · Sako · 2002 [cited by applicant]
US 20020051119A1 · Sherman · 2002 [cited by applicant]
US 20020056117A1 · Hasegawa · 2002 [cited by applicant]
US 20020082731A1 · Pitman · 2002 [cited by applicant]
US 20020091847A1 · Curtin · 2002 [cited by applicant]
US 20020177994A1 · Chang · 2002 [cited by applicant]
US 20030014262A1 · Kim · 2003 [cited by applicant]
US 20030099347A1 · Ford · 2003 [cited by applicant]
US 20030100965A1 · Sitrick · 2003 [cited by applicant]
US 20030117531A1 · Rovner · 2003 [cited by applicant]
US 20030164924A1 · Sherman · 2003 [cited by applicant]
US 20040159215A1 · Tohgi · 2004 [cited by applicant]
US 20040263664A1 · Aratani · 2004 [cited by applicant]
US 20050120865A1 · Tada · 2005 [cited by applicant]
US 20050123887A1 · Joung · 2005 [cited by applicant]
US 20050252362A1 · McHale · 2005 [cited by applicant]
US 20060165240A1 · Bloom · 2006 [cited by applicant]
US 20060206582A1 · Finn · 2006 [cited by applicant]
US 20070065794A1 · Mangum · 2007 [cited by applicant]
US 20070140510A1 · Redmann · 2007 [cited by applicant]
US 20070150082A1 · Yang · 2007 [cited by applicant]
US 20070245881A1 · Egozy · 2007 [cited by applicant]
US 20070245882A1 · Odenwald · 2007 [cited by applicant]
US 20070250323A1 · Dimkovic · 2007 [cited by applicant]
US 20070260690A1 · Coleman · 2007 [cited by applicant]
US 20070287141A1 · Milner · 2007 [cited by applicant]
US 20070294374A1 · Tamori · 2007 [cited by applicant]
US 20080026690A1 · Foxenland · 2008 [cited by examiner]
US 20080033585A1 · Zopf · 2008 [cited by applicant]
US 20080105109A1 · Li · 2008 [cited by applicant]
US 20080156178A1 · Georges · 2008 [cited by applicant]
US 20080190271A1 · Taub · 2008 [cited by applicant]
US 20080267443A1 · Aarabi · 2008 [cited by examiner]
US 20080270541A1 · Keener · 2008 [cited by applicant]
US 20080312914A1 · Rajendran · 2008 [cited by applicant]
US 20090003659A1 · Forstall · 2009 [cited by applicant]
US 20090038467A1 · Brennan · 2009 [cited by applicant]
US 20090106429A1 · Siegal · 2009 [cited by applicant]
US 20090107320A1 · Willacy · 2009 [cited by applicant]
US 20090165634A1 · Mahowald · 2009 [cited by applicant]
US 20090191521A1 · Paul · 2009 [cited by examiner]
US 20100087240A1 · Egozy · 2010 [cited by applicant]
US 20100126331A1 · Golovkin · 2010 [cited by applicant]
US 20100142926A1 · Coleman · 2010 [cited by applicant]
US 20100157016A1 · Sylvain · 2010 [cited by applicant]
US 20100192753A1 · Gao · 2010 [cited by applicant]
US 20100203491A1 · Yoon · 2010 [cited by applicant]
US 20100326256A1 · Emmerson · 2010 [cited by applicant]
US 20100332222A1 · Bai · 2010 [cited by examiner]
US 20110126103A1 · Cohen · 2011 [cited by examiner]
US 20110144981A1 · Salazar · 2011 [cited by examiner]
US 20110144982A1 · Salazar · 2011 [cited by applicant]
US 20110144983A1 · Salazar et al. · 2011 [cited by applicant]
US 20110154197A1 · Hawthorne · 2011 [cited by applicant]
US 20110251841A1 · Cook · 2011 [cited by applicant]
US 20110251842A1 · Cook · 2011 [cited by applicant]
US 20130006625A1 · Gunatilake · 2013 [cited by applicant]
US 20130254231A1 · Decker · 2013 [cited by applicant]
US 20140007147A1 · Anderson · 2014 [cited by applicant]
US 20140229831A1 · Chordia et al. · 2014 [cited by applicant]
US 20140282748A1 · McNamee · 2014 [cited by applicant]
US 20140290465A1 · Salazar · 2014 [cited by applicant]
US 20150201161A1 · Lachapelle · 2015 [cited by applicant]
US 20150279427A1 · Godfrey et al. · 2015 [cited by applicant]
US 20160057316A1 · Godfrey · 2016 [cited by applicant]
US 20160358595A1 · Sung · 2016 [cited by applicant]
CN 102456340A · 2012 [cited by applicant]
CN 104580838A · 2015 [cited by applicant]
CN 108040497A · 2018 [cited by applicant]
EP 2018058A1 · 2009 [cited by applicant]
EP 1065651B1 · 2016 [cited by applicant]
GB 2554322A · 2018 [cited by applicant]
IN 102456340A · 2012 [cited by applicant]
JP 2006047754A · 2006 [cited by applicant]
JP 2006311079A · 2006 [cited by applicant]
JP 2010060627A · 2010 [cited by applicant]
JP 2016206575A · 2016 [cited by applicant]
KR 20070016901A · 2007 [cited by applicant]
KR 20140023665A · 2014 [cited by applicant]
KR 20150033757A · 2015 [cited by applicant]
KR 101605497 · 2016 [cited by applicant]
WO 2003030143A1 · 2003 [cited by applicant]
WO 2011075446A1 · 2011 [cited by applicant]
WO 2011130325A1 · 2011 [cited by applicant]
WO 2015103415A1 · 2015 [cited by applicant]
WO 2016070080A1 · 2016 [cited by applicant]
WO 2016196987A1 · 2016 [cited by applicant]
WO 2018187360A2 · 2018 [cited by applicant]
PCT International Search Report of International Search Authority for counterpart application, mailed Feb. 17, 2016, of PCT/US2015/058373 filed Oct. 30, 2015. [cited by applicant]
PCT International Search Report/Written Opinion of International Search Authority for Counterpart application, mailed Oct. 17, 2016 of PCT/US2016/035810. [cited by applicant]
Kuhn, William. “A Real-Time Pitch Recognition Algorithm for Music Applications.” Computer Music Journal, vol. 14, No. 3, Fall 1990, Massachusetts Institute of Technology, Print pp. 60-71. [cited by applicant]
Johnson. Joel. “Glee on iPhone More than Good-It—s Fabulous.” Apr. 15, 2010. Web. http://gizmodo.com/5518067/glee-on-iphone-more-than-goodits-fabulous. Accessed Jun. 28, 2011. p. 1-3. [cited by applicant]
Wortham, Jenna. “Unleash Your Inner Gleek on the iPad.” Bits, The New York Times. Apr. 15, 2010. Web. http://bits.blogs.nytimes.com/2010/04/15/unleash-your-inner-gleek-on-the-ipad/?pagemod. [cited by applicant]
Gerhard, David. “Pitch Extraction and Fundamental Frequency: History and Current Techniques.” Department of Computer Science, University of Regina, Saskatchewan, Cananda. Nov. 2003 Print. p. 1-22. [cited by applicant]
“Auto-Tune: Intonation Correcting Plug-In.” User—s Manual. Antares Audio Technologies. 2000. Print. p. 1-52. [cited by applicant]
Trueman, Daniel et al.“PLOrk: the Princton Laptop Orchestra, Year 1. ”Music Department, Princeton University. 2009 Print. 10 pages. [cited by applicant]
Conneally, Tim. “The Age of Egregious Auto-tuning: 1998-2009.” Tech Gear News-Betanews. Jun. 15, 2009. Web.www.betanews.com/artical/the-age-of-egregious-autotuning-19982009/1245090927 Accessed Dec. 10, 2009. [cited by applicant]
Baran, Tom. “Autotalent v0.2: Pop Music in a Can” Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology. May 22, 2011. Web. http//web.mit.edu/tbaran/www/autotalent.html Accesse… [cited by applicant]
Atal, Bishnu S. “The History of Linear Prediction.” IEEE Signal Processing Magazine. vol. 154, Mar. 2006 Print. p. 154-161. [cited by applicant]
Shaffer, H. and Ross, M. and Cohen, A. “AMDF Pitch Extractor.” 85th Meeting Acoustical Society of America. vol. 54:1, Apr. 13, 1973. Print. p. 340. [cited by applicant]
Kumparak , Greg. “Gleeks Rejoice Smule Packs Fox—s Glee Into A Fantastic IPhone Application” MobilCrunch. Apr. 15, 2010 Web. Accessed Jun. 28, 2011; http://www.mobilecrunch.com/2010/04/15gleeks-rejoice-smule-packs-foxs-… [cited by applicant]
Rabiner, Lawrence R. “On the Use of Autocorrelation Analysis for Pitch Detection.” IEEE Transactions on Acoustics, Speech, and Signal Processing. vol. Assp. -25:Feb. 1, 1977. Print p. 24-33. [cited by applicant]
Wang, Ge. “Designing Smule—s IPhone Ocarina.” Center for Computer Research in Music and Acoustics, Standford University. Jun. 2009. Print 5 pages. [cited by applicant]
Clark, Don; “MuseAmi Hopes to Take Music Automation to New Level.” The Wall Street Journal, Digits, Technology News and Insigts, Mar. 19, 2010 Web. Accessed Jul. 6, 2011 http://blogs.wsj.com/digits/2010/19/museami-hopes… [cited by applicant]
Ananthapadmanabha, Tirupattur V. et al. “Epoch Extraction from Linear Prediction Residual for Identification of Closed Glottis Interval.” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. ASSP-27:4. Au… [cited by applicant]
Cheng, M.J. “Some Comparisons Among Several Pitch Detection Algorithms.” Bell Laboratories. Murray Hill, NJ. 1976. p. 332-335. [cited by applicant]
International Search Report and Written Opinion mailed in International Application No. PCT/US2010/60135 on Feb. 8, 2011, 17 pages. [cited by applicant]
International Search Report mailed in International Application No. PCT/US2011/032185 on Aug. 17, 2011, 6 pages. [cited by applicant]
Johnson-Bristow, Robert. “A Detailed Analysis of a Time-Domain Formant Corrected Pitch Shifting Alogorithm” AES: An Audio Engineering Society Preprint. Oct. 1993. Print. 24 pages. [cited by applicant]
Lent, Keith. “An Efficient Method for Pitch Shifting Digitally Sampled Sounds.” Departments of Music and Electrical Engineering, University of Texas at Austin. Computer Music Journal, vol. 13:4, Winter 1989, Massachuset… [cited by applicant]
McGonegal, Carol A. et al. “A Semiautomatic Pitch Detector (SAPD).” Bell Laboratories. Murray Hill, NJ. May 19, 1975. Print. p. 570-574. [cited by applicant]
Ying, Goangshiuan S. et al. “A Probabilistic Approach to AMDF Pitch Detection.” School of Electrical and Computer Engineering, Purdue University. 1996. Web. http://purcell.ecn.purdue.edu/˜speechg . Accessed Jul. 5, 2011… [cited by applicant]
Movie Maker, “Windows Movie Maker: Transitions and Video Effects”, [online], published Jan. 2007. [cited by applicant]
PCT International Search Report/Written Opinion of International Search Authority for Counterpart application, mailed Sep. 28, 2018 of PCT/US2018/025937. [cited by applicant]
PCT International Search Report and Written Opinion for counterpart application mailed Jan. 18, 2019 for PCT/US2018/047325 filed Aug. 21, 2018. [cited by applicant]