IP Library › Granted Patent US 12,640,127
Granted Patent B2
US 12,640,127 · App. 17/556,178 · Granted May 26, 2026

Interactive movement audio engine

Inventors: Bochen Li (Los Angeles, CA); Daiyu Zhang (Los Angeles, CA); Shawn Chan Zhen Yi (Los Angeles, CA); Jitong Chen (Los Angeles, CA)
Assignee: LEMON INC.
G10H1/0008G06V40/174G06V40/20G10H2210/105G10H2210/325G10H2210/571G10H2220/106G10H2220/201G10H2220/455G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,127
App. No.
17/556,178
Filed
Dec 20, 2021
Granted
May 26, 2026
Kind
B2
Art Unit
2837
USPC
84/609
Abstract

A method for generating an audio output is described. Image inputs of interactive movements by a user captured by an image sensor are received. The interactive movements are mapped to a sequence of audio element identifiers. The sequence of audio element identifiers are processed to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers. An audio output that represents the musical sequence is generated.

Claims (48)

1 . A method for generating an audio output, the method comprising:

receiving image inputs of interactive movements by a user captured by an image sensor;

displaying to the user an output image that comprises a graphical user interface overlaid on the image inputs, wherein the interactive movements comprise a plurality of user elements of the user that overlap with a corresponding plurality of icons the graphical user interface;

mapping each of the plurality of user elements of the interactive movements so that the user element overlaps with the corresponding icon of the plurality of icons of the graphical user interface to a sequence of audio element identifiers;

processing the sequence of audio element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers, wherein processing the audio element identifiers comprises:

receiving one or more selectable music theory rules that enforce corresponding elements of music theory;

modifying at least one audio element identifier of the sequence of audio element identifiers that violates a music theory rule when the violated music theory rule is at least one of the one or more received selectable music theory rules; and

generating the musical sequence based on the modified audio element identifier; and

generating an audio output that represents the musical sequence.

2 . The method of claim 1 , wherein modifying the at least one audio element identifier comprises changing a pitch associated with the audio element identifier.

3 . The method of claim 2 , wherein changing the pitch comprises matching a chord progression that satisfies the music theory rule.

4 . The method of claim 1 , wherein modifying the at least one audio element identifier comprises omitting the at least one audio element identifier when generating the musical sequence.

5 . The method of claim 1 , wherein modifying the at least one audio element identifier comprises changing a duration of the at least one audio element identifier.

6 . The method of claim 1 , wherein mapping the interactive movements comprises:

selecting a set of predetermined musical instruments from a plurality of instrument sets;

mapping the interactive movements to instruments within the selected set of predetermined musical instruments.

7 . The method of claim 6 , the method further comprising generating the plurality of instrument sets using a neural network engine that identifies sets of predetermined musical instruments from music samples.

8 . The method of claim 1 , wherein the plurality of user elements of the user are fingers, hands, arms, feet, and/or legs of the user.

9 . The method of claim 1 , wherein the plurality of icons corresponding to a plurality of predetermined audio element identifiers.

10 . The method of claim 9 , wherein the plurality of predetermined audio element identifiers includes single-element identifiers and multi-element identifiers.

11 . The method of claim 1 , wherein the interactive movements are facial expression elements performed by the user.

12 . The method of claim 1 , wherein the interactive movements are gestures performed by the user.

13 . A system for generating an audio output, the system comprising:

one or more hardware processors configured by machine-readable instructions to:

receive image inputs of interactive movements by a user captured by an image sensor;

display to the user an output image that comprises a graphical user interface overlaid on the image inputs, wherein the interactive movements comprise a plurality of user elements of the user that overlap with a corresponding plurality of icons the graphical user interface;

map each of the plurality of user elements of the interactive movements so that the user element overlaps with the corresponding icon of the plurality of icons of the graphical user interface to a sequence of audio element identifiers;

process the sequence of audio element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers, wherein processing the audio element identifiers comprises:

receive one or more selectable music theory rules that enforce corresponding elements of music theory;

modify at least one audio element identifier of the sequence of audio element identifiers that violates a music theory rule when the violated music theory rule is at least one of the one or more received selectable music theory rules; and

generate the musical sequence based on the modified audio element identifier; and

generate an audio output that represents the musical sequence.

14 . The system of claim 13 , wherein the one or more hardware processors are further configured by machine-readable instructions to:

change a pitch associated with the audio element identifier.

15 . The system of claim 13 , wherein the plurality of icons corresponding to a plurality of predetermined audio element identifiers.

16 . The system of claim 13 , wherein modifying the at least one audio element identifier comprises changing a pitch associated with the audio element identifier.

17 . The system of claim 16 , wherein changing the pitch comprises matching a chord progression that satisfies the music theory rule.

18 . The system of claim 13 , wherein modifying the at least one audio element identifier comprises omitting the at least one audio element identifier when generating the musical sequence.

19 . A non-transient computer-readable storage medium comprising instructions being executable by one or more processors, that when executed by the one or more processors, cause the one or more processors to:

receive image inputs of interactive movements by a user captured by an image sensor;

display to the user an output image that comprises a graphical user interface overlaid on the image inputs, wherein the interactive movements comprise a plurality of user elements of the user that overlap with a corresponding plurality of icons the graphical user interface;

map each of the plurality of user elements of the interactive movements so that the user element overlaps with the corresponding icon of the plurality of icons of the graphical user interface to a sequence of audio element identifiers;

process the sequence of audio element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers, wherein processing the audio element identifiers comprises:

receive one or more selectable music theory rules that enforce corresponding elements of music theory;

modify at least one audio element identifier of the sequence of audio element identifiers that violates a music theory rule when the violated music theory rule is at least one of the one or more received selectable music theory rules; and

generate the musical sequence based on the modified audio element identifier; and

generate an audio output that represents the musical sequence.

20 . The computer-readable storage medium of claim 19 , wherein the instructions are executable by the one or more processors to cause the one or more processors to: change a pitch associated with the audio element identifier.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2023
From: LI, BOCHEN; ZHANG, DAIYU; YI, SHAWN CHAN ZHEN; CHEN, JITONG
To: BYTEDANCE INC.
Reel/Frame 062738/0767 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2022
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 061210/0610 →
Continuity (1)
Related Publication 20230197040A1 · Jun 22, 2023
References Cited (15)
US 11551652B1 · Pajjuri · 2023 [cited by examiner]
US 20190147229A1 · Zatepyakin et al. · 2019 [cited by applicant]
US 20190295323A1 · Gutierrez · 2019 [cited by examiner]
US 20220164204A1 · Kwatra · 2022 [cited by examiner]
US 20220335974A1 · Butera · 2022 [cited by examiner]
US 20220375362A1 · Canberk · 2022 [cited by examiner]
CN 109413351A · 2019 [cited by applicant]
JP 2002311951A · 2002 [cited by applicant]
International Search Report mailed Jul. 7, 2023 in International Application No. PCT/SG2022/050853. [cited by applicant]
Prasad J. S. et al., Gesture Based Music Generation. Second International Conference on Emerging Trends in Engineering and Technology, ICETET-09, Jan. 30, 2010, pp. 209-214 [Retrieved on Jun. 30, 2023] <DOI: 10.1109/ICE… [cited by applicant]
IP H. H. S. et al., Cyber Composer: Hand Gesture-Driven Intelligent Music Composition and Generation. Proceedings of the 11th International Multimedia Modelling Conference (MMM'05), Feb. 14, 2005, pp. 1-7 [Retrieved on … [cited by applicant]
Beyer et al., “Music Interfaces for Novice Users: Composing Music on a Public Display with Hand Gestures”, Proceedings of the International Conference on New Interfaces for Musical Expression, Jun. 1, 2011, pp. 1-4. [cited by applicant]
Bresson et al., “From Motion to Musical Gesture: Experiments with Machine Learning in Computer-Aided Composition”, CNRS, 2018, pp. 1-5. [cited by applicant]
European Search Report for EP Patent Application No. 229121124, Issued on Feb. 7, 2025, 11 pages. [cited by applicant]
Office action received from Japanese patent application No. 2024-537405 mailed on Apr. 15, 2025, 5 pages. [cited by applicant]