IP Library › Granted Patent US 12,119,026
Granted Patent B2
US 12,119,026 · App. 17/722,202 · Granted Oct 15, 2024

Multimedia music creation using visual input

Inventor: Michael V. Butera (Nashville, TN)
Assignee: Artiphon, Inc.
G11B27/031G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,119,026
App. No.
17/722,202
Granted
Oct 15, 2024
Kind
B2
Abstract

A system for creating music using visual input. The system detects events and metrics (e.g., objects, gestures, etc.) in user input (e.g., video, audio, music data, touch, motion, etc.) and generates music and visual effects that are synchronized with the detected events and correspond to the detected metrics. To generate the music, the system selects parts from a library of stored music data and assigns each part to the detected events and metrics (e.g., using heuristics to match musical attributes to visual attributes in the user input). To generate the visual effects, the system applies rules (e.g., that map musical attributes to visual attributes) to translate the generated music data to visual effects. Because the visual effects are generated using music data that is generated using the detected events/metrics, both the generated music and the visual effects are synchronized with—and correspond to—the user input.

Claims (65)

1. A computer-implemented method for creating music using visual input, the method comprising:

storing a library of stored music data, the stored music data comprising a plurality of parts, each part comprising musical notes;

receiving, from a user, user input data that includes visual input;

detecting, in the user input data, events and metrics; and

generating audio that is synchronized with the detected events and corresponds to the detected metrics by:

selecting, from the library of stored music data, parts of the stored music data;

generating music data by assigning each of the selected parts of the stored music data to the detected events and metrics; and

translating the events and metrics detected in the user input data to audio by synthesizing the musical notes of the assigned parts of the generated music data;

generating visual effects that are synchronized with the events detected in the user input data and correspond to the metrics detected in the user input data by:

storing rules for translating the stored music data to visual effects; and

using the stored rules to translate the assigned parts to visual effects; and

generating video that includes the visual input, the generated audio, and the generated visual effects.

2. The method of claim 1 , wherein detecting events and metrics in the user input data comprises detecting objects or gestures in the visual input.

3. The method of claim 2 , wherein assigning each of the selected parts of the stored music data to each of the detected events and metrics comprises:

storing music data assignment heuristics that associate musical characteristics with visual characteristics;

identifying musical characteristics of each selected part of the stored music data;

identifying visual characteristics of each detected object or gesture;

using the music data assignment heuristics to assign each selected part to a detected object or gesture based on the musical characteristics of the selected part and the visual characteristics of the detected object or gesture.

4. The method of claim 3 , wherein storing the music data assignment heuristics comprises:

storing music data assignment training data that includes examples of musical attributes associated with visual attributes;

using a machine learning model, trained on the music data assignment training data, to generate the music data assignment heuristics.

5. The method of claim 2 , wherein detecting the events and metrics in the user input data further comprises:

generating virtual objects corresponding to the detected objects or gestures;

outputting those virtual objects to the user via a user interface;

detecting user interaction with the virtual objects via the user interface.

6. The method of claim 1 , wherein the library of stored music data includes notes, musical phrases, or musical effects.

7. The method of claim 6 , wherein synthesizing the generated music data comprises:

applying musical effects selected from the library of stored music data to parts selected from the library of stored music data; or

applying musical effects selected from the library of stored music data to input audio or input music data included in the user input data.

8. The method of claim 1 , wherein the user input data further includes touch input received via a touchpad.

9. The method of claim 1 , wherein the user input data further includes audio or music data.

10. The method of claim 1 , wherein storing the library of stored music data comprises:

storing music generation training data that includes compositions;

using a machine learning model, trained on the music data assignment training data, to generate the stored music data.

11. A system for creating music using visual input, comprising:

non-transitory computer readable storage media that stores a library of stored music data, the stored music data comprising a plurality of parts, each part comprising musical notes;

an event/metric detection unit that:

receives user input data, from a user, that includes visual input; and

detects events and metrics in the user input data;

a music data translation unit that translates the events and metrics detected in the user input data to generated music data by:

selecting parts of the stored music data; and

assigning each of the selected parts to the detected events and metrics;

an audio engine that generates audio that is synchronized with the detected events and corresponds to the detected metrics by synthesizing the musical notes of the assigned parts of the generated music data; and

a video engine that:

generates visual effects that are synchronized with the events detected in the user input data and correspond to the metrics detected in the user input data by applying rules to translate the generated music data to visual effects; and

generates video includes the generated audio, the visual input, and the generated visual effects.

12. The system of claim 11 , wherein the event/metric detection unit detects the events and metrics in the user input data by detecting objects or gestures in the visual input.

13. The system of claim 12 , wherein the music data translation unit assigns each of the selected parts of the stored music data to each of the detected events and metrics by:

identifying musical characteristics of each selected part of the stored music data;

identifying visual characteristics of each detected object or gesture;

using music data assignment heuristics, which associate musical characteristics with visual characteristics, to assign each selected part to a detected object or gesture based on the musical characteristics of the selected part and the visual characteristics of the detected object or gesture.

14. The system of claim 13 , further comprising:

a music data association model that uses a machine learning model, trained on music data assignment training data that includes examples of musical attributes associated with visual attributes, to generate the music data assignment heuristics.

15. The system of claim 12 , wherein the event/metric detection unit further detects the events and metrics in the user input data by:

generating virtual objects corresponding to the detected objects or gestures;

outputting those virtual objects to the user via a user interface;

detecting user interaction with the virtual objects via the user interface.

16. The system of claim 11 , wherein the library of stored music data includes notes, musical phrases, or musical effects.

17. The system of claim 16 , wherein the audio engine synthesizes the generated music data by:

applying musical effects selected from the library of stored music data to parts selected from the library of stored music data; or

applying musical effects selected from the library of stored music data to input audio or input music data included in the user input data.

18. The system of claim 11 , wherein the user input data includes touch input received via a touchpad.

19. The system of claim 11 , wherein the user input data further includes audio or music data.

20. The system of claim 11 , further comprising:

a music generation model that uses a machine learning model, trained on music generation training data that includes compositions, to generate the stored music data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2022
From: BUTERA, MICHAEL V.
To: ARTIPHON, INC.
Reel/Frame 059630/0900 →
Continuity (2)
Provisional Application 63175156 · Apr 15, 2021
Related Publication 20220335974A1 · Oct 20, 2022