IP Library › Granted Patent US 11,947,628
Granted Patent B2
US 11,947,628 · App. 17/217,139 · Granted Apr 2, 2024

Neural networks for accompaniment extraction from songs

Inventor: Gurunandan Krishnan Gorumkonda (Seattle, WA)
Assignee: Snap Inc.
G06F18/2148G06N3/04G06N3/08G10L21/14G10L25/30G06F2218/08G06F2218/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,947,628
App. No.
17/217,139
Granted
Apr 2, 2024
Kind
B2
Abstract

A messaging system that extracts accompaniment portions from songs. Methods of accompaniment extraction from songs includes receiving an input song that includes a vocal portion and an accompaniment portion, transforming the input song to an input image, where the input image represents the frequencies and intensities of the input song, processing the input image using a convolutional neural network (CNN) to generate an output image, and transforming the output image to an output accompaniment, where the output accompaniment includes the accompaniment of the input song.

Claims (51)

1. A method comprising:

training a second convolutional neural network (CNN) using a ground truth input song paired with a modified ground truth output song, wherein the modified ground truth output song is a ground truth output song with added vocal portions and an estimate of a clean loss for the modified ground truth output song;

processing the output image using a second CNN to generate a clean loss, the clean loss indicating an amount of the output image that is not attributed to the accompaniment portion;

training a first CNN based on the clean loss;

receiving an input song comprising a vocal portion and an accompaniment portion;

transforming the input song to an input image, the input image representing the frequencies and intensities of the input song;

processing the input image using the first CNN to generate an output image;

transforming the output image to an output accompaniment, the output accompaniment comprising the accompaniment of the input song.

2. The method of claim 1 further comprising:

determining a loss between the output song and a ground truth output song, wherein the input song is a ground truth input song paired with the ground truth output song; and

training the first CNN further based on the determined loss.

3. The method of claim 1 wherein the accompaniment portion of the input song is blank or the vocal portion of the input song is blank.

4. The method of claim 1 wherein the estimate of the clean loss is based on an amount of the added vocal portions in proportion to the ground truth output song.

5. The method of claim 1 further comprising:

storing first feature values determined in processing the output image by the second CNN;

processing the ground truth output song with the second CNN;

storing second feature values determined in processing the ground truth output song by the second CNN;

determining a feature loss based on differences between the second feature values and the first feature values; and

training the first CNN further based on the feature loss.

6. The method of claim 5 , wherein the input image has a pixel height and a pixel width, and wherein processing the output image by the second CNN further comprises:

determining a first layer of the first features values by convolving a plurality of kernels with the output image to determine feature values for the first layer, wherein the plurality of kernels have kernel pixel heights less than the pixel height of the input image and kernel pixel widths less than the pixel width of the input image, and wherein each kernel of the plurality of kernels generates a sublayer of the first layer.

7. The method of claim 1 wherein the second CNN does not comprise a maximum pooling layer.

8. The method of claim 5 wherein training the first CNN further based on the feature loss further comprises:

training the first CNN using stochastic gradient descent to minimize a weighted combination of the feature loss, the clean loss, and the determined loss.

9. The method of claim 1 wherein the first CNN comprises multiple convolution layers, a maximum pooling layer, an up-conversion layer, and a fully-connected layer.

10. A system comprising:

one or more computer processors; and

one or more computer-readable mediums storing instructions that, when executed by the one or more computer processors, cause the system to perform operations comprising:

training a second convolutional neural network (CNN) using a ground truth input song paired with a modified ground truth output song, wherein the modified ground truth output song is a ground truth output song with added vocal portions and an estimate of a clean loss for the modified ground truth output song;

processing the output image using a second CNN to generate a clean loss, the clean loss indicating an amount of the output image that is not attributed to the accompaniment portion;

training a first CNN based on the clean loss;

receiving an input song comprising a vocal portion and an accompaniment portion;

transforming the input song to an input image, the input image representing the frequencies and intensities of the input song;

processing the input image using the first CNN to generate an output image;

transforming the output image to an output accompaniment, the output accompaniment comprising the accompaniment of the input song.

11. The system of claim 10 wherein the operations further comprise:

determining a loss between the output song and a ground truth output song, wherein the input song is a ground truth input song paired with the ground truth output song; and

training the first CNN further based on the determined loss.

12. The system of claim 10 wherein the accompaniment portion of the input song is blank or the vocal portion of the input song is blank.

13. A non-transitory computer-readable storage medium including instructions that, when processed by a computer, configure the computer to perform operations comprising:

training a second convolutional neural network (CNN) using a ground truth input song paired with a modified ground truth output song, wherein the modified ground truth output song is a ground truth output song with added vocal portions and an estimate of a clean loss for the modified ground truth output song;

processing the output image using a second CNN to generate a clean loss, the clean loss indicating an amount of the output image that is not attributed to the accompaniment portion;

training a first CNN based on the clean loss;

receiving an input song comprising a vocal portion and an accompaniment portion;

transforming the input song to an input image, the input image representing the frequencies and intensities of the input song;

processing the input image using the first CNN to generate an output image;

transforming the output image to an output accompaniment, the output accompaniment comprising the accompaniment of the input song.

14. The non-transitory computer-readable storage medium of claim 13 wherein the operations further comprise:

determining a loss between the output song and a ground truth output song, wherein the input song is a ground truth input song paired with the ground truth output song; and

training the first CNN further based on the determined loss.

15. The non-transitory computer-readable storage medium of claim 13 wherein the accompaniment portion of the input song is blank or the vocal portion of the input song is blank.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2024
From: KRISHNAN GORUMKONDA, GURUNANDAN
To: SNAP INC.
Reel/Frame 066066/0499 →
Continuity (1)
Related Publication 20220318566A1 · Oct 6, 2022
Cited By (1)
US 12,524,497