IP Library › Granted Patent US 11,862,187
Granted Patent B2
US 11,862,187 · App. 17/751,471 · Granted Jan 2, 2024

Systems and methods for jointly estimating sound sources and frequencies from audio

Inventors: Andreas Jansson (New York, NY); Rachel Bittner (New York, NY)
Assignee: Spotify AB
G10L25/51G06N3/045G06N3/08G06N20/00H04L65/75
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,862,187
App. No.
17/751,471
Granted
Jan 2, 2024
Kind
B2
Abstract

An electronic device receives a first audio content item that includes a plurality of sound sources. The electronic device generates a representation of the first audio content item. The electronic device determines, from the representation of the first audio content item: a representation of an isolated sound source, and frequency data associated with the isolated sound source. Determining the representation of the isolated sound source and the frequency data associated with the isolated sound source includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source. The electronic device determines that a portion of a second audio content item matches the first audio content item using the representation of the isolated sound source and/or the frequency data associated with the isolated sound source.

Claims (65)

1. A method, comprising:

at a first electronic device, the first electronic device having one or more processors and memory storing instructions for execution by the one or more processors:

receiving a first audio content item that includes a plurality of sound sources;

generating a representation of the first audio content item; and

determining, from the representation of the first audio content item:

a representation of an isolated sound source, and

frequency data associated with the isolated sound source, wherein the step of determining the representation of the isolated sound source and the frequency data associated with the isolated sound source includes using a neural network system to jointly determine the representation of the isolated sound source using a first neural network and the frequency data associated with the isolated sound source using a second neural network, wherein weights of the first neural network and weights of the second neural network are trained simultaneously; and

determining that a portion of a second audio content item matches the first audio content item using the representation of the isolated sound source or the frequency data associated with the isolated sound source.

2. The method of claim 1 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

3. The method of claim 1 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining a representation of the isolated sound source for the second audio content item; and

comparing the representation of the isolated sound source for the second audio content item with the representation of the isolated sound source of the first audio content item.

4. The method of claim 1 , wherein the neural network system comprises a plurality of U-nets.

5. The method of claim 1 , wherein:

the neural network system comprises a first source network, a first pitch network, a second source network, and a second pitch network,

the second source network is fed a concatenation of an output of the first source network with an output of the first pitch network, and

the output of the second source network is fed to the second pitch network.

6. The method of claim 1 , wherein generating the representation of the first audio content item comprises:

determining a first set of weights for a source network of a source-to-pitch network;

feeding a pitch network of the source-to-pitch network an output of the source network of the source-to-pitch network; and

determining a second set of weights for the pitch network of the source-to-pitch network.

7. The method of claim 1 , wherein the isolated sound source comprises a vocal source.

8. The method of claim 1 , wherein the isolated sound source comprises an instrumental source.

9. A first electronic device comprising:

one or more processors; and

memory storing instructions for execution by the one or more processors, the instructions including instructions for:

receiving a first audio content item that includes a plurality of sound sources;

generating a representation of the first audio content item; and

determining, from the representation of the first audio content item:

a representation of an isolated sound source, and

frequency data associated with the isolated sound source, wherein the step of determining the representation of the isolated sound source and the frequency data associated with the isolated sound source includes using a neural network system to jointly determine the representation of the isolated sound source using a first neural network and the frequency data associated with the isolated sound source using a second neural network, wherein weights of the first neural network and weights of the second neural network are trained simultaneously; and

determining that a portion of a second audio content item matches the first audio content item using the representation of the isolated sound source or the frequency data associated with the isolated sound source.

10. The first electronic device of claim 9 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

11. The first electronic device of claim 9 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining a representation of the isolated sound source for the second audio content item; and

comparing the representation of the isolated sound source for the second audio content item with the representation of the isolated sound source of the first audio content item.

12. The first electronic device of claim 9 , wherein the neural network system comprises a plurality of U-nets.

13. The first electronic device of claim 9 , wherein:

the neural network system comprises a first source network, a first pitch network, a second source network, and a second pitch network,

the second source network is fed a concatenation of an output of the first source network with an output of the first pitch network, and

the output of the second source network is fed to the second pitch network.

14. The first electronic device of claim 9 , wherein generating the representation of the first audio content item comprises:

determining a first set of weights for a source network of a source-to-pitch network;

feeding a pitch network of the source-to-pitch network an output of the source network of the source-to-pitch network; and

determining a second set of weights for the pitch network of the source-to-pitch network.

15. The first electronic device of claim 9 , wherein the isolated sound source comprises a vocal source.

16. The first electronic device of claim 9 , wherein the isolated sound source comprises an instrumental source.

17. A non-transitory computer-readable storage medium storing instructions, which when executed by an electronic device, cause the electronic device to:

receive a first audio content item that includes a plurality of sound sources;

generate a representation of the first audio content item; and

determine, from the representation of the first audio content item:

a representation of an isolated sound source, and

frequency data associated with the isolated sound source, wherein the step of determining the representation of the isolated sound source and the frequency data associated with the isolated sound source includes using a neural network system to jointly determine the representation of the isolated sound source using a first neural network and the frequency data associated with the isolated sound source using a second neural network, wherein weights of the first neural network and weights of the second neural network are trained simultaneously; and

determine that a portion of a second audio content item matches the first audio content item using the representation of the isolated sound source or the frequency data associated with the isolated sound source.

18. The non-transitory computer-readable storage medium of claim 17 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

19. The non-transitory computer-readable storage medium of claim 17 , wherein determining that the portion of the second audio content item matches the first audio content item includes:

determining a representation of the isolated sound source for the second audio content item; and

comparing the representation of the isolated sound source for the second audio content item with the representation of the isolated sound source of the first audio content item.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the neural network system comprises a plurality of U-nets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2022
From: JANSSON, ANDREAS; BITTNER, RACHEL
To: SPOTIFY AB
Reel/Frame 060006/0898 →
Continuity (2)
Continuation 16596554 · Oct 8, 2019
Related Publication 20220351747A1 · Nov 3, 2022