IP Library Granted Patent US 11,355,137
Granted Patent B2
US 11,355,137 · App. 16/596,554 · Granted Jun 7, 2022

Systems and methods for jointly estimating sound sources and frequencies from audio

Inventors: Andreas Jansson (New York, NY); Rachel Bittner (New York, NY)
Assignee: Spotify AB
G10L25/51G06N3/0454G06N3/08G06N20/00H04L65/601
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,355,137
App. No.
16/596,554
Granted
Jun 7, 2022
Kind
B2
Abstract

An electronic device receives a first audio content item that includes a plurality of sound sources. The electronic device generates a representation of the first audio content item. The electronic device determines, from the representation of the first audio content item, a representation of an isolated sound source and frequency data associated with the isolated sound source. The determining includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source.

Claims (69)

1. A method, comprising:

at a first electronic device, the first electronic device having one or more processors and memory storing instructions for execution by the one or more processors:

receiving a first audio content item that includes a plurality of sound sources;

generating a representation of the first audio content item; and

determining, from the representation of the first audio content item:

a representation of an isolated sound source, and

frequency data associated with the isolated sound source,

wherein the determining includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source, the neural network comprising a pitch network and a source network, wherein an output of the pitch network is fed to the source network.

2. The method of claim 1 , further comprising, at the first electronic device, determining that a portion of a second audio content item matches the first audio content item by:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

3. The method of claim 1 , further comprising, at the first electronic device, determining that a portion of a third audio content item matches the first audio content item by:

determining a representation of an isolated sound source for the third audio content item; and

comparing the representation of the isolated sound source for the third audio content item with the representation of the isolated sound source of the first audio content item.

4. The method of claim 1 , wherein the neural network comprises a plurality of U-nets.

5. The method of claim 1 , wherein:

the neural network comprises a first source network, a first pitch network, a second source network, and a second pitch network,

the second source network is fed a concatenation of an output of the first source network with an output of the first pitch network, and

the output of the second source network is fed to the second pitch network.

6. The method of claim 1 , wherein generating the representation of the first audio content item comprises:

determining a first set of weights for a source network of a source-to-pitch network;

feeding a pitch network of the source-to-pitch network an output of the source network of the source-to-pitch network; and

determining a second set of weights for the pitch network of the source-to-pitch network.

7. The method of claim 1 , wherein the isolated sound source comprises a vocal source.

8. The method of claim 1 , wherein the isolated sound source comprises an instrumental source.

9. A first electronic device comprising:

one or more processors; and

memory storing instructions for execution by the one or more processors, the instructions including instructions for:

receiving a first audio content item that includes a plurality of sound sources;

generating a representation of the first audio content item; and

determining, from the representation of the first audio content item, a representation of an isolated sound source and frequency data associated with the isolated sound source, wherein the determining includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source, the neural network comprising a pitch network and a source network, wherein an output of the pitch network is fed to the source network.

10. The first electronic device of claim 9 , the instructions further including instructions for, at the first electronic device, determining that a portion of a second audio content item matches the first audio content item by:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

11. The first electronic device of claim 9 , the instructions further including instructions for, at the first electronic device, determining that a portion of a third audio content item matches the first audio content item by:

determining a representation of an isolated sound source for the third audio content item; and

comparing the representation of the isolated sound source for the third audio content item with the representation of the isolated sound source of the first audio content item.

12. The first electronic device of claim 9 , wherein the neural network comprises a plurality of U-nets.

13. The first electronic device of claim 9 , wherein:

the neural network comprises a first source network, a first pitch network, a second source network, and a second pitch network,

the second source network is fed a concatenation of an output of the first source network with an output of the first pitch network, and

the output of the second source network is fed to the second pitch network.

14. The first electronic device of claim 9 , wherein generating the representation of the first audio content item comprises:

determining a first set of weights for a source network of a source-to-pitch network;

feeding a pitch network of the source-to-pitch network an output of the source network of the source-to-pitch network; and

determining a second set of weights for the pitch network of the source-to-pitch network.

15. The first electronic device of claim 9 , wherein the isolated sound source comprises a vocal source.

16. The first electronic device of claim 9 , wherein the isolated sound source comprises an instrumental source.

17. A non-transitory computer-readable storage medium storing instructions, which when executed by an electronic device, cause the electronic device to:

receive a first audio content item that includes a plurality of sound sources;

generate a representation of the first audio content item; and

determine, from the representation of the first audio content item, a representation of an isolated sound source and frequency data associated with the isolated sound source, wherein the determining includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source, the neural network comprising a pitch network and a source network, wherein an output of the pitch network is fed to the source network.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the instructions further cause the electronic device to determine that a portion of a second audio content item matches the first audio content item by:

determining frequency data for a representation of the second audio content item; and

comparing the frequency data of the second audio content item with the frequency data of the first audio content item.

19. The non-transitory computer-readable storage medium of claim 17 , wherein the instructions further cause the electronic device to determine that a portion of a third audio content item matches the first audio content item by:

determining a representation of an isolated sound source for the third audio content item; and

comparing the representation of the isolated sound source for the third audio content item with the representation of the isolated sound source of the first audio content item.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the neural network comprises a plurality of U-nets.

21. The non-transitory computer-readable storage medium of claim 17 , wherein:

the neural network comprises a first source network, a first pitch network, a second source network, and a second pitch network,

the second source network is fed a concatenation of an output of the first source network with an output of the first pitch network, and

the output of the second source network is fed to the second pitch network.

22. The non-transitory computer-readable storage medium of claim 17 , wherein generating the representation of the first audio content item comprises:

determining a first set of weights for a source network of a source-to-pitch network;

feeding a pitch network of the source-to-pitch network an output of the source network of the source-to-pitch network; and

determining a second set of weights for the pitch network of the source-to-pitch network.

23. The non-transitory computer-readable storage medium of claim 17 , wherein the isolated sound source comprises a vocal source.

24. The non-transitory computer-readable storage medium of claim 17 , wherein the isolated sound source comprises an instrumental source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2019
From: JANSSON, ANDREAS; BITTNER, RACHEL
To: SPOTIFY AB
Reel/Frame 050807/0235 →
Continuity (1)
Related Publication 20210104256A1 · Apr 8, 2021