IP Library › Granted Patent US 11,257,480
Granted Patent B2
US 11,257,480 · App. 16/807,851 · Granted Feb 22, 2022

Unsupervised singing voice conversion with pitch adversarial network

Inventors: Chengzhu Yu (Bellevue, WA); Heng Lu (Sammamish, WA); Chao Weng (Fremont, CA); Dong Yu (Bothell, WA)
Assignee: TENCENT AMERICA LLC
G10L13/0335G10L13/047G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,257,480
App. No.
16/807,851
Granted
Feb 22, 2022
Kind
B2
Abstract

A method, a computer readable medium, and a computer system are provided for singing voice conversion. Data corresponding to a singing voice is received. One or more features and pitch data are extracted from the received data using one or more adversarial neural networks. One or more audio samples are generated based on the extracted pitch data and the one or more features.

Claims (30)

1. A method for singing voice conversion performed by one or more computer processors, comprising:

receiving data corresponding to a singing voice;

extracting one or more features from the received data;

extracting pitch data from the received data based on a pitch regression adversarial neural network including a dropout layer, two convolutional neural networks, and a fully connected layer, the dropout layer being employed at a beginning of each of the two convolutional neural networks; and

generating one or more audio samples based on the extracted pitch data and the one or more features.

2. The method of claim 1 , wherein the features are extracted based on an identification of a singer associated with the singing voice.

3. The method of claim 2 , wherein the identification is performed by a singer classification adversarial neural network.

4. The method of claim 3 , wherein the singer classification adversarial neural network comprises a dropout layer, two convolutional neural networks, and a fully connected layer.

5. The method of claim 1 , further comprising calculating a singer classification loss value and a pitch regression loss value.

6. The method of claim 5 , wherein the singer classification loss value and pitch regression loss value are used as training values based on minimizing the singer classification loss value and pitch regression loss value.

7. The method of claim 1 , wherein the received singing voice data is compressed using an average pooling function.

8. The method of claim 1 , wherein the audio samples are generated without parallel data and without changing the content associated with the singing voice.

9. A computer system for singing voice conversion, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

receiving code configured to cause the one or more computer processors to receive data corresponding to a singing voice;

first extracting code configured to cause the one or more computer processors to extract one or more features from the received data;

second extracting code configured to cause the one or more computer processors to extract pitch data from the received data based on a pitch regression adversarial neural network including a dropout layer, two convolutional neural networks, and a fully connected layer, the dropout layer being employed at a beginning of each of the two convolutional neural networks; and

generating code configured to cause the one or more computer processors to generate one or more audio samples based on the extracted pitch data and the one or more features.

10. The computer system of claim 9 , wherein the features are extracted based on an identification of a singer associated with the singing voice.

11. The computer system of claim 10 , wherein the identification is performed by a singer classification adversarial neural network.

12. The computer system of claim 11 , wherein the singer classification adversarial neural network comprises a dropout layer, two convolutional neural networks, and a fully connected layer.

13. The computer system of claim 9 , further comprising calculating code configured to cause the one or more computer processors to calculate a singer classification loss value and a pitch regression loss value, wherein the singer classification loss value and pitch regression loss value are used as training values based on minimizing the singer classification loss value and pitch regression loss value.

14. The computer system of claim 9 , wherein the received singing voice data is compressed using an average pooling function.

15. The computer system of claim 9 , wherein the audio samples are generated without parallel data and without changing the content associated with the singing voice.

16. A non-transitory computer readable medium having stored thereon a computer program for singing voice conversion, the computer program configured to cause one or more computer processors to:

receive data corresponding to a singing voice;

extract one or more features from the received data;

extract pitch data from the received data based on a pitch regression adversarial neural network including a dropout layer, two convolutional neural networks, and a fully connected layer, the dropout layer being employed at a beginning of each of the two convolutional neural networks; and

generate one or more audio samples based on the extracted pitch data and the one or more features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: YU, CHENGZHU; LU, HENG; WENG, CHAO; YU, DONG
To: TENCENT AMERICA LLC
Reel/Frame 052079/0522 →
Continuity (1)
Related Publication 20210280165A1 · Sep 9, 2021
Cited By (1)
US 12,646,493