IP Library Granted Patent US 10,453,476
Granted Patent B1
US 10,453,476 · App. 15/657,003 · Granted Oct 22, 2019

Split-model architecture for DNN-based small corpus voice conversion

Inventor: Sandesh Aryal (Pasadena, CA)
Assignee: OBEN, INC.
G10L25/30G06N3/04G10L21/013G10L25/24G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,453,476
App. No.
15/657,003
Granted
Oct 22, 2019
Kind
B1
Abstract

A voice conversion system suitable for encoding small and large corpuses is disclosed. The voice conversion system comprises hardware including a neural network for generating estimated target speech data based on source speech data. The neural network includes an input layer, an output layer, and a novel split-model hidden layer. The input layer comprises a first portion and a second portion. The output layer comprises a third portion and a fourth portion. The hidden layer comprises a first subnet and a second subnet, wherein the first subnet is directly connected to the first portion of the input layer and the third portion of the output layer, and wherein the second subnet is directly connected to the second portion of the input layer and the fourth portion of the output layer. The first subnet and second subnet operate in parallel, and link to different but overlapping nodes of the input layer.

Claims (19)

1. A voice conversion system comprising:

a microphone for recording source speech data;

a neural network for generating estimated target speech data based on the source speech data, wherein the neural network comprises:

a) an input layer comprising a first portion and a second portion, wherein the first portion is associated with a first plurality of audio features, and the second portion is associated with a second plurality of audio features; wherein second plurality of audio features overlaps with the first plurality of audio features and comprises each of the first plurality of audio features; wherein the first plurality of audio features comprises MCEP coefficients one through 10; and the second plurality of audio features comprises MCEP coefficients one through 30;

b) an output layer comprising a third portion and a fourth portion, wherein the third portion is associated with a third plurality of audio features, and the fourth portion is associated with a fourth plurality of audio features; and

c) a hidden layer comprising a first subnet and a second subnet; wherein the first subnet is directly connected to the first portion of the input layer and the third portion of the output layer; and wherein the second subnet is directly connected to the second portion of the input layer and the fourth portion of the output layer;

a waveform generator configured to generate a target voice signal based on the estimated target speech data; and

a speaker configured to play the target voice signal.

2. The voice conversion system of claim 1 , wherein the first subnet comprises a single layer of nodes, and wherein the second subnet comprises two layers of nodes.

3. The voice conversion system of claim 2 , wherein the first subnet comprises a single layer of nodes comprising 64 nodes, and wherein the second subnet comprises two layers of 256 nodes each.

4. A voice conversion system comprising:

a microphone for recording source speech data;

a neural network for generating estimated target speech data based on the source speech data, wherein the neural network comprises:

a) an input layer comprising a first portion and a second portion, wherein the first portion is associated with a first plurality of audio features, and the second portion is associated with a second plurality of audio features; wherein second plurality of audio features overlaps with the first plurality of audio features and comprises each of the first plurality of audio features;

b) an output layer comprising a third portion and a fourth portion, wherein the third portion is associated with a third plurality of audio features, and the fourth portion is associated with a fourth plurality of audio features; and

c) a hidden layer comprising a first subnet and a second subnet; wherein the first subnet is directly connected to the first portion of the input layer and the third portion of the output layer; and wherein the second subnet is directly connected to the second portion of the input layer and the fourth portion of the output layer; wherein the first subnet comprises a single layer of nodes comprising 64 nodes, and wherein the second subnet comprises two layers of 256 nodes each;

a waveform generator configured to generate a target voice signal based on the estimated target speech data; and

a speaker configured to play the target voice signal.

5. The voice conversion system of claim 4 , wherein the first plurality of audio features comprises MCEP coefficients one through 10; and the second plurality of audio features comprises MCEP coefficients one through 30.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2019
From: ARYAL, SANDESH
To: OBEN, INC.
Reel/Frame 050323/0924 →
Continuity (1)
Provisional Application 62365022 · Jul 21, 2016
Cited By (3)
US 12,341,619 US 12,412,588 US 12,676,138