IP Library › Granted Patent US 11,908,447
Granted Patent B2
US 11,908,447 · App. 17/596,037 · Granted Feb 20, 2024

Method and apparatus for synthesizing multi-speaker speech using artificial neural network

Inventors: Joon Hyuk Chang (Seoul, KR); Jae Uk Lee (Seoul, KR)
Assignee: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
G10L13/047G10L17/04G10L17/06G10L17/18G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,447
App. No.
17/596,037
Granted
Feb 20, 2024
Kind
B2
Abstract

According to an aspect, method for synthesizing multi-speaker speech using an artificial neural network comprises generating and storing a speech learning model for a plurality of users by subjecting a synthetic artificial neural network of a speech synthesis model to learning, based on speech data of the plurality of users, generating speaker vectors for a new user who has not been learned and the plurality of users who have already been learned by using a speaker recognition model, determining a speaker vector having the most similar relationship with the speaker vector of the new user according to preset criteria out of the speaker vectors of the plurality of users who have already been learned, and generating and learning a speaker embedding of the new user by subjecting the synthetic artificial neural network of the speech synthesis model to learning, by using a value of a speaker embedding of a user for the determined speaker vector as an initial value and based on speaker data of the new user.

Claims (34)

1. A method for synthesizing multi-speaker speech using an artificial neural network, comprising:

generating and storing a speech learning model for a plurality of users by subjecting a synthetic artificial neural network of a speech synthesis model to learning, based on speech data of the plurality of users;

generating speaker vectors for a new user who has not been learned and the plurality of users who have already been learned by using a speaker recognition model;

determining a speaker vector having a most similar relationship with the speaker vector of the new user according to preset criteria out of the speaker vectors of the plurality of users who have already been learned;

generating and learning a speaker embedding of the new user by subjecting the synthetic artificial neural network of the speech synthesis model to learning, by using a value of a speaker embedding of a user for the determined speaker vector as an initial value and based on speaker data of the new user;

wherein the generating a speech learning model for the new user comprises performing the learning of the synthetic artificial neural network of the speech synthesis model only for a preset time that prevents overfitting,

wherein the preset time comprises a range of 10 seconds to 60 seconds, and

wherein the generating speaker vectors comprises generating the speaker vectors using an artificial neural network of the speaker recognition model, by using a speech signal of the user as an input value.

2. The method for synthesizing multi-speaker speech using the artificial neural network of claim 1 , wherein the determining a speaker vector having the most similar relationship according to preset criteria comprises:

determining based on values calculated by taking inner products of the speaker vector of the new user and the speaker vectors of the plurality of users who have already been learned.

3. The method for synthesizing multi-speaker speech using the artificial neural network of claim 2 , wherein the determining a speaker vector having the most similar relationship according to preset criteria comprises:

calculating cosine similarity values based on the values of the inner products calculated, and determining a speaker vector of a user having a highest cosine similarity value as the speaker vector having the most similar relationship with the speaker vector of the new user.

4. The method for synthesizing multi-speaker speech using the artificial neural network of claim 1 , further comprising:

converting a mel-scale spectrogram calculated through the synthetic artificial neural network of the speech synthesis model into speech through a Griffin-Lim algorithm.

5. The method for synthesizing multi-speaker speech using the artificial neural network of claim 1 , wherein the synthetic artificial neural network model of the speech synthesis model comprises a Tacotron 2 algorithm.

6. An apparatus for synthesizing multi-speaker speech using an artificial neural network, comprising:

a speech synthesizer configured to generate a speech learning model for a plurality of users by subjecting a synthetic artificial neural network of a speech synthesis model to learning, based on speech data of the plurality of users;

a storage configured to store information on the speech learning model generated;

a speaker vector generator configured to generate speaker vectors for a new user who has not been learned and the plurality of users who have already been learned by using a speaker recognition model; and

a similar-vector determinator configured to determine a speaker vector having a most similar relationship with the speaker vector of the new user according to preset criteria out of the speaker vectors of the plurality of users who have already been learned, wherein the speech synthesizer generates and learns a speaker embedding of the new user by subjecting the synthetic artificial neural network of the speech synthesis model to learning, by using a value of a speaker embedding of a user for the determined speaker vector as an initial value and based on speaker data of the new user,

wherein the speech synthesizer is configured to perform the learning of the synthetic artificial neural network of the speech synthesis model only for a preset time that prevents overfitting,

wherein the preset time comprises a range of 10 seconds to 60 seconds, and

wherein the speaker vector generator is configured to generate the speaker vectors using an artificial neural network of the speaker recognition model, by using a speech signal of the user as an input value.

7. The apparatus for synthesizing multi-speaker speech using the artificial neural network of claim 6 , wherein the similar-vector determinator determines the speaker vector having the most similar relationship with the speaker vector of the new user based on values calculated by taking inner products of the speaker vector of the new user and the speaker vectors of the plurality of users who have already been learned.

8. The apparatus for synthesizing multi-speaker speech using the artificial neural network of claim 7 , wherein the similar-vector determinator calculates cosine similarity values based on the values of the inner products calculated, and then determines a speaker vector of a user having a highest cosine similarity value as the speaker vector having the most similar relationship with the speaker vector of the new user.

9. An apparatus for synthesizing multi-speaker speech using an artificial neural network, comprising:

a communicator configured to receive, from an external server, information on a speech learning model for a plurality of users by subjecting a synthetic artificial neural network of a speech synthesis model to learning, based on speech data of the plurality of users;

a storage configured to store information on the speech learning model generated;

a speaker vector generator configured to generate speaker vectors for a new user who has not been learned and the plurality of users who have already been learned by using a speaker recognition model; and

a similar-vector determinator configured to determine a speaker vector having a most similar relationship with the speaker vector of the new user according to preset criteria out of the speaker vectors of the plurality of users who have already been learned; and

a speech synthesizer configured to generate and learn a speaker embedding of the new user by subjecting the synthetic artificial neural network of the speech synthesis model to learning, by using a value of a speaker embedding of a user for the determined speaker vector as an initial value and based on speaker data of the new user,

wherein the speech synthesizer is configured to perform the learning of the synthetic artificial neural network of the speech synthesis model only for a preset time that prevents overfitting,

wherein the preset time comprises a range of 10 seconds to 60 seconds, and

wherein the speaker vector generator is configured to generate the speaker vectors using an artificial neural network of the speaker recognition model, by using a speech signal of the user as an input value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2021
From: CHANG, JOON HYUK; LEE, JAE UK
To: IUCF-HYU (INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY)
Reel/Frame 058271/0283 →
Priority Claims (1)
KR 10-2020-0097585 · Aug 4, 2020 · national
Continuity (1)
Related Publication 20230178066A1 · Jun 8, 2023