IP Library › Granted Patent US 12,536,987
Granted Patent B2
US 12,536,987 · App. 18/284,146 · Granted Jan 27, 2026

Method and device for speech synthesis based on multi-speaker training data sets

Inventors: Joon Hyuk Chang (Seoul, KR); Jae Uk Lee (Seoul, KR)
Assignee: Industry-University Cooperation Foundation Hanyang University
G10L13/08G10L13/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,987
App. No.
18/284,146
Granted
Jan 27, 2026
Kind
B2
Abstract

An exemplary embodiment of the present disclosure is a speech synthesis method based on multi-speaker training dataset of a speech synthesis apparatus including pre-training a speech synthesis model using a previously stored neural network with a single speaker training dataset with most spoken sentences, among training datasets of a plurality of speakers, fine-tuning the pre-trained speech synthesis model with a training dataset of a plurality of speakers, and applying a target speech dataset to the fine-tuned speech synthesis model to be converted into a mel spectrogram.

Claims (35)

1 . A speech synthesis method based on a multi-speaker training dataset of a speech synthesis apparatus, comprising:

pre-training by applying a single speaker training dataset with most spoken sentences, among training datasets of a plurality of speakers, to a speech synthesis model using a previously stored neural network;

fine-tuning the pre-trained speech synthesis model with a training dataset of the plurality of speakers; and

converting to a mel spectrogram by applying a target speech dataset to the fine-tuned speech synthesis model,

wherein the speech synthesis model conditions a target speaker based on a speaker embedding and a score vector, the score vector being calculated by:

calculating similarity scores between the target speaker and the plurality of speakers with respect to a speaking characteristic;

selecting a speaker, from among the plurality of speakers, based on the similarity scores, such that the selected speaker has the most similar speaking characteristic to the target speaker;

multiplying (i) a one-hot vector identifying the selected speaker from among the plurality of speakers, with (ii) a similarity score between the target speaker and the selected speaker; and

using a result of the multiplication as the score vector.

2 . The speech synthesis method based on a multi-speaker training dataset of claim 1 , further comprising:

generating a waveform type speech file based on the mel spectrogram.

3 . The speech synthesis method based on a multi-speaker training dataset of claim 1 , wherein the converting to a mel spectrogram further includes:

selecting a speaker with a high similarity to a target speech, among previously stored training datasets of the plurality of speakers, based on the target speech dataset; and

setting speaker embedding of the selected speaker as an initial value of the fine-tuned speech synthesis model.

4 . The speech synthesis method based on a multi-speaker training dataset of claim 3 , wherein in the selecting a speaker with a high similarity, the similarity is calculated based on a tone color and a speech rate between speakers.

5 . The speech synthesis method based on a multi-speaker training dataset of claim 4 , wherein the selecting a speaker with a high similarity includes:

calculating a similarity of a tone color by extracting a feature vector from the speech between the speakers and using the inner product;

calculating a phoneme duration of each speaker and calculating a similarity of a speech rate by dividing the phoneme duration of the trained speaker by a phoneme duration of the target speaker;

calculating a tone color similarity score between two speakers based on the similarity of the tone color and the similarity of the speech rate; and

selecting a speaker with the highest tone color similarity score as a speaker with the most similar speaking characteristic.

6 . The speech synthesis method based on a multi-speaker training dataset of claim 1 , wherein the speech synthesis model conditions each speaker based on a trainable speaker embedding and a one-hot vector.

7 . A speech synthesis apparatus based on a multi-speaker training dataset, comprising:

a memory which stores one or more instructions; and

a processor which executes the one or more instructions which are stored in the memory,

wherein the processor executes the one or more instructions to pre-train a speech synthesis model using a previously stored neural network with a single speaker training dataset with most spoken sentences, among training datasets of a plurality of speakers, fine-tune the pre-trained speech synthesis model with a training dataset of the plurality of speakers, and apply a target speech dataset to the fine-tuned speech synthesis model to be converted into a mel spectrogram, and

wherein the speech synthesis model conditions a target speaker based on a speaker embedding and a score vector, the score vector being calculated by:

calculating similarity scores between the target speaker and the plurality of speakers with respect to a speaking characteristic;

selecting a speaker, from among the plurality of speakers, based on the similarity scores, such that the selected speaker has the most similar speaking characteristic to the target speaker;

multiplying (i) a one-hot vector identifying the selected speaker from among the plurality of speakers, with (ii) a similarity score between the target speaker and the selected speaker; and

using a result of the multiplication as the score vector.

8 . The speech synthesis apparatus based on a multi-speaker training dataset of claim 7 , wherein the processor generates a waveform type speech file based on the mel spectrogram.

9 . The speech synthesis apparatus based on a multi-speaker training dataset of claim 7 , wherein the processor is configured to select a speaker with a high similarity to a target speech, among previously stored training datasets of the plurality of speakers, based on the target speech dataset; and set a speaker embedding of the selected speaker as an initial value of the fine-tuned speech synthesis model.

10 . The speech synthesis apparatus based on a multi-speaker training dataset of claim 9 , wherein the processor calculates the similarity based on the tone color and the speech rate between the speakers.

11 . The speech synthesis apparatus based on a multi-speaker training dataset of claim 10 , wherein the processor is configured to calculate a similarity of a tone color by extracting a feature vector from the speech between the speakers and using the inner product, calculate a phoneme duration of each of the speakers, calculate a similarity of a speech rate by dividing the phoneme duration of the trained speaker by a phoneme duration of the target speaker, calculate a tone color similarity score between two speakers based on the similarity of the tone color and the similarity of the speech rate, and select a speaker with the highest tone color similarity score as a speaker with the most similar speaking characteristic.

12 . The speech synthesis apparatus based on a multi-speaker training dataset of claim 7 , wherein the speech synthesis model conditions each speaker based on a trainable speaker embedding and a one-hot vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2023
From: CHANG, JOON HYUK; LEE, JAE UK
To: INDUSTRY-UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
Reel/Frame 065027/0719 →
Priority Claims (1)
KR 10-2021-0039888 · Mar 26, 2021 · national
Continuity (1)
Related Publication 20240169973A1 · May 23, 2024
References Cited (62)
US 6813604B1 · Shih · 2004 [cited by examiner]
US 10699697B2 · Qian · 2020 [cited by examiner]
US 10930263B1 · Mahyar · 2021 [cited by examiner]
US 11276389B1 · Aryal · 2022 [cited by examiner]
US 11410684B1 · Klimkov · 2022 [cited by examiner]
US 11605388B1 · Gupta · 2023 [cited by examiner]
US 20050182630A1 · Miro · 2005 [cited by examiner]
US 20070168189A1 · Tamura · 2007 [cited by examiner]
US 20070256189A1 · Tian · 2007 [cited by examiner]
US 20080195386A1 · Proidl · 2008 [cited by examiner]
US 20090177473A1 · Aaron · 2009 [cited by examiner]
US 20100198600A1 · Masuda · 2010 [cited by examiner]
US 20110087488A1 · Morinaka · 2011 [cited by examiner]
US 20110218804A1 · Chun · 2011 [cited by examiner]
US 20120059654A1 · Nishimura · 2012 [cited by examiner]
US 20120095767A1 · Hirose · 2012 [cited by examiner]
US 20140052447A1 · Tachibana · 2014 [cited by examiner]
US 20150058015A1 · Mitsufuji · 2015 [cited by examiner]
US 20150112687A1 · Bredikhin · 2015 [cited by examiner]
US 20150127350A1 · Agiomyrgiannakis · 2015 [cited by examiner]
US 20150187356A1 · Aronowitz · 2015 [cited by examiner]
US 20150228271A1 · Morita · 2015 [cited by examiner]
US 20170301340A1 · Yassa · 2017 [cited by examiner]
US 20180211649A1 · Li · 2018 [cited by examiner]
US 20190019500A1 · Jang · 2019 [cited by examiner]
US 20190244623A1 · Hall · 2019 [cited by examiner]
US 20200058290A1 · Chae · 2020 [cited by examiner]
US 20200135209A1 · Delfarah · 2020 [cited by examiner]
US 20200342852A1 · Kim · 2020 [cited by examiner]
US 20200349922A1 · Peyser · 2020 [cited by examiner]
US 20200372897A1 · Battenberg · 2020 [cited by examiner]
US 20200380952A1 · Zhang · 2020 [cited by examiner]
US 20200394998A1 · Kim · 2020 [cited by examiner]
US 20200410976A1 · Zhou · 2020 [cited by examiner]
US 20210020161A1 · Gao · 2021 [cited by examiner]
US 20210209315A1 · Jia · 2021 [cited by examiner]
US 20210217404A1 · Jia · 2021 [cited by examiner]
US 20210248997A1 · Yu · 2021 [cited by examiner]
US 20220013106A1 · Deng · 2022 [cited by examiner]
US 20220051655A1 · Kanagawa · 2022 [cited by examiner]
US 20220068256A1 · Jia · 2022 [cited by examiner]
US 20220068257A1 · Biadsy · 2022 [cited by examiner]
US 20220068259A1 · Pan · 2022 [cited by examiner]
US 20220122579A1 · Biadsy · 2022 [cited by examiner]
US 20220148562A1 · Park · 2022 [cited by examiner]
US 20220189455A1 · Pekar · 2022 [cited by examiner]
US 20220230625A1 · Zhu · 2022 [cited by examiner]
US 20220230629A1 · Zhu · 2022 [cited by examiner]
US 20220238116A1 · Gao · 2022 [cited by examiner]
US 20220293091A1 · Pan · 2022 [cited by examiner]
US 20220301542A1 · Sung · 2022 [cited by examiner]
US 20220310058A1 · Zhao · 2022 [cited by examiner]
US 20220406289A1 · Kanagawa · 2022 [cited by examiner]
US 20230081659A1 · Pan · 2023 [cited by examiner]
US 20230178066A1 · Chang · 2023 [cited by examiner]
US 20230186937A1 · Uhlich · 2023 [cited by examiner]
US 20230298567A1 · Tan · 2023 [cited by examiner]
KR 1020190008137A · 2019 [cited by applicant]
KR 1020190100095A · 2019 [cited by applicant]
Kurniawati Azizah et al., “Hierarchical Transfer Learning for Multilingual, Multi-Speaker, and Style Transfer DNN-Based TTS on Low-Resource Languages”, IEEE Access, Sep. 29, 2020, vol. 8, pp. 179798-179812. [cited by applicant]
Henry B. Moss et al., “Boffin TTS: Few-Shot Speaker Adaptation by Bayesian Optimization”, Amazon Research, 5 Pages. [cited by applicant]
International Search Report for PCT/KR2021/017121 dated Jul. 19, 2022. [cited by applicant]