IP Library › Granted Patent US 12,322,377
Granted Patent B2
US 12,322,377 · App. 17/856,377 · Granted Jun 3, 2025

Method and apparatus for data augmentation

Inventors: Byung-Ok Kang (Daejeon, KR); Jeon-Gue Park (Daejeon, KR); Hyung-Bae Jeon (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
G10L15/063G10L15/16G10L25/51G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,377
App. No.
17/856,377
Granted
Jun 3, 2025
Kind
B2
Abstract

Disclosed herein is a method for data augmentation, which includes pretraining latent variables using first data corresponding to target speech and second data corresponding to general speech, training data augmentation parameters by receiving the first data and the second data as input, and augmenting target data using the first data and the second data through the pretrained latent variables and the trained parameters.

Claims (36)

1. A method for data augmentation by an apparatus, the apparatus including a processor and a memory operably coupled to the processor, wherein the memory stores program instructions to be executed by the processor, the method comprising:

pretraining latent variables using first data corresponding to target speech and second data corresponding to general speech;

training data augmentation parameters by receiving the first data and the second data as input; and

augmenting target data using the first data and the second data through the pretrained latent variables and the trained parameters,

wherein the latent variables include

a first latent variable corresponding to content attributes including a phone sequence of refined speech corpus of speech data; and

a second latent variable corresponding to environment information of speech data,

wherein the environment information of the speech data includes information about at least one of a channel, noise, an accent, a tone, a rhythm, and a speech tempo of a speaker.

2. The method of claim 1 , wherein pretraining the latent variables includes

inferring, by an encoder of a variational autoencoder, latent variables using the first data and the second data; and

performing, by a decoder of the variational autoencoder, training so as to generate input data by receiving the latent variables as input, and

wherein training the data augmentation parameters comprises performing training using a structure in which a generative adversarial network is added to the variational autoencoder.

3. The method of claim 2 , wherein training the data augmentation parameters includes

generating first output through the variational autoencoder by receiving the first data as input;

generating second output through the decoder by receiving the second latent variable corresponding to the first data and the first latent variable corresponding to the second data; and

differentiating, by a discriminator, the first output and the second output from each other.

4. The method of claim 3 , wherein augmenting the target data comprises augmenting the target data through the decoder by receiving the second latent variable corresponding to the first data and the first latent variable corresponding to the second data.

5. The method of claim 4 , wherein training the data augmentation parameters comprises training the parameters using a first loss function corresponding to the encoder, a second loss function corresponding to the decoder, and a third loss function corresponding to the discriminator.

6. The method of claim 5 , wherein the parameters include a first parameter corresponding to the encoder, a second parameter corresponding to the decoder, and a third parameter corresponding to the discriminator.

7. An apparatus for data augmentation, comprising:

one or more processors; and

executable memory for storing at least one program executed by the one or more processors,

wherein the at least one program is configured to

pretrain latent variables using first data corresponding to target speech and second data corresponding to general speech;

train data augmentation parameters by receiving the first data and the second data as input; and

augment target data using the first data and the second data through the pretrained latent variables and the trained parameters,

wherein the latent variables include

a first latent variable corresponding to content attributes including a phone sequence of refined speech corpus of speech data; and

a second latent variable corresponding to environment information of speech data,

wherein the environment information of the speech data includes information about at least one of a channel, noise, an accent, a tone, a rhythm, and a speech tempo of a speaker.

8. The apparatus of claim 7 , wherein the at least one program performs training such that an encoder of a variational autoencoder infers latent variables using the first data and the second data and such that a decoder of the variational autoencoder generates input data by receiving the latent variables as input, and

wherein the at least one program trains the parameters using a structure in which a generative adversarial network is added to the variational autoencoder.

9. The apparatus of claim 8 , wherein the at least one program trains the parameters by generating first output through the variational autoencoder receiving the first data as input, generating second output through the decoder receiving the second latent variable corresponding to the first data and the first latent variable corresponding to the second data, and making a discriminator differentiate between the first output and the second output.

10. The apparatus of claim 9 , wherein the at least one program augments the target data through the decoder by receiving the second latent variable corresponding to the first data and the first latent variable corresponding to the second data as input.

11. The apparatus of claim 10 , wherein the at least one program trains the parameters using a first loss function corresponding to the encoder, a second loss function corresponding to the decoder, and a third loss function corresponding to the discriminator.

12. The apparatus of claim 11 , wherein the parameters include a first parameter corresponding to the encoder, a second parameter corresponding to the decoder, and a third parameter corresponding to the discriminator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2022
From: KANG, BYUNG-OK; PARK, JEON-GUE; JEON, HYUNG-BAE
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 060423/0430 →
Priority Claims (1)
KR 10-2021-0088496 · Jul 6, 2021 · national
Continuity (1)
Related Publication 20230009771A1 · Jan 12, 2023
References Cited (28)
US 9721559B2 · Cui et al. · 2017 [cited by applicant]
US 10387765B2 · Mailhe · 2019 [cited by examiner]
US 10410113B2 · Clayton · 2019 [cited by examiner]
US 10565758B2 · Hadap · 2020 [cited by examiner]
US 11514888B2 · Finkelstein · 2022 [cited by examiner]
US 11830476B1 · Karanasou · 2023 [cited by examiner]
US 20190026631A1 · Carr · 2019 [cited by examiner]
US 20190304480A1 · Narayanan et al. · 2019 [cited by applicant]
US 20200034661A1 · Kim et al. · 2020 [cited by applicant]
US 20200151963A1 · Lee et al. · 2020 [cited by applicant]
US 20200335086A1 · Paraskevopoulos · 2020 [cited by examiner]
US 20210103721A1 · Im et al. · 2021 [cited by applicant]
US 20220101121A1 · Vahdat · 2022 [cited by examiner]
US 20230072255A1 · Klein · 2023 [cited by examiner]
US 20230104417A1 · Vechtomova · 2023 [cited by examiner]
KR 1020190106861A · 2019 [cited by applicant]
KR 1020200044337A · 2020 [cited by applicant]
KR 1020200107389A · 2020 [cited by applicant]
KR 102158743B1 · 2020 [cited by applicant]
KR 20210036692A · 2021 [cited by applicant]
KR 20210041567A · 2021 [cited by applicant]
Hsu, Wei-Ning, Yu Zhang, and James Glass. “Unsupervised learning of disentangled and interpretable representations from sequential data,” Advances in neural information processing systems 30 (2017) (Year: 2017). [cited by examiner]
B. He, S. Wang, W. Yuan, J. Wang and M. Unoki, “Data Augmentation for Monaural Singing Voice Separation Based on Variational Autoencoder-Generative Adversarial Network,” 2019 IEEE International Conference on Multimedia … [cited by examiner]
X. Xia, R. Togneri, F. Sohel and D. Huang, “Auxiliary Classifier Generative Adversarial Network With Soft Labels in Imbalanced Acoustic Event Detection,” in IEEE Transactions on Multimedia, vol. 21, No. 6, pp. 1359-1371… [cited by examiner]
Byung Ok Kang et.al, “Speech Recognition for Task Domains with Sparse Matched Training Data”, Applied sciences, Sep. 2020. [cited by applicant]
Chin-Cheng Hsu et al., “Voice Conversion from Unaligned Corpora using Variational Autoencoding Wasserstein Generative Adversarial Networks”., Interspeech, Apr. 2017. [cited by applicant]
Marc'Aurelio Ranzato et.al, “Semi-supervised Learning of Compact Document Representations with Deep Networks”, International Conference on Machine Learning (ICML), Jul. 2008. [cited by applicant]
Wei-Ning Hsu et al., “Unsupervised learning of disentangled and interpretable representations from sequential data”., Neural Information Processing Systems, Sep. 22, 2017. [cited by applicant]