IP Library › Granted Patent US 12,488,779
Granted Patent B2
US 12,488,779 · App. 18/199,670 · Granted Dec 2, 2025

Method and apparatus for voice synthesis based on brain waves during imagined speech

Inventors: Seong Whan Lee (Seoul, KR); Young Eun Lee (Seongnam-si, KR); Seo Hyun Lee (Seoul, KR); Soo Won Kim (Daegu, KR); Sang Ho Kim (Seoul, KR); Byung Kwan Ko (Donghae-si, KR); Ji Won Lee (Seoul, KR); Jung Sun Lee (Changwon-si, KR)
Assignee: Korea University Research and Business Foundation
G10L13/027G06F3/015G10L13/047G10L15/26G10L25/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,779
App. No.
18/199,670
Granted
Dec 2, 2025
Kind
B2
Abstract

The present invention relates to a method and an apparatus for synthesizing the voice based on brain waves during imagined speech. The method for synthesizing the voice based on brain waves during imagined speech according to an embodiment of the present invention may include the following steps: a step to obtain the user's brain waves during imagined speech; a step to convert the above-mentioned brain waves of imagined speech into embedding vectors; a step to generate the mel-spectrograms based on the above-mentioned embedding vectors; a step to generate the voice using the above-mentioned mel-spectrograms; a step to output the above-mentioned voice.

Claims (37)

1 . A method for synthesizing a voice based on brain waves during imagined speech through an apparatus for synthesizing the voice based on brain waves during imagined speech, the method comprising the steps of:

a step to obtain user's brain waves during imagined speech;

a step to convert the obtained brain waves during imagined speech into embedding vectors;

a step to generate Mel-spectrograms based on the embedding vectors;

a step to generate the voice using the Mel-spectrograms;

a step to output the generated voice; and

a step to train a generator and a discriminator,

wherein the generator can generate the Mel-spectrograms using the embedding vectors, and the discriminator can distinguish between the Mel-spectrograms generated from brain waves and the Mel-spectrograms generated from voice, and

wherein the step to train the generator and the discriminator comprises the steps of:

a step to make the generator and the discriminator learn based on the brain waves and voice signals generated during speech;

a step to perform a transfer learning for the generator and the discriminator that have learned based on the brain waves and the voice signals generated during speech; and

a step to make the generator and the discriminator re-learn based on brain waves generated during imagined speech.

2 . The method according to claim 1 , wherein the step to make the generator and the discriminator learn based on the brain waves and the voice signals generated during speech comprises the steps of:

a substep to obtain brain waves and voice signals generated during the speech;

a substep to convert brain waves generated during the speech into embedding vectors;

a substep to generate the Mel-spectrograms based on the embedding vectors converted from brain waves during the speech; and

a substep to train the generator and the discriminator based on the Mel-spectrograms of voice signals and the Mel-spectrograms based on the embedding vectors converted from the brain waves generated during the speech.

3 . The method according to claim 1 ,

wherein the step to convert user's brain waves obtained during imagined speech into the embedding vectors is to convert user's imagined speech brain waves into an embedding space that maximizes the difference between phonemes or words of the imagined speech brain wave.

4 . The method according to claim 1 , further comprising:

a step to convert the voice into letters, phonemes, or pronunciation based on a voice recognition model.

5 . The method according to claim 4 ,

wherein the converted letters, phonemes, or pronunciations are used for learning of the generator.

6 . An apparatus for synthesizing a voice based on brain waves during imagined speech, the apparatus comprising a processor-based brain wave and voice signal recognition module configured to obtain the user's brain waves during imagined speech; a processor-based feature vector conversion module including an embedding transformer that converts the obtained imagined speech brain waves into embedding vectors; a processor-based voice synthesis module including a Mel-spectrogram generator and a vocoder, the voice synthesis module being configured to generate Mel-spectrograms based on the embedding vectors and generates the voice using the Mel-spectrograms; and a voice output module including a speaker or audio interface configured to output the generated voice, wherein the voice synthesis module comprises: a Mel-spectrogram generation component including a neural network generator comprising convolutional layers, residual blocks, or attention modules configured to generate the Mel-spectrograms from the embedding vectors; and a discrimination component including a neural network discriminator trained to distinguish between the Mel-spectrograms generated from brain waves and the Mel-spectrograms generated from voice signals, wherein the Mel-spectrogram generation component and the discrimination component are processor-implemented neural networks trained based on brain waves and voice signals generated during speech, wherein for the Mel-spectrogram generation component and the discrimination component that have been trained based on brain waves and voice signals generated during speech, a transfer learning process is performed by fine-tuning the neural networks, and wherein the Mel-spectrogram generation component and the discrimination component are retrained based on the brain waves generated during imagined speech.

7 . The apparatus according to claim 6 ,

wherein the processor-based Mel-spectrogram generation component and discrimination component are configured to:

obtain the brain waves and voice signals generated during speech,

convert the brain waves generated during speech into embedding vectors using the embedding transformation processor,

generate the Mel-spectrograms based on the embedding vectors converted from the brain waves during speech, and

train the Mel-spectrogram generation component and discrimination component based on Mel-spectrograms derived from the embedding vectors converted from brain waves during speech and on Mel-spectrograms of the voice signals.

8 . The apparatus according to claim 6 ,

wherein the feature vector conversion module includes an embedding transformation processor configured to convert the user's brain waves obtained during imagined speech into an embedding space that maximizes the difference between phonemes or words of the imagined speech brain waves.

9 . The apparatus according to claim 6 ,

wherein the voice synthesis module further comprises:

a letter recognition component including a pre-trained speech recognition model implemented in software or firmware and configured to convert the generated voice into letters, phonemes, or pronunciations.

10 . The apparatus according to claim 9 ,

wherein the converted letters, phonemes, or pronunciations produced by the letter recognition component including the pre-trained speech recognition model are used for learning and updating parameters of the Mel-spectrogram generation component.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2023
From: LEE, SEONG WHAN; LEE, YOUNG EUN; LEE, SEO HYUN; KIM, SOO WON; KIM, SANG HO; KO, BYUNG KWAN; LEE, JI WON; LEE, JUNG SUN
To: KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
Reel/Frame 063709/0432 →
Priority Claims (1)
KR 10-2022-0173551 · Dec 13, 2022 · national
Continuity (1)
Related Publication 20240194179A1 · Jun 13, 2024
References Cited (11)
US 12051428B1 · Petrochuk · 2024 [cited by examiner]
US 20150297106A1 · Pasley · 2015 [cited by examiner]
US 20150380009A1 · Chang · 2015 [cited by examiner]
US 20190295566A1 · Moghadamfalahi · 2019 [cited by examiner]
US 20200142481A1 · Lee · 2020 [cited by examiner]
US 20210050004A1 · Whiting · 2021 [cited by examiner]
US 20230025518A1 · Pazderski · 2023 [cited by examiner]
KR 1020200052807A · 2020 [cited by applicant]
KR 1020220004272A · 2022 [cited by applicant]
Herff, Christian, et al., “Towards direct speech synthesis from ECoG: A pilot study”, 38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), 2016, (p. 1540-1543). [cited by applicant]
Berezutskaya, Julia, et al., “Towards Naturalistic Speech Decoding from Intracranial Brain Data”, 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 2022, (p. 3100-3104). [cited by applicant]