IP Library Granted Patent US 12,548,550
Granted Patent B2
US 12,548,550 · App. 18/576,322 · Granted Feb 10, 2026

Generating method, generating program, and generating device

Inventor: Hiroki Kanagawa (Tokyo, JP)
Assignee: NTT, Inc.
G10L13/06G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,550
App. No.
18/576,322
Granted
Feb 10, 2026
Kind
B2
Abstract

A generation device generates intermediate representation information of a subband signal based on an acoustic feature value of a speech waveform, and simultaneously generates a plurality of subband signals corresponding to a plurality of different times and a plurality of different bands by inputting the intermediate representation information to a plurality of probability distribution generation models that outputs information on subband signals corresponding to times and bands allocated respectively, the plurality of probability distribution generation models corresponding to the number of channels of the subband signals and the number of samples to be simultaneously generated; and generates the speech waveform based on the plurality of subband signals.

Claims (38)

1 . A computer implemented method comprising:

generating intermediate representation information of a subband signal based on an acoustic feature value of a speech waveform;

simultaneously generating a plurality of subband signals corresponding to a plurality of different times and a plurality of different bands by inputting the intermediate representation information to a plurality of probability distribution generation models, wherein the plurality of probability distribution generation models outputs information on subband signals corresponding to times and bands allocated respectively, and the plurality of probability distribution generation models corresponds to a number of channels of the subband signals and a number of samples to be simultaneously generated; and

generating the speech waveform based on the plurality of subband signals.

2 . The computer implemented method according to claim 1 , further comprising:

converting the acoustic feature value of the speech waveform into the intermediate representation information of the acoustic feature value by using a first intermediate representation model, wherein the first intermediate representation model outputs the intermediate representation information of the acoustic feature value in a case where the acoustic feature value is input.

3 . The computer implemented method according to claim 2 , wherein the generating intermediate representation information further comprises generating the intermediate representation information of the subband signal using a second intermediate representation model, wherein the second intermediate representation model outputs the intermediate representation information of the subband signal in a case where the intermediate representation information of the acoustic feature value is input.

4 . The computer implemented method according to claim 3 , further comprising:

calculating a loss value based on the plurality of subband signals, wherein the plurality of subband signals is calculated from the speech waveform; and

executing learning of at least one model, wherein the at least one model is among the first intermediate representation model, the second intermediate representation model, and the plurality of probability distribution generation models based on the loss value.

5 . The computer implemented method according to claim 2 , wherein the first intermediate representation model includes a neural network.

6 . The computer implemented method according to claim 1 , wherein the simultaneously generating a plurality of subband signals further comprises generating the plurality of subband signals using a simultaneous probability distribution generation model, and the simultaneous probability distribution generation model simultaneously outputs information of subband signals corresponding to a plurality of time zones and a plurality of bands from one model.

7 . The computer implemented method according to claim 1 , wherein the intermediate representation information is in multi-dimensional vector form.

8 . A computer-readable non-transitory recording medium storing a computer-executable program instructions that when executed by a processor cause a computer to execute operations comprising:

generating intermediate representation information of a subband signal based on an acoustic feature value of a speech waveform;

simultaneously generating a plurality of subband signals corresponding to a plurality of different times and a plurality of different bands by inputting the intermediate representation information to a plurality of probability distribution generation models, wherein the plurality of probability distribution generation model outputs information on subband signals corresponding to times and bands allocated respectively, and the plurality of probability distribution generation models corresponds to a number of channels of the subband signals and a number of samples to be simultaneously generated; and

generating the speech waveform based on the plurality of subband signals.

9 . The computer-readable non-transitory recording medium according to claim 8 , the computer-executable program instructions when executed further causing the computer to execute operations comprising:

converting the acoustic feature value of the speech waveform into the intermediate representation information of the acoustic feature value by using a first intermediate representation model, wherein the first intermediate representation model that outputs the intermediate representation information of the acoustic feature value in a case where the acoustic feature value is input.

10 . The computer-readable non-transitory recording medium according to claim 9 , wherein the generating intermediate representation information further comprises generating the intermediate representation information of the subband signal using a second intermediate representation model, wherein the second intermediate representation model outputs the intermediate representation information of the subband signal in a case where the intermediate representation information of the acoustic feature value is input.

11 . The computer-readable non-transitory recording medium according to claim 10 , the computer-executable program instructions when executed further causing the computer to execute operations comprising:

calculating a loss value based on the plurality of subband signals, wherein the plurality of subband signals is calculated from the speech waveform; and

executing learning of at least one model, wherein the at least one model is among the first intermediate representation model, the second intermediate representation model, and the plurality of probability distribution generation models based on the loss value.

12 . The computer-readable non-transitory recording medium according to claim 9 , wherein the first intermediate representation model includes a neural network.

13 . The computer-readable non-transitory recording medium according to claim 8 , wherein the simultaneously generating a plurality of subband signals further comprises generating the plurality of subband signals using a simultaneous probability distribution generation model, and the simultaneous probability distribution generation model simultaneously outputs information of subband signals corresponding to a plurality of time zones and a plurality of bands from one model.

14 . The computer-readable non-transitory recording medium according to claim 8 , wherein the intermediate representation information is in multi-dimensional vector form.

15 . A device comprising a processor configured to execute operations comprising:

generating intermediate representation information of a subband signal based on an acoustic feature value of a speech waveform;

generating a plurality of subband signals corresponding to a plurality of different times and a plurality of different bands by inputting the intermediate representation information to a plurality of probability distribution generation models, wherein the plurality of probability distribution generation models outputs information on subband signals corresponding to times and bands allocated respectively, and the plurality of probability distribution generation models corresponds to a number of channels of the subband signals and a number of samples to be simultaneously generated; and

generating the speech waveform based on the plurality of subband signals.

16 . The device according to claim 15 , further comprising:

converting the acoustic feature value of the speech waveform into the intermediate representation information of the acoustic feature value by using a first intermediate representation model, wherein the first intermediate representation model that outputs the intermediate representation information of the acoustic feature value in a case where the acoustic feature value is input.

17 . The device according to claim 16 , wherein the generating intermediate representation information further comprises generating the intermediate representation information of the subband signal using a second intermediate representation model, wherein the second intermediate representation model outputs the intermediate representation information of the subband signal in a case where the intermediate representation information of the acoustic feature value is input.

18 . The device according to claim 17 , further comprising:

calculating a loss value based on the plurality of subband signals, wherein the plurality of subband signals is calculated from the speech waveform; and

executing learning of at least one model, wherein the at least one model is among the first intermediate representation model, the second intermediate representation model, and the plurality of probability distribution generation models based on the loss value.

19 . The device according to claim 15 , wherein the simultaneously generating a plurality of subband signals further comprises generating the plurality of subband signals using a simultaneous probability distribution generation model, and the simultaneous probability distribution generation model simultaneously outputs information of subband signals corresponding to a plurality of time zones and a plurality of bands from one model.

20 . The device according to claim 15 , wherein the intermediate representation information is in multi-dimensional vector form.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0725 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2024
From: KANAGAWA, HIROKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 066011/0204 →
Continuity (1)
Related Publication 20240339105A1 · Oct 10, 2024
References Cited (22)
US 11790884B1 · Shakeri · 2023 [cited by examiner]
US 12288566B1 · Ganguly · 2025 [cited by examiner]
US 20030055640A1 · Burshtein · 2003 [cited by examiner]
US 20060085194A1 · Okutani · 2006 [cited by examiner]
US 20130238337A1 · Kamai · 2013 [cited by examiner]
US 20140163991A1 · Short · 2014 [cited by examiner]
US 20180075343A1 · van den Oord · 2018 [cited by examiner]
US 20180174570A1 · Tamura · 2018 [cited by examiner]
US 20190189115A1 · Hori · 2019 [cited by examiner]
US 20200401874A1 · Kalchbrenner · 2020 [cited by examiner]
US 20210014039A1 · Zhang · 2021 [cited by examiner]
US 20230169953A1 · Zhang · 2023 [cited by examiner]
US 20230395061A1 · Biadsy · 2023 [cited by examiner]
US 20250252958A1 · Chang · 2025 [cited by examiner]
JP 2019045856A · 2019 [cited by examiner]
WO 2018048934A1 · 2018 [cited by applicant]
WO 2019155054A1 · 2019 [cited by applicant]
Lajos Hanzo; F. Clare A. Somerville; Jason P. Woodward, “Speech Signals and Introduction to Speech Coding,” in Voice Compression and Communications: Principles and Applications for Fixed and Wireless Channels , IEEE, 20… [cited by examiner]
Kawahara et al. (1999) “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based FO extraction: Possible role of a repetitive structure in sounds” Speech … [cited by applicant]
Morise et al. (2016) “World: a vocoder-based high-quality speech synthesis system for real-time applications” IEICE Transactions on Information and Systems, vol. E99-D, No. 7, pp. 1877-1884. [cited by applicant]
Valin et al. (2019) “LPCNET: Improving Neural Speech Synthesis through Linear Prediction” ICASSP 2019, May 12, 2019, pp. 5891-5895. [cited by applicant]
Yu et al. (2020) “DurIAN: Duration Informed Attention Network for Speech Synthesis” Interspeech 2020, Oct. 25, 2020, pp. 2027-2031. [cited by applicant]