IP Library › Granted Patent US 12,620,408
Granted Patent B2
US 12,620,408 · App. 18/519,986 · Granted May 5, 2026

Generating audio using neural networks

Inventors: Aaron Gerard Antonius van den Oord (London, GB); Sander Etienne Lea Dieleman (London, GB); Nal Emmerich Kalchbrenner (Amsterdam, NL); Karen Simonyan (London, GB); Oriol Vinyals (London, GB)
Assignee: GDM Holding LLC
G10L25/30G06N3/045G06N3/0464G06N3/048G10L13/06G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,408
App. No.
18/519,986
Granted
May 5, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output sequence of audio data that comprises a respective audio sample at each of a plurality of time steps. One of the methods includes, for each of the time steps: providing a current sequence of audio data as input to a convolutional subnetwork, wherein the current sequence comprises the respective audio sample at each time step that precedes the time step in the output sequence, and wherein the convolutional subnetwork is configured to process the current sequence of audio data to generate an alternative representation for the time step; and providing the alternative representation for the time step as input to an output layer, wherein the output layer is configured to: process the alternative representation to generate an output that defines a score distribution over a plurality of possible audio samples for the time step.

Claims (62)

1 . A method performed by one or more computers, the method comprising:

receiving data characterizing a sequence of text;

processing a network input comprising the data characterizing the sequence of text using a neural network to generate a neural network output that defines audio data that is a verbalization of the sequence of text characterized by the network input, wherein:

the neural network comprises a sequence of one or more convolutional neural network layers; and

each convolutional neural network layer in the sequence of one or more convolutional neural network layers is configured to:

receive a layer input;

apply one or more convolution operations to the layer input; and

generate a layer output based at least in part on a result of applying the one or more convolution operations to the layer input;

wherein the neural network output directly comprises amplitude values of an audio waveform that is the verbalization of the sequence of text characterized by the network input,

wherein the amplitude values of the audio waveform are directly generated as an output of an output neural network layer of the neural network, and

wherein the neural network has been trained by applying a supervised learning technique that depends on (i) ground truth output audio waveforms for each of a set of training examples for the neural network and (ii) corresponding output audio waveforms generated by the neural network.

2 . The method of claim 1 , wherein the data characterizing the sequence of text comprises a sequence of phonemes corresponding to the sequence of text.

3 . The method of claim 1 , wherein for each of one or more convolutional neural network layers in the sequence of one or more convolutional neural network layers, applying one or more convolution operations to the layer input comprises:

applying one or more dilated convolution operations to the layer input.

4 . The method of claim 3 , wherein the sequence of one or more convolutional neural network layers comprises a plurality of convolutional neural network layers that each implement respective dilated convolution operations associated with a respective different dilation rate.

5 . The method of claim 4 , wherein the sequence of convolutional neural network layers comprises a plurality of convolutional neural network layers that each have a dilation rate that is a constant multiple of a dilation rate of a preceding convolutional neural network layer.

6 . The method of claim 1 , wherein for each of one or more convolutional neural network layers in the sequence of one or more convolutional neural network layers, the convolutional neural network layer comprises a gated activation unit that is configured to:

process the layer input by a main convolutional operation to generate a main convolutional output;

process the layer input by a gate convolutional operation to generate a gating convolutional output; and

generate a gated activation unit output by element-wise multiplying the main convolutional output and the gating convolutional output.

7 . The method of claim 1 , wherein the sequence of one or more convolutional neural network layers comprises one or more residual connections, wherein each residual connection is configured to route the layer input to the convolutional neural network layer to a summer that sums the layer input with an intermediate output generated by the convolutional neural network layer, wherein the layer output is based at least in part on the sum of the layer input with the intermediate output.

8 . The method of claim 1 , wherein the sequence of one or more convolutional neural network layers comprises a plurality of convolutional neural network layers.

9 . The method of claim 1 , wherein the neural network is conditioned on speaker identity data; and

wherein the audio data that is the verbalization of the sequence of text is expressed in a voice associated with the speaker identity data.

10 . The method of claim 1 , further comprising:

evaluating an objective function that measures an error in the audio waveform; and

backpropagating gradients of the objective function through the neural network.

11 . The method of claim 1 , wherein:

the network input comprises data characterizing a conditioning input; and

wherein processing the network input comprising the data characterizing the sequence of text using the neural network to generate the neural network output that defines audio data that is the verbalization of the sequence of text characterized by the network input comprises:

processing the network input comprising the data characterizing the sequence of text using the neural network and the data characterizing the conditioning input to generate the neural network output that defines audio data that is the verbalization of the sequence of text characterized by the network input, as conditioned on the conditioning input.

12 . The method of claim 11 , wherein the conditioning input comprises image data.

13 . The method of claim 11 , wherein the conditioning input comprises video data.

14 . The method of claim 11 , wherein the conditioning input comprises data characterizing a particular speaker for the verbalization of the sequence of text.

15 . The method of claim 11 , wherein the conditioning input comprises data characterizing a particular language for the verbalization of the sequence of text.

16 . The method of claim 11 , wherein the conditioning input comprises data characterizing particular music for the verbalization of the sequence of text.

17 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving data characterizing a sequence of text;

processing a network input comprising the data characterizing the sequence of text using a neural network to generate a neural network output that defines audio data that is a verbalization of the sequence of text characterized by the network input, wherein:

the neural network comprises a sequence of one or more convolutional neural network layers; and

each convolutional neural network layer in the sequence of one or more convolutional neural network layers is configured to:

receive a layer input;

apply one or more convolution operations to the layer input; and

generate a layer output based at least in part on a result of applying the one or more convolution operations to the layer input;

wherein the neural network output directly comprises amplitude values of an audio waveform that is the verbalization of the sequence of text characterized by the network input, and

wherein the amplitude values of the audio waveform are directly generated as an output of an output neural network layer of the neural network, and

wherein the neural network has been trained by applying a supervised learning technique that depends on (i) ground truth output audio waveforms for each of a set of training examples for the neural network and (ii) corresponding output audio waveforms generated by the neural network.

18 . The system of claim 17 , wherein the data characterizing the sequence of text comprises a sequence of phonemes corresponding to the sequence of text.

19 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving data characterizing a sequence of text;

processing a network input comprising the data characterizing the sequence of text using a neural network to generate a neural network output that defines audio data that is a verbalization of the sequence of text characterized by the network input, wherein:

the neural network comprises a sequence of one or more convolutional neural network layers; and

each convolutional neural network layer in the sequence of one or more convolutional neural network layers is configured to:

receive a layer input;

apply one or more convolution operations to the layer input; and

generate a layer output based at least in part on a result of applying the one or more convolution operations to the layer input;

wherein the neural network output directly comprises amplitude values of an audio waveform that is the verbalization of the sequence of text characterized by the network input,

wherein the amplitude values of the audio waveform are directly generated as an output of an output neural network layer of the neural network, and

wherein the neural network has been trained by applying a supervised learning technique that depends on (i) ground truth output audio waveforms for each of a set of training examples for the neural network and (ii) corresponding output audio waveforms generated by the neural network.

20 . The non-transitory computer storage media of claim 19 , wherein the data characterizing the sequence of text comprises a sequence of phonemes corresponding to the sequence of text.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2023
From: VAN DEN OORD, AARON GERARD ANTONIUS; DIELEMAN, SANDER ETIENNE LEA; KALCHBRENNER, NAL EMMERICH; SIMONYAN, KAREN; VINYALS, ORIOL
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 065936/0685 →
Continuity (7)
Continuation 17838985 · Jun 13, 2022
Continuation 17020348 · Sep 14, 2020
Continuation 16390549 · Apr 22, 2019
Continuation 16030742 · Jul 9, 2018
Continuation PCTUS2017050320 · Sep 6, 2017
Provisional Application 62384115 · Sep 6, 2016
Related Publication 20240135955A1 · Apr 25, 2024
References Cited (174)
US 2810457A · Halliday · 1957 [cited by applicant]
US 5377302A · Tsiang · 1994 [cited by applicant]
US 5668926A · Karaali et al. · 1997 [cited by applicant]
US 5913194A · Karaali et al. · 1999 [cited by applicant]
US 7409340B2 · Holzapfel · 2008 [cited by applicant]
US 8527276B1 · Senior · 2013 [cited by applicant]
US 8645137B2 · Bellegarda et al. · 2014 [cited by applicant]
US 9058811B2 · Wang · 2015 [cited by applicant]
US 9190053B2 · Penn · 2015 [cited by applicant]
US 9595002B2 · Leeman-Munk et al. · 2017 [cited by applicant]
US 9734824B2 · Penn · 2017 [cited by applicant]
US 9824684B2 · Yu et al. · 2017 [cited by applicant]
US 9953634B1 · Pearce · 2018 [cited by applicant]
US 9972314B2 · Wang et al. · 2018 [cited by applicant]
US 10043512B2 · Jaitly et al. · 2018 [cited by applicant]
US 10049106B2 · Goyal · 2018 [cited by applicant]
US 10304477B2 · Van den Oord et al. · 2019 [cited by applicant]
US 10354015B2 · Kalchbrenner et al. · 2019 [cited by applicant]
US 10403269B2 · Sainath et al. · 2019 [cited by applicant]
US 10460747B2 · Roblek · 2019 [cited by applicant]
US 10714077B2 · Song et al. · 2020 [cited by applicant]
US 10803884B2 · van den Oord · 2020 [cited by applicant]
US 10977529B2 · Szegedy et al. · 2021 [cited by applicant]
US 11080591B2 · Van Den Oord et al. · 2021 [cited by applicant]
US 11386914B2 · van den Oord et al. · 2022 [cited by applicant]
US 20020011024A1 · Baldwin et al. · 2002 [cited by applicant]
US 20020110248A1 · Kovales · 2002 [cited by applicant]
US 20060064177A1 · Tian et al. · 2006 [cited by applicant]
US 20090210218A1 · Collobert et al. · 2009 [cited by applicant]
US 20120166198A1 · Lin · 2012 [cited by applicant]
US 20120323521A1 · De Foras et al. · 2012 [cited by applicant]
US 20150032449A1 · Sainath et al. · 2015 [cited by applicant]
US 20150356075A1 · Rao et al. · 2015 [cited by applicant]
US 20160023244A1 · Zhuang et al. · 2016 [cited by applicant]
US 20160035344A1 · Gonzalez-Dominguez et al. · 2016 [cited by applicant]
US 20160063359A1 · Szegedy et al. · 2016 [cited by applicant]
US 20160093278A1 · Esparza · 2016 [cited by applicant]
US 20160099010A1 · Sainath et al. · 2016 [cited by applicant]
US 20160140951A1 · Agionnyrgiannakis · 2016 [cited by applicant]
US 20160180162A1 · Cetintas et al. · 2016 [cited by applicant]
US 20160232440A1 · Gregor · 2016 [cited by applicant]
US 20160343366A1 · Fructuoso et al. · 2016 [cited by applicant]
US 20170011738A1 · Senior · 2017 [cited by applicant]
US 20170103752A1 · Senior · 2017 [cited by applicant]
US 20170148431A1 · Catanzaro et al. · 2017 [cited by applicant]
US 20170262737A1 · Rabinovich · 2017 [cited by applicant]
US 20170330586A1 · Roblek et al. · 2017 [cited by applicant]
US 20180025257A1 · van den Oord et al. · 2018 [cited by applicant]
US 20180025721A1 · Li · 2018 [cited by applicant]
US 20180068207A1 · Szegedy et al. · 2018 [cited by applicant]
US 20180075343A1 · Van den Oord et al. · 2018 [cited by applicant]
US 20180322891A1 · Van den Oord et al. · 2018 [cited by applicant]
US 20180329897A1 · Kalchbrenner et al. · 2018 [cited by applicant]
US 20180365554A1 · Van den Oord et al. · 2018 [cited by applicant]
US 20190043516A1 · Germain et al. · 2019 [cited by applicant]
US 20190066713A1 · Mesgarani et al. · 2019 [cited by applicant]
US 20190378498A1 · Sainath et al. · 2019 [cited by applicant]
US 20200051583A1 · Wu · 2020 [cited by applicant]
US 20200342183A1 · Kalchbrenner et al. · 2020 [cited by applicant]
CA 2810457 · 2014 [cited by applicant]
CN 102047321 · 2011 [cited by applicant]
CN 104681034 · 2015 [cited by applicant]
CN 105068998 · 2015 [cited by applicant]
CN 105096939 · 2015 [cited by applicant]
CN 105144164 · 2015 [cited by applicant]
CN 105159890 · 2015 [cited by applicant]
CN 105210064 · 2015 [cited by applicant]
CN 105321525 · 2016 [cited by applicant]
CN 105513591 · 2016 [cited by applicant]
JP 1991071196 · 1991 [cited by applicant]
JP H08512150 · 1996 [cited by applicant]
JP H10333699 · 1998 [cited by applicant]
JP 1999282484 · 1999 [cited by applicant]
JP 2002123280 · 2002 [cited by applicant]
JP 6750121 · 2020 [cited by applicant]
KR 1020080026951 · 2008 [cited by applicant]
KR 1020080042083 · 2008 [cited by applicant]
KR 1020140056368 · 2014 [cited by applicant]
KR 1020150104111 · 2015 [cited by applicant]
KR 1020160013710 · 2016 [cited by applicant]
KR 101855597 · 2018 [cited by applicant]
WO WO2018048945 · 2018 [cited by applicant]
Paszke et al. “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation”. arXiv: 1606.02147v1 [cs.CV] Jun. 7, 2016 (Year: 2016). [cited by examiner]
Tokuda et al. “Directly Modeling Speech Waveforms by Neural Networks for Statistical Parametric Speech Synthesis”. ICASSP 2015 (Year: 2015). [cited by examiner]
Keren et al. “Convolutional RNN: An enhanced model for extracting features from sequential data”. 2016 International Joint Conference on Neural Networks, 24-29, Jul. 2016, Vancouver, BC, Canada, pp. 3412-3419 (Year: 201… [cited by examiner]
Veit et al. “Residual Networks are Exponential Ensembles of Relatively Shallow Networks”. arXiv:1605.06431v1 [cs.CV] May 20, 2016 (Year: 2016). [cited by examiner]
Graves, Alex. “Generating Sequences with Recurrent Neural Networks”. arXiv:1308.0850v5 [cs.NE] Jun. 5, 2014 (Year: 2014). [cited by examiner]
Agiomyrgiannakis, “Vocain the vocoder and applications is speech synthesis,” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, 5 pages. [cited by applicant]
Allowance of Patent in Korean Appln. No. 10-2019-7008257, dated Apr. 18, 2022, 3 pages (with English translation). [cited by applicant]
Allowance of Patent in Korean Appln. No. 10-2019-7013231, dated Jan. 8, 2022, 3 pages (with English translation). [cited by applicant]
Allowance of Patent in Korean Appln. No. 10-2022-7003520, dated Aug. 31, 2022, 4 pages (with English translation). [cited by applicant]
Bahdanau et al., “Neural machine translation by jointly learning to align and translation,” arXiv 1409.0473v7, May 19, 2016, 15 pages. [cited by applicant]
Berglund et al., “Bidirectional Recurrent Neural Networks as Generative Models,” Advances in Neural Information Processing Systems 28 (NIPS 2015), 2015, 9 pages. [cited by applicant]
Bishop, “Mixture density networks,” Technical Report NCRG/94/004, Neural Computing Research Group, Aston University, 1994, 26 pages. [cited by applicant]
Chen et al., “Semantic image segmentation with deep convolutional nets and fully connected CRF's,” arXiv 1412.7062v4, Jun. 7, 2016, 14 pages. [cited by applicant]
CN Office Action in Chinese Appln. No. 201780065523.6, dated Jan. 6, 2020, 14 pages (with English translation). [cited by applicant]
Decision to Grant a Patent in Japanese Appln. No. 2019-150456, dated Apr. 26, 2021, 5 pages (with English translation). [cited by applicant]
Decision to Grant a Patent in Japanese Appln. No. 2021-087708, dated Dec. 19, 2022, 5 pages (with English translation). [cited by applicant]
Decision to Grant Patent in Japanese Appln. No. 2020-135790, dated Apr. 4, 2022, 5 pages (with English translation). [cited by applicant]
EP Extended Search Report in European Appln. No. 20192441.2, dated Dec. 23, 2020, 8 pages. [cited by applicant]
EP Extended Search Report in European Appln. No. 20195353.6, dated Apr. 21, 2021, 7 pages. [cited by applicant]
EP Office Action in European Appln. 17794596.1, dated Jun. 5, 2019, 3 pages. [cited by applicant]
Fan et al., “TTS synthesis with bidirectional LSTM based recurrent neural networks,” Fifteenth Annual Conference of the International Speech Communication Association, 2014, 5 pages. [cited by applicant]
Fant et al., “TTS synthesis with bidirectional LSTM based recurrent neural networks,” Interspeech, Sep. 2014, 5 pages. [cited by applicant]
Fisher et al., “WaveMedic: Convolutional Neural Networks for Speech Audio Enhancement,” Stanford University, Sep. 2020, 6 pages. [cited by applicant]
Gonzalvo et al., “Recent advances in Google real-time HMM-driven unit selection synthesizer,” In Proc. Interspeech, Sep. 2016, 5 pages. [cited by applicant]
He et al., “Deep residual learning for image recognition,” arXiv 1512.03385v1, Dec. 10, 2015, 12 pages. [cited by applicant]
Hertel et al., “Comparing Time and Frequency Domain for Audio Event Recognition Using Deep Learning,” 2016 International Joint Conference on Neural Networks, Mar. 2016, 5 pages. [cited by applicant]
Hochreiter et al., “Long short-term memory,” Neural Computation 9(8), Nov. 1997, 46 pages. [cited by applicant]
Homepages.inf.ed.ac.uk [online] “CSTR VCTK Corpus English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit,” available on or before Mar. 13, 2013, via Internet Archive Wayback Machine at URL<https://web.archive.org/w… [cited by applicant]
Hoshen et al., “Speech acoustic modeling from raw multichannel waveforms,” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, 5 pages. [cited by applicant]
Hunt et al., “Unit selection in a concatenative speech synthesis system using a large speech database,” IEEE International Conference on Acoustics, Speech and Signal Processing, May 7, 1996, 4 pages. [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2017/050320, mailed on Dec. 14, 2018, 7 pages. [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2017/050335, mailed on Dec. 14, 2018, 7 pages. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2017/050320, mailed on Jan. 2, 2018, 16 pages. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2017/050335, mailed on Jan. 2, 2018, 16 pages. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2017/058046, mailed on Jan. 31, 2018, 15 pages. [cited by applicant]
Jozefowicz et al., “Exploring the limits of language modeling,” arXiv 1602.02410v2 Feb. 11, 2016, 11 pages. [cited by applicant]
Kalchbrenner et al., “Neural Machine Translation in Linear Time,” CoRR, Oct. 2016, arxiv.org/abs/1610.10099, 9 pages. [cited by applicant]
Kalchbrenner et al., “Video pixel networks,” arXiv 1610.00527v1, Oct. 3, 2016, 16 pages. [cited by applicant]
Karaali, “text-to-speech conversion with neural networks: A recurrent TDNN approach,” cs.NE/9811032, Nov. 24, 1998, 4 pages. [cited by applicant]
Kawahara et al., “Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system straight,” Second International Worksh… [cited by applicant]
Kawahara et al., “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: possible role of a repetitive structure in sounds,” Speech Comm.… [cited by applicant]
Lamb et al., “Convolutional encoders for neural machine translation,” Project reports of 2015 CS224d course at Stanford university, Jun. 22, 2015, 8 pages. [cited by applicant]
Law et al., “Input-agreement: a new mechanism for collecting data using human computation games,” Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, ACM, Apr. 4, 2009, 10 pages. [cited by applicant]
Lee et al., “Fully character-level neural machine translation without explicit segmentation,” arXiv 1610.03017v1, Oct. 10, 2016, 14 pages. [cited by applicant]
Maia et al., “Statistical parametric speech synthesis with joint estimation of acoustic and excitation model parameters,” ISCA SSW7, Sep. 24, 2010, 6 pages. [cited by applicant]
Meng et al., “Encoding source Language with convolutional neural network for machine translation,” Proceedings of the 53rd annual meeting of the Association for Computational Linguistics and the 7th International Joint … [cited by applicant]
Morise et al., “World: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Tranc. Inf. Syst., Jul. 1, 2016, E99-D(7):8 pages. [cited by applicant]
Moulines et al., “Pitch synchronous waveform processing techniques for text-to-speech synthesis using diphones,” Speech Comm., Dec. 1990, (9):15 pages. [cited by applicant]
Muthukumar et al., “A deep learning approach to data-driven parameterizations for statistical parametric speech synthesis,” arXiv 1409.8558v1, Sep. 30, 2014, 5 pages. [cited by applicant]
Nair et al., “Rectified linear units improve restricted Boltzmann machines,” Proceedings of the 37th International Conference on Machine Learning, Jun. 21-24, 2010, 8 pages. [cited by applicant]
Nakamura et al., “Integration of spectral feature extraction and modeling for HMM-based speech synthesis,” IFICE Transaction on Information and Systems 97(6), Jun. 2014, 11 pages. [cited by applicant]
Notice of Allowance in Chinese Appln. No. 202011082855.5, dated Jan. 10, 2024, 8 pages (with English translation). [cited by applicant]
Office Action in Brazilian Appln. No. BR112019004524-4, dated Jun. 23, 2023, 6 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201780065178.6, dated Mar. 31, 2023, 8 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201780065178.6, dated Sep. 13, 2022, 16 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201780073530.0, dated Sep. 2, 2022, 14 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 17794596.1, dated Oct. 8, 2021, 11 pages. [cited by applicant]
Office Action in Indian Appln. No. 201927011194, dated Mar. 10, 2021, 7 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2019-150456, dated Nov. 30, 2020, 4 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2020-135790, dated Sep. 6, 2021, 8 pages (with English translation). [cited by applicant]
Office Action in Japanese AppIn, No. 2021-087708, dated Jun. 27, 2022, 4 pages (with English translation). [cited by applicant]
Palaz et al., “Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks,” arXiv 1304.1018v2, Jun. 12, 2013, 5 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Appln. No. PCT/US2017/058046, dated May 9, 2019, 9 pages. [cited by applicant]
Peltonen et al., “Nonlinear filter design: methodologies and challenges,” ISPA Proceedings of the 2nd International Symposium on Image Signal Processing and Analysis, Jun. 2001, 6 pages. [cited by applicant]
Sagisaka et al., “ATR v-talk speech synthesis system,” Second International Conference on Spoken Language Processing, Oct. 1992, 4 pages. [cited by applicant]
Sainath et al., “Learning the speech front-end with raw waveform CLDNNs,” Sixteenth Annual Conference of the International Speech Communication Association, Sep. 2015, 5 pages. [cited by applicant]
Takaki et al., “A deep auto-encoder based low dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis,” IEEE International Conference on Acoustics Speech and Signal Process… [cited by applicant]
Theis et al., “Generative image modeling using spatial LSTMs,” Advances in Neural Information Processing Systems, Dec. 2015, 9 pages. [cited by applicant]
Toda et al., “A speech parameter generation algorithm generation algorithm considering global variance for HMM-based speech synthesis,” IEICE Trans. Inf. Syst., May 1, 2007, 90(5):9 pages. [cited by applicant]
Toda et al., “Statistical approach to vocal tract transfer function estimation based on factor analyzed trajectory hmm,” IEEE International Conference on Acoustics, Speech and Signal Processing, Mar. 31, 2008, 4 pages. [cited by applicant]
Tokuda et al., “Directly modeling speech waveforms by neural networks for statistical parametric speech synthesis,” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, 5 pages. [cited by applicant]
Tokuda et al., “Directly modeling voiced and unvoiced components in speech waveforms by neural networks,” IEEE International Conference on Acoustics, Speech and Signal Processing, Mar. 20, 2016, 5 pages. [cited by applicant]
Tuerk et al., “Speech synthesis using artificial neural networks trained on cepstral coefficients,” Proc. Eurospeech, Sep. 1993, 4 pages. [cited by applicant]
Tuske et al., “Acoustic modeling with deep neural networks using raw time signal for LVCSR,” Fifteenth Annual Conference of the International Speech Communication Association, Sep. 2014, 5 pages. [cited by applicant]
Uria et al., “Modeling acoustic feature dependencies with artificial neural networks: Trajectory-RNADE,” Proceedings of the International Conference on Acoustics Speech and Signal Processing, Apr. 19, 2015, 5 pages. [cited by applicant]
Van den Oord et al., “Conditional image generation with pixelenn decoders,” arXiv 1606.05328v2, Jun. 18, 2016, 13 pages. [cited by applicant]
Van den Oord et al., “Pixel recurrent neural networks,” arXiv 1601.06759v3, Aug. 19, 2016, 11 pages. [cited by applicant]
Van den Oord et al., “WaveNet: A Generative Model for Raw Audio,” arXiv 1609.03499v2, Sep. 19, 2016, 15 pages. [cited by applicant]
Written Opinion issued in International Application No. PCT/US2017/050320, mailed on Aug. 3, 2018, 7 pages. [cited by applicant]
Written Opinion issued in International Application No. PCT/US2017/050335, mailed on Aug. 3, 2018, 7 pages. [cited by applicant]
Wu et al., “Minimum generation error training with direct log spectral distortion on LSPs for HMM-based speech synthesis,” Interspeech, Sep. 2008, 4 pages. [cited by applicant]
www.itu.int [online] “ITU-T Recommendation G 711, Pulse Code Modulation (PCM) of voice frequencies,” Nov. 1988, retrieved on Jul. 9, 2017, retrieved from URL<https://www.itu.int/rec/T-REC-G.711-198811-I/en> 1 page. [cited by applicant]
www.sp.nitech.ac [online], “Speech synthesis as a statistical machine learning problem,” Dec. 2011, retrieved on Jul. 9, 2018, retrieved from Internet: URL<http://www.sp.nitech.ac.jp/˜tokuda/tokuda_asru2011_for_pdf.pdf,… [cited by applicant]
www.sp.nitech.ac.jp [online] “Speech Synthesis as A Statistical Machine Learning Problem,” Dec. 14, 2011, retrieved on Feb. 19, 2021, retrieved from URL <http://hts.sp.nitech.as.jp/>, 66 pages. [cited by applicant]
Yoshimura, “Simultaneous modeling of phonetic and prosodic parameters and characteristic conversion for HMM-based text-to-speech systems,” PhD thesis, Nagoya Institute of Technology, Jan. 2002, 109 pages. [cited by applicant]
Yu et al., “Multi-scale context aggregation by dilated convolutions,” arXiv 1511.07122, Nov. Apr. 30, 2016, 13 pages. [cited by applicant]
Zen et al., “Fast, Compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices,” arXiv 1606.06061v2, Jun. 22, 2016, 14 pages. [cited by applicant]
Zen et al., “Statistical parametric speech synthesis using deep neural networks,” IEEE International Conference on Acoustics, Speech and Signal Processing, May 26, 2013, 5 pages. [cited by applicant]
Zen et al., “Statistical parametric speech synthesis,” Speech Comm. 51(11) Nov. 30, 2009, 24 pages. [cited by applicant]
Zen, “An example of context-dependent label format for HMM-based speech synthesis in English,” The HTS CMUARCTIC demo 133, Mar. 2, 2006, 1 page. [cited by applicant]
Office Action in Canadian Appln. No. 3,155,320, dated Aug. 21, 2024, 4 pages. [cited by applicant]
Office Action in European Appln. No. 24188470.9, dated Jan. 21, 2025, 12 pages. [cited by applicant]