Method and system for a parametric speech synthesis
Embodiments of the present systems and methods may provide techniques for synthesizing speech in any voice in any language in any accent. For example, in an embodiment, a text-to-speech conversion system may comprise a text converter adapted to convert input text to at least one phoneme selected from a plurality of phonemes stored in memory, a machine-learning model storing voice patterns for a plurality of individuals and adapted to receive the at least one phoneme and an identity of a speaker and to generate acoustic features for each phoneme, and a decoder adapted to receive the generated acoustic features and to generate a speech signal simulating a voice of the identified speaker in a language.
1. A speech conversion system comprising;
a machine-learning model storing voice patterns for a plurality of individuals and adapted to receive at least one phoneme and an identity of a speaker and to generate and enhance acoustic features for each phoneme, wherein the voice patterns comprise a matrix of components equal to a number of components of a production approach to speech synthesis times a number of components of an acoustic approach to speech synthesis, and wherein the enhanced acoustic features comprise at least one of spectral enhancement or focal enhancement; and
a decoder adapted to receive the generated acoustic features and to generate a speech signal simulating a voice of the identified speaker in a language.
2. The system of claim 1 , wherein the at least one phenome comprises a phoneme of the International Phonetic Alphabet and silence and breath.
3. The system of claim 1 , wherein the machine-learning model comprises a neural network model.
4. The system of claim 1 , wherein spectral enhancement comprises increasing a peak of a spectral envelope or decreasing a trough of the spectral envelope and focal enhancement comprises emphasizing the difference between a first frame and a second frame.
5. The system of claim 1 , wherein the generated acoustic features include accent acoustic features and the generated speech signal further simulates a voice of the identified speaker in a language and in an accent.
6. The system of claim 5 , wherein the accent corresponds to a native accent of the identified speaker.
7. The system of claim 1 , wherein the production approach comprises phones, coarticulation, prosody from linguistic features, and prosody from extra-linguistic features and the acoustic approach comprises spectrum, power, duration, and pitch.
8. A conversion method implemented in a computer system comprising a processor, memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor, the method comprising:
storing, in a machine-learning model at the computer system, voice patterns for a plurality of individuals, receiving, at the machine-learning model at the computer system, at least one phoneme and an identity of a speaker, and generating and enhancing, with the machine-learning model at the computer system, acoustic features for each phoneme, wherein the voice patterns comprise a matrix of components equal to a number of components of a production approach to speech synthesis times a number of components of an acoustic approach to speech synthesis, and wherein enhancing acoustic features comprises at least one of spectral enhancement or focal enhancement; and
receiving, at the computer system, the generated acoustic features and generating, at the computer system, a speech signal simulating a voice of the identified speaker in a language.
9. The method of claim 8 , wherein the at least one phenome comprises a phoneme of the International Phonetic Alphabet and silence and breath.
10. The method of claim 8 , wherein the machine-learning model comprises a neural network model.
11. The method of claim 8 , wherein spectral enhancement comprises increasing a peak of a spectral envelope or decreasing a trough of the spectral envelope and focal enhancement comprises emphasizing the difference between a first frame and a second frame.
12. The method of claim 8 , wherein the generated acoustic features include accent acoustic features and the generated speech signal further simulates a voice of the identified speaker in a language and in an accent.
13. The method of claim 12 , wherein the accent corresponds to a native accent of the identified speaker.
14. The method of claim 8 , wherein the production approach comprises phones, coarticulation, prosody from linguistic features, and prosody from extra-linguistic features and the acoustic approach comprises spectrum, power, duration, and pitch.