LATENT-SEGMENTATION INTONATION MODEL
The intonation model of the present technology disclosed herein assigns different words within a sentence to be prominent, analyzes multiple prominence possibilities (in some cases, all prominence possibilities), and learns parameters of the model using large amounts of data. Unlike previous systems, intonation patterns are discovered from data. Speech data is sub-segmented into words, the different segments are analyzed and used for learning, and a determination is made as to whether the segmentations predict pitch
1 . A method for performing speech synthesis, comprising:
receiving, by an application on a computing device, data for a collection of words;
marking one or more of the collection of words with prominence data by the application;
determining parameters based on the prominence data; and
generating by the application synthesized speech data based on the determined parameters.
2 . The method of claim 1 , further comprising marking, by the application on the computing device, one or more syllables of the words with prominence data.
3 . The method of claim 2 , wherein the syllables are prominent syllables.
4 . The method of claim 1 , further comprising assigning one or more of the parameters to a word of the collection of words.
5 . The method of claim 1 , wherein the computing device includes a mobile device, the application including a mobile application in communication with remote server.
6 . The method of claim 1 , wherein the computing device includes a server, the server in communication with a mobile device.
7 . A non-transitory computer readable medium for performing speech synthesis, comprising:
receiving, by an application on a computing device, data for a collection of words;
marking one or more of the collection of words with prominence data by the application;
determining parameters based on the prominence data; and
generating by the application synthesized speech data based on the determined parameters.
8 . The non-transitory computer readable medium of claim 7 , further comprising marking, by the application on the computing device, one or more syllables of the words with prominence data.
9 . The non-transitory computer readable medium of of claim 8 , wherein the syllables are prominent syllables.
10 . The non-transitory computer readable medium of claim 7 , further comprising assigning one or more of the parameters to a word of the collection of words.
11 . The non-transitory computer readable medium of claim 7 , wherein the computing device includes a mobile device, the application including a mobile application in communication with remote server.
12 . The non-transitory computer readable medium of claim 7 , wherein the computing device includes a server, the server in communication with a mobile device.