Speech synthesis apparatus and method utilizing acquisition of at least two speech unit waveforms acquired from a continuous memory region by one access
View Patent ↗A speech synthesizing apparatus includes a selector configured to select a plurality of speech units for synthesizing a speech of a phoneme sequence by referring to speech unit information stored in an information memory. Speech unit waveforms corresponding to the speech units are acquired from a plurality of speech unit waveforms stored in a waveform memory, and the speech is synthesized by utilizing the speech unit waveforms acquired. When acquiring the speech unit waveforms, at least two speech unit waveforms from a continuous region of the waveform memory are copied onto a buffer by one access, wherein a data quantity of the at least two speech unit waveforms is less than or equal to a size of the buffer.
1. An apparatus for synthesizing a speech of a phoneme sequence, comprising:
a selector configured to select a plurality of speech units for synthesizing the speech of the phoneme sequence by referring to speech unit information stored in an information memory;
an acquisition unit configured to acquire a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory; and
a synthesizing unit configured to synthesize the speech by utilizing the speech unit waveform acquired by the acquisition unit;
wherein the acquisition unit copies at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units.
2. The apparatus according to claim 1 , further comprising:
a text input unit by which text to be converted to the phoneme sequence is input.
3. The apparatus according to claim 1 , further comprising:
at least one of a speaker and a head phone by which the speech is output.
4. The apparatus according to claim 1 , further comprising:
a computer operation system configured to execute copying by said acquisition unit of said at least two speech unit waveforms.
5. The apparatus according to claim 1 , further comprising:
a middle-ware (MW) program unit configured to execute copying by said acquisition unit of said at least two speech unit waveforms.
6. The apparatus according to claim 1 , wherein
the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.
7. The apparatus according to claim 1 , wherein
an access speed of the information memory is quicker than an access speed of the waveform memory.
8. The apparatus according to claim 1 , wherein
the synthesizing unit synthesizes the speech by concatenating the speech unit waveform with modification or without modification.
9. A method for synthesizing a speech of an input a phoneme sequence, comprising:
selecting a plurality of speech units for synthesizing the speech of the phoneme sequence by referring to speech unit information stored in an information memory;
acquiring by an acquisition unit a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory, including
copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units; and
synthesizing the speech by utilizing the speech unit waveform acquired by the acquisition unit.
10. The method according to claim 9 , wherein
the copying is executed by a computer operation system.
11. The method according to claim 9 , wherein
the copying is executed by a middle-ware (MW) program unit.
12. The method according to claim 9 , further comprising:
outputting the speech by at least one of a speaker and a head phone.
13. The method according to claim 9 , wherein
the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.
14. The method according to claim 9 , wherein
the synthesizing includes
synthesizing the speech by concatenating the speech unit waveform with modification or without modification.
15. A non-transitory computer readable memory storing program instructions, which when executed by a computer, results in performance of a method for synthesizing a speech of a phoneme sequence, comprising steps of:
selecting a plurality of speech units for synthesizing the speech of the input phoneme sequence by referring to speech unit information stored in an information memory;
acquiring a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory, including
copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units; and
synthesizing the speech by utilizing the speech unit waveform being acquired.
16. The non-transitory computer readable medium according to claim 15 , wherein
the copying is executed by a computer operation system.
17. The non-transitory computer readable medium according to claim 15 , wherein
the copying is executed by a middle-ware (MW) program unit.
18. The non-transitory computer readable medium according to claim 15 , further comprising:
outputting the speech by at least one of a speaker and a head phone.
19. The non-transitory computer readable medium according to claim 15 , wherein
the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.
20. The non-transitory computer readable medium according to claim 15 , wherein
the synthesizing includes
synthesizing the speech by concatenating the speech unit waveform with modification or without modification.