IP Library Granted Patent US 8,731,933
Granted Patent B2
US 8,731,933 · App. 13/860,319 · Granted May 20, 2014

Speech synthesis apparatus and method utilizing acquisition of at least two speech unit waveforms acquired from a continuous memory region by one access

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,731,933
App. No.
13/860,319
Granted
May 20, 2014
Kind
B2
Abstract

A speech synthesizing apparatus includes a selector configured to select a plurality of speech units for synthesizing a speech of a phoneme sequence by referring to speech unit information stored in an information memory. Speech unit waveforms corresponding to the speech units are acquired from a plurality of speech unit waveforms stored in a waveform memory, and the speech is synthesized by utilizing the speech unit waveforms acquired. When acquiring the speech unit waveforms, at least two speech unit waveforms from a continuous region of the waveform memory are copied onto a buffer by one access, wherein a data quantity of the at least two speech unit waveforms is less than or equal to a size of the buffer.

Claims (51)

1. An apparatus for synthesizing a speech of a phoneme sequence, comprising:

a selector configured to select a plurality of speech units for synthesizing the speech of the phoneme sequence by referring to speech unit information stored in an information memory;

an acquisition unit configured to acquire a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory; and

a synthesizing unit configured to synthesize the speech by utilizing the speech unit waveform acquired by the acquisition unit;

wherein the acquisition unit copies at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units.

2. The apparatus according to claim 1 , further comprising:

a text input unit by which text to be converted to the phoneme sequence is input.

3. The apparatus according to claim 1 , further comprising:

at least one of a speaker and a head phone by which the speech is output.

4. The apparatus according to claim 1 , further comprising:

a computer operation system configured to execute copying by said acquisition unit of said at least two speech unit waveforms.

5. The apparatus according to claim 1 , further comprising:

a middle-ware (MW) program unit configured to execute copying by said acquisition unit of said at least two speech unit waveforms.

6. The apparatus according to claim 1 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.

7. The apparatus according to claim 1 , wherein

an access speed of the information memory is quicker than an access speed of the waveform memory.

8. The apparatus according to claim 1 , wherein

the synthesizing unit synthesizes the speech by concatenating the speech unit waveform with modification or without modification.

9. A method for synthesizing a speech of an input a phoneme sequence, comprising:

selecting a plurality of speech units for synthesizing the speech of the phoneme sequence by referring to speech unit information stored in an information memory;

acquiring by an acquisition unit a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory, including

copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units; and

synthesizing the speech by utilizing the speech unit waveform acquired by the acquisition unit.

10. The method according to claim 9 , wherein

the copying is executed by a computer operation system.

11. The method according to claim 9 , wherein

the copying is executed by a middle-ware (MW) program unit.

12. The method according to claim 9 , further comprising:

outputting the speech by at least one of a speaker and a head phone.

13. The method according to claim 9 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.

14. The method according to claim 9 , wherein

the synthesizing includes

synthesizing the speech by concatenating the speech unit waveform with modification or without modification.

15. A non-transitory computer readable memory storing program instructions, which when executed by a computer, results in performance of a method for synthesizing a speech of a phoneme sequence, comprising steps of:

selecting a plurality of speech units for synthesizing the speech of the input phoneme sequence by referring to speech unit information stored in an information memory;

acquiring a speech unit waveform corresponding to each speech unit of the plurality of speech units from a plurality of speech unit waveforms stored in a waveform memory, including

copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the speech unit waveforms being less than or equal to a size of the buffer, the speech unit waveforms corresponding to at least two speech units included in the plurality of speech units; and

synthesizing the speech by utilizing the speech unit waveform being acquired.

16. The non-transitory computer readable medium according to claim 15 , wherein

the copying is executed by a computer operation system.

17. The non-transitory computer readable medium according to claim 15 , wherein

the copying is executed by a middle-ware (MW) program unit.

18. The non-transitory computer readable medium according to claim 15 , further comprising:

outputting the speech by at least one of a speaker and a head phone.

19. The non-transitory computer readable medium according to claim 15 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, and a pitch waveform.

20. The non-transitory computer readable medium according to claim 15 , wherein

the synthesizing includes

synthesizing the speech by concatenating the speech unit waveform with modification or without modification.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →