IP Library Granted Patent US 9,666,179
Granted Patent B2
US 9,666,179 · App. 14/191,082 · Granted May 30, 2017

Speech synthesis apparatus and method utilizing acquisition of at least two speech unit waveforms acquired from a continuous memory region by one access

Inventor: Takehiko Kagoshima (Kanagawa-ken, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L13/02G10L13/04G10L13/047G10L13/06G10L13/08G10L15/063G11C7/16G10L13/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,666,179
App. No.
14/191,082
Granted
May 30, 2017
Kind
B2
Abstract

A waveform memory that stores a plurality of speech unit waveforms corresponding to respective speech units, wherein an address order of the speech unit waveforms is determined by a sort order of speech units included in a speech unit sequence corresponding to a phoneme sequence of training data, and the speech units included in the speech unit sequence are selected so as to synthesize a speech of the phone sequence.

Claims (51)

1. An apparatus for synthesizing a speech, comprising:

processing circuitry configured to

acquire, from a waveform memory, speech unit waveforms corresponding to a plurality of speech units for synthesizing the speech, and

synthesize the speech by utilizing the speech unit waveforms,

wherein the processing circuitry copies at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the at least two speech unit waveforms being less than or equal to a size of the buffer, the at least two speech unit waveforms corresponding to at least two speech units included in the plurality of speech units.

2. The apparatus according to claim 1 , wherein

the processing circuitry selects the plurality of speech units for synthesizing the speech by referring to speech unit information stored in an information memory.

3. The apparatus according to claim 2 , wherein

an access speed of the information memory is quicker than an access speed of the waveform memory.

4. The apparatus according to claim 1 , further comprising:

a text input unit by which a text to be converted to the speech is input.

5. The apparatus according to claim 1 , further comprising:

at least one of a speaker or a head phone by which the speech is output.

6. The apparatus according to claim 1 , further comprising:

a computer operation system configured to execute copying of the at least two speech unit waveforms by the processing circuitry.

7. The apparatus according to claim 1 , further comprising:

a middle-ware (MW) program configured to execute copying of the at least two speech unit waveforms by the processing circuitry.

8. The apparatus according to claim 1 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, or a pitch waveform.

9. The apparatus according to claim 1 , wherein

the processing circuitry synthesizes the speech by concatenating the speech unit waveforms with modification or without modification.

10. A method for synthesizing a speech, comprising:

acquiring by processing circuitry, from a waveform memory, speech unit waveforms corresponding to a plurality of speech units for synthesizing the speech; and

synthesizing by the processing circuitry, the speech by utilizing the speech unit waveforms,

wherein the acquiring includes

copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the at least two speech unit waveforms being less than or equal to a size of the buffer, the at least two speech unit waveforms corresponding to at least two speech units included in the plurality of speech units.

11. The method according to claim 10 , further comprising:

selecting by the processing circuitry, the plurality of speech units for synthesizing the speech by referring to speech unit information stored in an information memory.

12. The method according to claim 10 , wherein

the copying is executed by a computer operation system.

13. The method according to claim 10 , wherein

the copying is executed by a middle-ware (MW) program.

14. The method according to claim 10 , further comprising:

outputting the speech by at least one of a speaker or a head phone.

15. The method according to claim 10 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, or a pitch waveform.

16. The method according to claim 10 , wherein

the synthesizing includes

synthesizing the speech by concatenating the speech unit waveforms with modification or without modification.

17. A non-transitory computer readable memory storing program instructions, which when executed by a computer, results in performance of a method for synthesizing a speech, the method comprising:

acquiring by processing circuitry, from a waveform memory, speech unit waveforms corresponding to a plurality of speech units for synthesizing the speech; and

synthesizing by the processing circuitry, the speech by utilizing the speech unit waveforms,

wherein the acquiring includes

copying at least two speech unit waveforms from a continuous region of the waveform memory onto a buffer by one access, a data quantity of the at least two speech unit waveforms being less than or equal to a size of the buffer, the at least two speech unit waveforms corresponding to at least two speech units included in the plurality of speech units.

18. The non-transitory computer readable medium according to claim 17 , further comprising:

selecting by the processing circuitry, the plurality of speech units for synthesizing the speech by referring to speech unit information stored in an information memory.

19. The non-transitory computer readable medium according to claim 17 , wherein

the speech unit waveform is one of a waveform itself, a parameter to generate the waveform, encoded data of the waveform, or a pitch waveform.

20. The non-transitory computer readable medium according to claim 17 , wherein

the synthesizing includes

synthesizing the speech by concatenating the speech unit waveforms with modification or without modification.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
Priority Claims (1)
JP 2006-139587 · May 18, 2006 · national
Continuity (3)
Division 13860319 · Apr 10, 2013
Continuation 11745785 · May 8, 2007
Related Publication 20140180681A1 · Jun 26, 2014