IP Library Granted Patent US 9,997,154
Granted Patent B2
US 9,997,154 · App. 14/275,349 · Granted Jun 12, 2018

System and method for prosodically modified unit selection databases

Inventors: Alistair D. Conkie (Morristown, NJ); Ladan Golipour (Morristown, NJ); Ann K. Syrdal (Morristown, NJ)
Assignee: AT&T Intellectual Property I, L.P.
G10L13/06G10L15/02G10L15/063G10L21/003G10L25/90G11C7/16G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,997,154
App. No.
14/275,349
Granted
Jun 12, 2018
Kind
B2
Abstract

Systems, methods, and computer-readable storage devices to improve the quality of synthetic speech generation. A system selects speech units from a speech unit database, the speech units corresponding to text to be converted to speech. The system identifies a desired prosodic curve of speech produced from the selected speech units, and also identifies an actual prosodic curve of the speech units. The selected speech units are modified such that a new prosodic curve of the modified speech units matches the desired prosodic curve. The system stores the modified speech units into the speech unit database for use in generating future speech, thereby increasing the prosodic coverage of the database with the expectation of improving the output quality.

Claims (41)

1. A method comprising:

selecting, via a processor, speech units from a speech unit database, where the speech units are used to generate speech correspond to text;

identifying a desired prosodic curve of the speech to be produced from the speech units;

identifying an actual prosodic curve of the speech units;

decomposing, via a residual-excited linear prediction algorithm, the speech units into residual coefficients and linear predictive coder coefficients;

determining a cost of modifying the residual coefficients to yield a determination;

modifying, via a pitch synchronous overlap and add algorithm, the residual coefficients, to yield modified residual coefficients based on the determination;

combining, via the residual-excited linear prediction algorithm, the modified residual coefficients with the linear predictive coder coefficients, to yield new speech units, such that a new prosodic curve corresponding to the new speech units conforms to the desired prosodic curve; and

generating the speech using the new speech units.

2. The method of claim 1 , wherein the desired prosodic curve corresponds to a type of intonation.

3. The method of claim 2 , wherein the type of intonation is one of a declarative intonation and a yes-no interrogative.

4. The method of claim 1 , wherein the modifying of the residual coefficients comprises scaling a pitch of the residual coefficients.

5. The method of claim 1 , wherein the cost of the modifying of the residual coefficients is higher than a cost of retrieving the new speech units from the speech unit database.

6. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

selecting speech units from a speech unit database, where the speech units are used to generate speech correspond to text;

identifying a desired prosodic curve of the speech to be produced from the speech units;

identifying an actual prosodic curve of the speech units;

decomposing, via a residual-excited linear prediction algorithm, the speech units into residual coefficients and linear predictive coder coefficients;

determining a cost of modifying the residual coefficients to yield a determination;

modifying, via a pitch synchronous overlap and add algorithm, the residual coefficients, to yield modified residual coefficients based on the determination;

combining, via the residual-excited linear prediction algorithm, the modified residual coefficients with the linear predictive coder coefficients, to yield new speech units, such that a new prosodic curve corresponding to the new speech units conforms to the desired prosodic curve; and

generating the speech using the new speech units.

7. The system of claim 6 , wherein the desired prosodic curve corresponds to a type of intonation.

8. The system of claim 7 , wherein the type of intonation is one of a declarative intonation and a yes-no interrogative.

9. The system of claim 6 , wherein the modifying of the residual coefficients comprises scaling a pitch of the residual coefficients.

10. The system of claim 6 , wherein the cost of the modifying of the residual coefficients is higher than a cost of retrieving the new speech units from the speech unit database.

11. A non-transitory computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

selecting speech units from a speech unit database, where the speech units are used to generate speech correspond to text;

identifying a desired prosodic curve of the speech to be produced from the speech units;

identifying an actual prosodic curve of the speech units;

decomposing, via a residual-excited linear prediction algorithm, the speech units into residual coefficients and linear predictive coder coefficients;

determining a cost of modifying the residual coefficients to yield a determination;

modifying, via a pitch synchronous overlap and add algorithm, the residual coefficients, to yield modified residual coefficients based on the determination;

combining, via the residual-excited linear prediction algorithm, the modified residual coefficients with the linear predictive coder coefficients, to yield new speech units, such that a new prosodic curve corresponding to the new speech units conforms to the desired prosodic curve; and

generating the speech using the new speech units.

12. The non-transitory computer-readable storage device of claim 11 , wherein the desired prosodic curve corresponds to a type of intonation.

13. The non-transitory computer-readable storage device of claim 12 , wherein the type of intonation is one of a declarative intonation and a yes-no interrogative.

14. The non-transitory computer-readable storage device of claim 11 , wherein the modifying of the residual coefficients comprises scaling a pitch of the residual coefficients.

15. The non-transitory computer-readable storage device of claim 11 , wherein the cost of the modifying of the residual coefficients is higher than a cost of retrieving the new speech units from the speech unit database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2014
From: CONKIE, ALISTAIR D.; GOLIPOUR, LADAN; SYRDAL, ANN K.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 033922/0838 →
Continuity (1)
Related Publication 20150325248A1 · Nov 12, 2015