IP Library Granted Patent US 7,369,994
Granted Patent B1
US 7,369,994 · App. 11/381,544 · Granted May 6, 2008

Methods and apparatus for rapid acoustic unit selection from a large speech corpus

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,369,994
App. No.
11/381,544
Granted
May 6, 2008
Kind
B1
Abstract

A speech synthesis system can select recorded speech fragments, or acoustic units, from a very large database of acoustic units to produce artificial speech. The selected acoustic units are chosen to minimize a combination of target and concatenation costs for a given sentence. However, as concatenation costs, which are measures of the mismatch between sequential pairs of acoustic units, are expensive to compute, processing can be greatly reduced by pre-computing and aching the concatenation costs. Accordingly, a method is disclosed for constructing an efficient concatenation cost database by synthesizing a large body of speech, identifying the acoustic unit sequential pairs generated and their respective concatention costs, and storing those concatenation costs likely to occur.

Claims (42)

1. A computer-implemented method of synthesizing speech, the method comprising:

selecting a pair of acoustic units from an acoustic unit database;

identifying a concatenation cost between the pair of acoustic units based on communication with a concatenation cost database; and

synthesizing speech using the concatenation cost for the selected pair of acoustic units.

2. The method of claim 1 , wherein the concatenation cost is a measure of the mismatch between the pair of acoustic units.

3. The method of claim 1 , wherein the concatenation cost database contains a subset of all possible acoustic unit sequential pairs.

4. The method of claim 1 , wherein the concatenation with the concatenation cost database comprises:

extracting a concatenation cost of the pair of acoustic units form the concatenation cost database if the concatenation cost database contains the concatenation cost of the pair of acoustic units; and

determining a value of the concatenation cost of the pair of acoustic units if the concatenation cost database does not contain the concatenation cost of the pair of acoustic units.

5. The method of claim 1 , wherein the concatenation cost database is derived at least in part using statistical techniques which predict acoustic unit sequential pairs likely to occur in speech.

6. The method of claim 1 , wherein the concatenation cost database is derived at least in part by assigning costs to acoustic unit sequential pairs.

7. The method of claim 1 , wherein selecting at least one acoustic unit from the acoustic unit database further uses at least one target cost of an acoustic unit, the target cost being a measure of the mismatch between an acoustic unit and a phoneme.

8. The method of claim 4 , wherein determining a value of the concatenation cost of the pair of acoustic units comprises computing the concatenation cost of the pair of acoustic units.

9. A concatenation cost database stored in a computer-readable medium, the concatenation cost database generated according to a method comprising:

identifying at least some acoustic units to prune an acoustic unit database; and

storing in a concatenation cost database, concatenation costs for sequential acoustic units associated with the pruned acoustic unit database.

10. A computer-readable medium storing instructions for controlling a computing device, the instructions comprising:

selecting a pair of acoustic units from an acoustic unit database;

identifying a concatenation cost between the pair of acoustic units based on communication with a concatenation cost database; and

synthesizing speech using the concatenation cost for the selected pair of acoustic units.

11. The computer-readable medium of claim 10 , wherein the concatenation cost is a measure of the mismatch between the pair of acoustic units.

12. The computer-readable medium of claim 10 , wherein the concatenation cost database contains a subset of all possible acoustic unit sequential pairs.

13. The computer-readable medium of claim 10 , wherein the communication with the concatenation cost database comprises:

extracting a concatenation cost of the pair of acoustic units from the concatenation cost database if the concatenation cost database contains the concatenation cost of the pair of acoustic units; and

determining a value of the concatenation cost of the pair of acoustic units if the concatenation cost database does not contain the concatenation cost of the pair of acoustic units.

14. The computer-readable medium of claim 13 , wherein determining a value of the concatenation cost of the pair of acoustic units comprises computing the concatenation cost of the pair of acoustic units.

15. The computer-readable medium of claim 10 , wherein the concatenation cost database is derived at least in part using statistical techniques which predict acoustic unit sequential pairs likely to occur in speech.

16. The computer-readable medium of claim 10 , wherein the concatenation cost database is derived at least in part by assigning costs to acoustic unit sequential pairs.

17. The computer-readable medium of claim 10 , wherein selecting at least one acoustic unit from the acoustic unit database further uses at least one target cost of an acoustic unit, the target cost being a measure of the mismatch between an acoustic unit and a phoneme.

18. A system for synthesizing speech, the system comprising:

a module configured to select a pair of acoustic units from an acoustic unit database;

a module configured to identify a concatenation cost between the pair of acoustic units based on communication with a concatenation cost database; and

a module configured to synthesize speech using the concatenation cost for the selected pair of acoustic units.

19. The system of claim 18 , wherein the concatenation cost is a measure of the mismatch between the pair of acoustic units.

20. The system of claim 18 , wherein the concatenation cost database contains a subset of all possible acoustic unit sequential pairs.

21. The system of claim 18 , wherein the communication with the concatenation cost database comprises:

extracting a concatenation cost of the pair of acoustic units from the concatenation cost database if the concatenation cost database contains the concatenation cost of the pair of acoustic units; and

determining a value of the concatenation cost of the pair of acoustic units if the concatenation cost database does not contain the concatenation cost of the pair of acoustic units.

22. The system of claim 18 , wherein the concatenation cost database is derived at least in part using statistical techniques which predict acoustic unit sequential pairs likely to occur in speech.

23. The system of claim 18 , wherein the concatenation cost database is derived at least in part by assigning costs to acoustic unit sequential pairs.

24. The system of claim 18 , wherein the module configured to select at least one acoustic unit from the acoustic unit database further uses at least one target cost of an acoustic unit, the target cost being a measure of the mismatch between an acoustic unit and a phoneme.

25. The system of claim 21 , wherein the module configured to determine a value of the concatenation cost of the pair of acoustic units comprises computing the concatenation cost of the pair of acoustic units.

Assignments (11)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041498/0316 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2016
From: BEUTNAGEL, MARK CHARLES; MOHRI, MEHRYAR; RILEY, MICHAEL DENNIS
To: AT&T CORP.
Reel/Frame 038289/0761 →