IP Library Granted Patent US 9,424,833
Granted Patent B2
US 9,424,833 · App. 14/572,451 · Granted Aug 23, 2016

Method and apparatus for providing speech output for speech-enabled applications

Inventors: Darren C. Meyer (Duxbury, MA); Corinne Bos-Plachez (Baisieux, FR); Martine Marguerite Staessen (Wervik, BE)
Assignee: Nuance Communications, Inc.
G10L13/02G10L13/04G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,424,833
App. No.
14/572,451
Granted
Aug 23, 2016
Kind
B2
Abstract

Techniques for providing speech output for speech-enabled applications. A synthesis system receives from a speech-enabled application a text input including a text transcription of a desired speech output. The synthesis system selects one or more audio recordings corresponding to one or more portions of the text input. In one aspect, the synthesis system selects from audio recordings provided by a developer of the speech-enabled application. In another aspect, the synthesis system selects an audio recording of a speaker speaking a plurality of words. The synthesis system forms a speech output including the one or more selected audio recordings and provides the speech output for the speech-enabled application.

Claims (22)

1. A method for providing a speech output for a speech-enabled application, the method comprising:

receiving from the speech-enabled application a text input comprising a text transcription of a desired speech output;

selecting, using at least one computer system, a sequence of audio recordings for concatenation to produce the desired speech output, the selected sequence of audio recordings comprising a first audio recording for concatenation with one or more other audio recordings in the selected sequence of audio recordings, the first audio recording selected for being of a speaker speaking a plurality of words in the text transcription, wherein selecting the sequence of audio recordings comprises applying one or more selection criteria that favor the selected sequence of audio recordings for being a smaller number of audio recordings than other candidate sequences of audio recordings for producing the desired speech output;

generating a speech output by concatenating the selected sequence of audio recordings; and

providing the generated speech output for the speech-enabled application.

2. The method of claim 1 , wherein the first audio recording is of the speaker reading at least a portion of a script, the at least a portion of the script corresponding exactly to the plurality of words, the plurality of words corresponding exactly to words of the text transcription.

3. The method of claim 1 , wherein the first audio recording is stored in a single audio file.

4. The method of claim 1 , wherein the plurality of words were spoken consecutively by the speaker when forming the first audio recording.

5. The method of claim 1 , wherein the first audio recording comprises the plurality of words spoken naturally by the speaker.

6. The method of claim 1 , wherein applying the one or more selection criteria to select the sequence of audio recordings comprises identifying the plurality of words matched by the first audio recording as being a longer sequence of contiguous words in the text transcription than a second plurality of words matched by a second audio recording in another candidate sequence of audio recordings for producing the desired speech output.

7. The method of claim 1 , wherein applying the one or more selection criteria to select the sequence of audio recordings comprises optimizing a cost function maximizing an average length of audio recordings in the sequence of audio recordings for concatenation.

8. The method of claim 1 , wherein applying the one or more selection criteria to select the sequence of audio recordings comprises optimizing a cost function minimizing a number of concatenations between audio recordings in the sequence of audio recordings.

9. A method for providing a speech output for a speech-enabled application, the method comprising:

receiving at least one input specifying a desired speech output;

selecting, using at least one computer system, at least one audio recording corresponding to at least a first portion of the desired speech output, the selecting comprising:

in response to identifying a desired contrastive stress pattern in the desired speech output, selecting the at least one audio recording based at least in part on metadata, belonging to the at least one audio recording, that identifies the at least one audio recording as carrying contrastive stress; and

providing for the speech-enabled application a speech output comprising the at least one audio recording.

10. At least one non-transitory computer-readable storage medium encoded with a plurality of computer-executable instructions that, when executed, perform a method for providing a speech output for a speech-enabled application, the method comprising:

receiving at least one input specifying a desired speech output;

selecting at least one audio recording corresponding to at least a first portion of the desired speech output, the selecting comprising:

in response to identifying a desired contrastive stress pattern in the desired speech output, selecting the at least one audio recording based at least in part on metadata, belonging to the at least one audio recording, that identifies the at least one audio recording as carrying contrastive stress; and

providing for the speech-enabled application a speech output comprising the at least one audio recording.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2016
From: MEYER, DARREN C.; BOS-PLACHEZ, CORRINE; STAESSEN, MARTINE MARGUERITE
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 037466/0397 →
Continuity (2)
Continuation 12704859 · Feb 12, 2010
Related Publication 20150106101A1 · Apr 16, 2015