IP Library Granted Patent US 10,192,541
Granted Patent B2
US 10,192,541 · App. 15/308,731 · Granted Jan 29, 2019

Systems and methods for generating speech of multiple styles from text

Inventors: Paolo Mairano (San Carlo Canavese, IT); Corinne Bos-Plachez (Baisieux, FR); Sourav Nandy (Lucknow, IN); Johan Wouters (Cham, CH); Silvia Maria Antonella Quazza (Turin, IT); Dong-Jian Yue (Shanghai, CN)
Assignee: Nuance Communications, Inc.
G10L13/10G10L13/047G10L13/07G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,192,541
App. No.
15/308,731
Granted
Jan 29, 2019
Kind
B2
Abstract

A text-to-speech (TTS) system includes components capable of supporting the generation of speech output in any of multiple styles, and may switch seamlessly from producing speech output in one style to producing speech output in another style. For example, a concatenative TTS system may include a speech base storing speech units associated with multiple speech styles, and a linguistic analysis component to generate a phonetic transcription specifying speech output in any of multiple styles. Text input may include a style indication associated with a particular segment of the input text. The linguistic analysis component may invoke encoded rules and/or components based upon the style indication, and generate a phonetic transcription specifying a speech style, which may be processed to generate output speech.

Claims (30)

1. A method for use in a text-to-speech system comprising a computer-implemented linguistic analysis component operative to generate a phonetic transcription based upon input text, a speech base comprising speech unit recordings associated with a plurality of styles of speech, and at least one computer-implemented speech generation component operative to generate output speech from stored speech unit recordings based at least in part on the phonetic transcription, the method comprising acts of:

(A) receiving, by the linguistic analysis component, input text produced by a text-producing application, wherein the text produced by a text-producing application comprises a speech style indication indicating a style of speech to be output by the text-to-speech system for an associated segment of the input text;

(B) generating, by the linguistic analysis component, a phonetic transcription based at least in part on the input text, the phonetic transcription specifying a first style of speech of the plurality of styles of speech to be output by the at least one speech generation component for the segment of the input text; and

(C) generating, by the at least one speech generation component, output speech based at least in part on the phonetic transcription generated in the act (B), wherein the generating comprises the at least one speech generation component selecting, from the speech unit recordings in the speech base, speech unit recordings associated with a second style of speech of the plurality of styles of speech, the second style of speech being different than the first style of speech, and concatenating the selected speech unit recordings to generate output speech in the first style.

2. The method of claim 1 , wherein the first style is a didactic style, and the second style is a neutral style.

3. The method of claim 1 , wherein the act (C) comprises the at least one speech generation component slowing down an output speech rate and/or inserting at least one pause in the output speech.

4. The method of claim 1 , further comprising an act (D) of generating output speech for another segment of the input text by applying, to the other segment, a statistical model associated with a style of speech specified in the phonetic transcription for the other segment.

5. The method of claim 1 , wherein the act (A) comprises receiving input text comprising a plurality of segments each having an associated speech style indication, at least one of the speech style indications being different than at least one other of the speech style indications, the act (B) comprises generating a phonetic transcription specifying a style of speech to be output for each one of the plurality of segments according to the speech style indication associated with the one segment, and the act (C) comprises generating output speech for each one of the plurality of segments according to the speech style indication associated with the one segment.

6. The method of claim 5 , wherein the plurality of segments constitute a single sentence.

7. The method of claim 1 , wherein the act (B) comprises the linguistic analysis component invoking one or more rules and/or components specific to a style of speech indicated by the speech style indication.

8. A text-to-speech system, comprising:

at least one storage facility storing a speech base comprising speech unit recordings associated with a plurality of styles of speech; and

at least one computer processor programmed to;

receive input text produced by a text-producing application, wherein the text produced by a text-producing application comprises a speech style indication indicating a style of speech to be output by the text-to-speech system for an associated segment of the input text;

generate a phonetic transcription based at least in part on the input text, the phonetic transcription specifying a first style of speech of the plurality of styles of speech to be output for the segment of the input text; and

generate output speech based at least in part on the generated phonetic transcription, the generating comprising selecting, from the speech unit recordings in the speech base, speech unit recordings associated with a second style of speech of the plurality of styles of speech, the second style of speech being different than the first style of speech, and concatenating the selected speech unit recordings to generate output speech in the first style.

9. The text-to-speech system of claim 8 , wherein the first style is a didactic style, and the second style is a neutral style.

10. The text-to-speech system of claim 8 , wherein the at least one computer processor is programmed to generate the output speech by slowing down an output speech rate and/or inserting at least one pause in the output speech.

11. The text-to-speech system of claim 8 , wherein the at least one computer processor is programmed to generate output speech for another segment of the input text by applying, to the other segment, a statistical model associated with a style of speech specified in the phonetic transcription for the other segment.

12. The text-to-speech system of claim 8 , wherein the at least one computer processor is programmed to:

receive input text comprising a plurality of segments each having an associated speech style indication, at least one of the speech style indications being different than at least one other of the speech style indications;

generate a phonetic transcription specifying a style of speech to be output for each one of the plurality of segments according to the speech style indication associated with the one segment; and

generate output speech for each one of the plurality of segments according to the speech style indication associated with the one segment.

13. The text-to-speech system of claim 12 , wherein the plurality of segments constitute a single sentence.

14. The text-to-speech system of claim 8 , wherein the at least one computer processor is programmed to generate the phonetic transcription by invoking one or more rules and/or components specific to a style of speech indicated by the speech style indication in the input text.

15. At least one non-transitory computer-readable storage medium having instructions encoded thereon which, when executed in a computer system, cause the computer system to perform a method comprising acts of:

(A) receiving input text produced by a text-producing application, wherein the text produced by a text-producing application comprises a speech style indication indicating a style of speech to be output by the text-to-speech system for an associated segment of the input text;

(B) generating a phonetic transcription based at least in part on the input text, the phonetic transcription specifying a first style of speech to be output for the segment of the input text; and

(C) generating output speech based at least in part on the phonetic transcription generated in the act (B), wherein the generating comprises selecting, from speech unit recordings in a speech base, speech unit recordings associated with a second style of speech that is different than the first style of speech, and concatenating the selected speech unit recordings to generate output speech in the first style.

16. The at least one non-transitory computer-readable storage medium of claim 15 , wherein the act (A) comprises receiving input text comprising a plurality of segments each having an associated speech style indication, at least one of the speech style indications being different than at least one other of the speech style indications, the act (B) comprises generating a phonetic transcription specifying a style of speech to be output for each one of the plurality of segments according to the speech style indication associated with the one segment, and the act (C) comprises generating output speech for each one of the plurality of segments according to the speech style indication associated with the one segment.

Assignments (6)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 064723/0519 →
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATIONS NUMBERS PREVIOUSLY RECORDED AT REEL: 055927 FRAME: 0620. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 16, 2021
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 056299/0078 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2018
From: BOS-PLACHEZ, CORINNE
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 047019/0857 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2018
From: MAIRANO, PAOLO; NANDY, SOURAV; WOUTERS, JOHAN; QUAZZA, SILVIA MARIA ANTONELLA; YUE, DONG JIAN
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 046997/0153 →
Continuity (1)
Related Publication 20170186418A1 · Jun 29, 2017
Cited By (1)
US 12,340,788