IP Library Granted Patent US 8,326,629
Granted Patent B2
US 8,326,629 · App. 11/164,415 · Granted Dec 4, 2012

Dynamically changing voice attributes during speech synthesis based upon parameter differentiation for dialog contexts

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,326,629
App. No.
11/164,415
Granted
Dec 4, 2012
Kind
B2
Abstract

A method of speech synthesis can include automatically identifying spoken passages and non-spoken passages within a text source and converting the text source to speech by applying different voice configurations to different portions of text within the text source according to whether each portion of text was identified as a spoken passage or a non-spoken passage. The method further can include identifying the speaker and/or the gender of the speaker and applying different voice configurations according to the speaker identity and/or speaker gender.

Claims (42)

1. A computer-implemented method of speech synthesis to create an audio recording from a text source comprising a story including a first character and a second character, the method comprising:

automatically identifying based, at least in part, on a content of the text source, at least one first spoken passage as being spoken by the first character, at least one second passage as being spoken by the second character, and at least one non-spoken passage within the text source from which speech is to be synthesized to create the audio recording;

automatically assigning a first voice configuration for the first character to the at least one first spoken passage, a second voice configuration for the second character to the at least one second spoken passage, and a third voice configuration to the at least one non-spoken passage;

automatically identifying at least one third spoken passage having a measure of certainty regarding an identity of the character speaking the at least one third spoken passage being less than a threshold value;

automatically assigning to the at least one third spoken passage, a voice configuration for a character assigned to a spoken passage preceding the at least one third spoken passage; and

creating the audio recording by converting the text source to speech by selectively applying the first voice configuration to the at least one first spoken passage, applying the second voice configuration to the at least one second spoken passage, and applying the third voice configuration to the at least one non-spoken passage.

2. The method of claim 1 , further comprising:

automatically determining a speaker gender for at least one fourth spoken passage based, at least in part, on gender specific pronouns identified in the text source.

3. The method of claim 1 , further comprising:

automatically determining a speaker gender for at least one fourth spoken passage based, at least in part, on gender specific proper names identified in the text source.

4. The method of claim 1 , wherein the audio recording is an audiobook of the story.

5. The method of claim 1 , wherein the audio recording is a podcast.

6. The method of claim 1 , wherein the at least one first spoken passage includes a plurality of first spoken passages identified as being spoken by the first character, wherein the method further comprises:

determining a confidence value for at least one of the plurality of first spoken passages that the at least one of the plurality of first spoken passages is associated with the first character in the story; and

visually indicating the confidence value on a display.

7. A text-to-speech system comprising:

at least one computer programmed to perform speech synthesis for creating an audio recording from a text source comprising a story including a first character and a second character, wherein the at least one computer is programmed to:

automatically identify based, at least in part, on a content of the text source, at least one first spoken passage as being spoken by the first character, at least one second passage as being spoken by the second character, and at least one non-spoken passage within the text source from which speech is to be synthesized to create the audio recording;

automatically assign a first voice configuration for the first character to the at least one first spoken passage, a second voice configuration for the second character to the at least one second spoken passage, and a third voice configuration to the at least one non-spoken passage;

automatically identify at least one third spoken passage having a measure of certainty regarding an identity of the character speaking the at least one third spoken passage being less than a threshold value;

automatically assign to the at least one third spoken passage, a voice configuration for a character assigned to a spoken passage preceding the at least one third spoken passage; and

create the audio recording by converting the text source to speech by selectively applying the first voice configuration to the at least one first spoken passage, applying the second voice configuration to the at least one second spoken passage, and applying the third voice configuration to the at least one non-spoken passage.

8. The text-to-speech system of claim 7 , wherein the at least one computer is programmed to automatically determine a speaker gender for at least one fourth spoken passage based, at least in part on gender specific pronouns identified in the text source.

9. The text-to-speech system of claim 7 , wherein the at least one computer is programmed to automatically determine a speaker gender for at least one fourth spoken passage based, at least in part on, gender specific proper names identified in the text source.

10. The text-to-speech system of claim 7 , wherein the audio recording is an audiobook of the story.

11. The text-to-speech system of claim 7 , wherein the audio recording is a podcast.

12. The text-to-speech system of claim 7 , wherein the at least one computer is further programmed to:

determine a confidence value for at least one of the plurality of first spoken passages that the at least one of the plurality of first spoken passages is associated with the first character in the story; and

visually indicate the confidence value on a display.

13. A machine readable storage having stored thereon a computer program having a plurality of code sections comprising:

code for automatically identifying based, at least in part, on a content of the text source, at least one first spoken passage as being spoken by a first character of a story, at least one second passage as being spoken by a second character of the story, and at least one non-spoken passage within the text source from which speech is to be synthesized to create the audio recording;

code for automatically assigning a first voice configuration for the first character to the at least one first spoken passage, a second voice configuration for the second character to the at least one second spoken passage, and a third voice configuration to the at least one non-spoken passage;

code for automatically identifying at least one third spoken passage having a measure of certainty regarding an identity of the character speaking the at least one third spoken passage being less than a threshold value;

code for automatically assigning to the at least one third spoken passage, a voice configuration for a character assigned to a spoken passage preceding the at least one third spoken passage; and

code for creating the audio recording by converting the text source to speech by selectively applying the first voice configuration to the at least one first spoken passage, applying the second voice configuration to the at least one second spoken passage, and applying the third voice configuration to the at least one non-spoken passage.

14. The machine readable storage of claim 13 , further comprising code for automatically determining a speaker gender for at least one fourth spoken passage based, at least in part, on gender specific pronouns identified in the text source.

15. The machine readable storage of claim 13 , wherein the code for automatically determining a speaker gender for at least one fourth spoken passage based, at least in part, on gender specific proper names identified in the text source.

16. The machine readable storage of claim 13 , wherein the audio recording is an audiobook of the story.

17. The machine readable storage of claim 13 , wherein the audio recording is a podcast.

18. The machine readable storage of claim 13 , further comprising:

code for determining a confidence value for at least one of the plurality of first spoken passages that the at least one of the plurality of first spoken passages is associated with the first character in the story; and

code for visually indicating the confidence value on a display.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2005
From: SKURATOVSKY, ILYA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 016808/0863 →