IP Library Granted Patent US 9,183,831
Granted Patent B2
US 9,183,831 · App. 14/227,462 · Granted Nov 10, 2015

Text-to-speech for digital literature

Inventors: Donna Karen Byron (Littleton, MA); Alexander Pikovsky (Littleton, MA); Eric Woods (Durham, NC)
Assignee: International Business Machines Corporation
G10L13/027G10L13/043
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,183,831
App. No.
14/227,462
Granted
Nov 10, 2015
Kind
B2
Abstract

A digital work of literature is vocalized using enhanced text-to-speech (TTS) controls by analyzing a digital work of literature using natural language processing to identify speaking character voice characteristics associated with context of each quote as extracted from the first work of literature; converting the character voice characteristics to audio metadata to control text-to-speech audio synthesis for each quote; transforming the audio metadata into text-to-speech engine commands, each quote being associated with audio synthesis control parameters for the TTS in the context of each the quotes in the work of literature; and inputting the commands to a text-to-speech engine to cause vocalization of the work of literature according to the words of each quote, character voice characteristics of corresponding to each quote, and context corresponding to each quote.

Claims (26)

1. A computer program product for vocalizing a digital work of literature, comprising:

a tangible, computer-readable storage memory device; and

program instructions encoded by the tangible, computer-readable storage memory device, for causing a processor to perform operations comprising:

analyzing a first digital work of literature using natural language processing to identify speaking character voice characteristics associated with context of each quote as extracted from the first work of literature;

converting the character voice characteristics to audio metadata to control text-to-speech audio synthesis for each quote;

transforming the audio metadata into text-to-speech engine commands, wherein each quote is associated with audio synthesis control parameters for rendering the corresponding character voice characteristics according to the context of the quotes in the first work of literature; and

inputting the commands to a text-to-speech engine to cause vocalization of the first work of literature, wherein spoken quotes by characters are vocalized according to the words of each quote, character voice characteristics corresponding to each quote, and context corresponding to each quote.

2. The computer program product as set forth in claim 1 wherein the program instructions further comprise instructions for inferring a narrator character for vocalization of text from the first digital work of literature which are not drawn from a character quote.

3. The computer program product as set forth in claim 1 wherein the first work of literature comprises at least one text source selected from the group consisting of a digital document, a digital book, a digital magazine, a data stream, and a web page, wherein the character characteristics are selected from the group consisting of gender, age, nationality, ethnicity, mood, profession, hero role, villain role, likability, and education level, and wherein the context is selected from the group consisting of genre, tempo, tone, pace, plot elements, subplot elements, and literary elements.

4. The computer program product as set forth in claim 3 wherein the vocalization is adjusted to reflect changes according to plot elements.

5. The computer program product as set forth in claim 1 wherein the program instructions further comprise instructions for:

searching a corpus of sources extrinsic to the first digital work of literature to the identify the same or similar character voice characteristics and same or similar context, wherein the converting and transforming include converting and transforming according to found character voice characteristics and context; and

responsive to detecting a current character in the first digital work of literature is derived from, same as or similar to a character in a previously-vocalized work of literature, retrieving and using character voice characteristics according to the previously-vocalized work of literature in the context of a second digital work of literature.

6. A system for vocalizing a digital work of literature, comprising:

a computer having a processor; and

a tangible, computer-readable storage memory device encoding program instructions for causing the processor to perform operations comprising:

analyzing a first digital work of literature using natural language processing to identify speaking character voice characteristics associated with context of each quote as extracted from the first work of literature;

converting the character voice characteristics to audio metadata to control text-to-speech audio synthesis for each quote;

transforming the audio metadata into text-to-speech engine commands, wherein each quote is associated with audio synthesis control parameters for rendering the corresponding character voice characteristics according to the context of the quotes in the first work of literature; and

inputting the commands to a text-to-speech engine to cause vocalization of the first work of literature, wherein spoken quotes by characters are vocalized according to the words of each quote, character voice characteristics corresponding to each quote, and context corresponding to each quote.

7. The system as set forth in claim 6 wherein the program instructions further comprise instructions for inferring a narrator character for vocalization of text from the first digital work of literature which are not drawn from a character quote.

8. The system as set forth in claim 6 wherein the first work of literature comprises at least one text source selected from the group consisting of a digital document, a digital book, a digital magazine, a data stream, and a web page, wherein the character characteristics are selected from the group consisting of gender, age, nationality, ethnicity, mood, profession, hero role, villain role, likability, and education level, and wherein the context is selected from the group consisting of genre, tempo, tone, pace, plot elements, subplot elements, and literary elements.

9. The system as set forth in claim 8 wherein the vocalization is adjusted to reflect changes according to plot elements.

10. The system as set forth in claim 6 wherein the program instructions further comprise instructions for:

searching a corpus of sources extrinsic to the first digital work of literature to the identify the same or similar character voice characteristics and same or similar context, wherein the converting and transforming include converting and transforming according to found character voice characteristics and context; and

responsive to detecting a current character in the first digital work of literature is derived from, same as or similar to a character in a previously-vocalized work of literature, retrieving and using character voice characteristics according to the previously-vocalized work of literature in the context of a second digital work of literature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2014
From: BYRON, DONNA K.; PIKOVSKY, ALEXANDER; WOODS, ERIC
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 032543/0031 →
Continuity (1)
Related Publication 20150279349A1 · Oct 1, 2015