IP Library Granted Patent US 9,240,180
Granted Patent B2
US 9,240,180 · App. 13/308,860 · Granted Jan 19, 2016

System and method for low-latency web-based text-to-speech without plugins

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,240,180
App. No.
13/308,860
Granted
Jan 19, 2016
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for reducing latency in web-browsing TTS systems without the use of a plug-in or Flash® module. A system configured according to the disclosed methods allows the browser to send prosodically meaningful sections of text to a web server. A TTS server then converts intonational phrases of the text into audio and responds to the browser with the audio file. The system saves the audio file in a cache, with the file indexed by a unique identifier. As the system continues converting text into speech, when identical text appears the system uses the cached audio corresponding to the identical text without the need for re-synthesis via the TTS server.

Claims (49)

1. A method comprising:

receiving, from a client, text associated with a request for text-to-speech synthesis;

performing, via a processor of a computing device, an analysis of the text to identify a plurality of intonational phrases in the text, wherein a size of the text being analyzed is based on a network latency;

generating, via the processor, a first file containing text-to-speech data for a first intonational phrase of the plurality of intonational phrases using a first text-to-speech voice, wherein the first text-to-speech voice is selected based on user preferences, and wherein the first intonational phrase is indexed by a first unique identifier;

generating, via the processor, a second file containing the text-to-speech data for a second intonational phrase of the plurality of intonational phrases using a second text-to-speech voice, wherein the second text-to-speech voice is selected based on the user preferences, and wherein the second intonational phrase is indexed by a second unique identifier;

storing the first file and the second file in a cache on a web-server;

transmitting the first file to the client in response to the request; and

while the client plays the first file, generating additional files containing additional text-to-speech data for remaining intonational phrases of the plurality of intonational phrases, wherein the remaining intonational phrases comprise the second intonational phrase, and wherein each of the additional files is indexed by the first unique identifier plus a respective offset.

2. The method of claim 1 , wherein an intonational phrase is a phrase in which intonation within the phrase only depends on text inside the phrase.

3. The method of claim 1 , wherein the first file is indexed by a unique identifier.

4. The method of claim 1 , wherein the first file contains notification information.

5. The method of claim 1 , wherein the unique identifier comprises a text identifier and an offset index.

6. The method of claim 1 , wherein the additional files contain additional notification information.

7. The method of claim 1 , wherein generating the additional files occurs while the web browser plays the text-to-speech data in the first file.

8. The method of claim 1 , wherein the receiving and the transmitting occur on the web server, wherein the web server deletes items saved in the cache within an expiration threshold.

9. The method of claim 1 , further comprising transmitting one of the first file and a supplemental file of the additional files to the web browser in response to an additional request.

10. The method of claim 4 , wherein the notification information comprises synchronization data.

11. The method of claim 1 , wherein boundaries between intonational phrases comprise silence.

12. The method of claim 1 , further comprising:

receiving text-to-speech settings from the client; and

generating the first file and the additional files based on the text-to-speech settings.

13. The method of claim 1 , further comprising:

generating parallel versions of the first file and the additional files using different text-to-speech voices.

14. A system comprising:

a processor;

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving, from a client, text associated with a request for text-to-speech synthesis;

performing, via a processor of a computing device, an analysis of the text to identify a plurality of intonational phrases in the text, wherein a size of the text being analyzed is based on a network latency;

generating, via the processor, a first file containing text-to-speech data for a first intonational phrase of the plurality of intonational phrases using a first text-to-speech voice, wherein the first text-to-speech voice is selected based on user preferences, and wherein the first intonational phrase is indexed by a first unique identifier;

generating, via the processor, a second file containing the text-to-speech data for a second intonational phrase of the plurality of intonational phrases using a second text-to-speech voice, wherein the second text-to-speech voice is selected based on the user preferences, and wherein the second intonational phrase is indexed by a second unique identifier;

storing the first file and the second file in a cache on a web-server;

transmitting the first file to the client in response to the request; and

while the client plays the first file, generating additional files containing additional text-to-speech data for remaining intonational phrases of the plurality of intonational phrases, wherein the remaining intonational phrases comprise the second intonational phrase, and wherein each of the additional files is indexed by the first unique identifier plus a respective offset.

15. The system of claim 14 , wherein the operations are associated with a web browser.

16. The system of claim 15 , wherein no browser plugin is required for the operations.

17. The system of claim 14 , wherein the computer-readable storage medium has additional instructions stored which, when executed by the processor, result in operations comprising:

receiving user input navigating to a different position within the text;

identifying a new offset for the different position; and

fetching a corresponding file from the server for playback based on the unique identifier and the new offset.

18. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving, from a client, text associated with a request for text-to-speech synthesis;

performing, via a processor of a computing device, an analysis of the text to identify a plurality of intonational phrases in the text, wherein a size of the text being analyzed is based on a network latency;

generating, via the processor, a first file containing text-to-speech data for a first intonational phrase of the plurality of intonational phrases using a first text-to-speech voice, wherein the first text-to-speech voice is selected based on user preferences, and wherein the first intonational phrase is indexed by a first unique identifier;

generating, via the processor, a second file containing the text-to-speech data for a second intonational phrase of the plurality of intonational phrases using a second text-to-speech voice, wherein the second text-to-speech voice is selected based on the user preferences, and wherein the second intonational phrase is indexed by a second unique identifier;

storing the first file and the second file in a cache on a web-server;

transmitting the first file to the client in response to the request; and

while the client plays the first file, generating additional files containing additional text-to-speech data for remaining intonational phrases of the plurality of intonational phrases, wherein the remaining intonational phrases comprise the second intonational phrase, and wherein each of the additional files is indexed by the first unique identifier plus a respective offset.

19. The computer-readable storage device of claim 18 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising:

generating parallel versions of the first file and the additional files using different text-to-speech voices.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2011
From: CONKIE, ALISTAIR D.; BEUTNAGEL, MARK CHARLES; MISHRA, TANIYA
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 027311/0001 →