Content and advertising service using one server for the content, sending it to another for advertisement and text-to-speech synthesis before presenting to user
Methods and systems for providing a network-accessible text-to-speech synthesis service are provided. The service accepts content as input. After extracting textual content from the input content, the service transforms the content into a format suitable for high-quality speech synthesis. Additionally, the service produces audible advertisements, which are combined with the synthesized speech. The audible advertisements themselves can be generated from textual advertisement content.
1. A method, comprising:
receiving, by a computer system at a first location of an information network, content from a second location of the information network in response to a request sent to the second location from a third location of the information network;
identifying advertising information based on the content;
synthesizing audible speech data by performing, at the computer system, text-to-speech synthesis using the content;
obtaining advertising audio data based on the advertising information by performing text-to-speech synthesis at the computer system;
combining, by the computer system, the audible speech data and the advertising audio data into a set of combined audio data; and
conveying the set of combined audio data to the third location via the information network.
2. The method of claim 1 further comprising:
analyzing one or more selection parameters to determine selection criteria;
wherein said identifying the advertising information is further based on the selection criteria.
3. The method of claim 1 , wherein the first location corresponds to a text-to-speech server, the second location corresponds to a content server, and the third location corresponds to a user device.
4. The method of claim 1 , wherein said synthesizing audible speech data further comprises using rules based on the content.
5. The method of claim 1 , wherein said combining the audible speech data and the advertising audio data into the set of combined audio data includes combining the audible speech data and data corresponding to a plurality of audio advertisements.
6. The method of claim 1 , the method further comprising extracting textual information from the content using character recognition.
7. The method of claim 1 , wherein said identifying the advertising information is further based on information related to a requester of the audible speech data.
8. A system, comprising:
at least one processor configured to implement:
a content module operable to receive content at a first location from a content server at a second location in response to a request sent to the content server from a device at a third location;
an advertising content module at the first location operable to obtain advertising information based on the content;
a synthesis module at the first location operable to synthesize audible speech data from the content;
and a presentation module operable to convey the audible speech data and the advertising information to the device at the third location via an information network.
9. The system of claim 8 , wherein the advertising content module is further operable to:
analyze one or more selection parameters to determine selection criteria; and
select the advertising information based on the selection criteria.
10. The system of claim 9 , wherein the selection parameters include information about a requester of the audible speech data.
11. The system of claim 8 , wherein the synthesis module is further operable to synthesize the audible speech data using a set of rules that correspond to a requester of the audible speech data, the second location, one or more topics, the requester's identity, or characteristics of the information network.
12. The system of claim 8 , wherein the presentation module is further operable to combine the audible speech data and data corresponding to a plurality of audible advertisements.
13. The system of claim 8 , wherein the system is configured to extract textual information from the content by performing character recognition.
14. The system of claim 8 , wherein the advertising content module is operable to obtain advertising information further based on information related to a requester of the audible speech data.
15. A method, comprising:
receiving content at a text-to-speech server computer system at a first location of an information network, wherein the received content was sent from a content server at a second location of the information network to the first location via the information network, and wherein the receiving is in response to a request sent to the second location from a user computing device at a third location of the information network;
extracting, by the computer system, one or more topics from the content;
selecting, by the computer system, advertising information based on the one or more topics;
synthesizing, by the computer system, audible speech data corresponding to the content;
obtaining advertising audio data based on the advertising information;
combining the audible speech data and the advertising audio data into a combined set of audio data;
and conveying the combined set of audio data to the third location via the information network.