Systems and methods for automated synthetic voice pipelines
Disclosed are example embodiments of systems, method, and devices for generating audio assets. The example method includes receiving an input including at least one of audio, text, video, Java Script Object Notation (JSON), Extensible Markup Language (XML), Really Simple Syndication (RSS). The example method includes receiving configuration inputs including at least one of configure language, gender, and persona. The example method includes preparing for processing based on the at least one of configure language, gender, and persona. The example method includes processing the input based on the configuration inputs, the processing including at least one of transcribing, translating, brand safety, enrichment, generating custom Speech Synthesis Markup Language (SSML), and including generating an audio clip. The example method includes delivering the audio clip.
1 . A method of generating real-time audio assets for an automated synthetic voice pipeline, the method comprising:
polling and listening for a textual input in accordance with a schedule for an online broadcast;
receiving configuration inputs including at least one of configure language, gender, and persona;
preparing for processing based on the at least one of configure language, gender, and persona;
processing the textual input based on the configuration inputs, the processing including at least one of transcribing, translating, brand safety, enrichment, generating custom Speech Synthesis Markup Language (SSML), and including generating a real-time audio clip; and
broadcasting the audio clip via the automated synthetic voice pipeline,
wherein the polling and listening comprises establishing and maintaining a network connection configured for real-time, event-driven polling or subscription to the online broadcast source, and wherein the processing and broadcasting are performed by the synthetic voice generation module to deliver the audio clip for broadcast in real time according to the defined schedule.
2 . The method of claim 1 , wherein broadcasting the audio clip includes delivering files.
3 . The method of claim 1 , wherein broadcasting the audio clip further includes delivering metadata.
4 . The method of claim 1 , wherein preparing for processing based on the at least one of configure language, gender, and persona comprises preparing for processing based on configure language and generating the clip comprises generating the clip in a predetermined language based on the configuration language selected.
5 . The method of claim 1 , wherein preparing for processing based on the at least one of configure language, gender, and persona comprises preparing for processing based on gender and generating the clip comprises generating the clip having a voice corresponding to a selected gender.
6 . The method of claim 1 , wherein preparing for processing based on the at least one of configure language, gender, and persona comprises preparing for processing based on persona and generating the clip comprises generating a clip based on a predetermined persona, an aspect of someone's character that is perceived by others, based on the persona selected.