Systems and methods for machine-generated avatars
Systems and methods are disclosed for creating a machine generated avatar. A machine generated avatar is an avatar generated by processing video and audio information extracted from a recording of a human speaking a reading corpora and enabling the created avatar to be able to say an unlimited number of utterances, i.e., utterances that were not recorded. The video and audio processing consists of the use of machine learning algorithms that may create predictive models based upon pixel, semantic, phonetic, intonation, and wavelets.
1. A method for creating a machine-generated avatar, comprising:
one or more reprocessing and training steps, comprising:
receiving video and audio recording data;
timestamping the recording data with phoneme times;
extracting, based on the timestamping, phoneme clips and viseme clips;
storing individual phoneme instances and individual viseme instances;
extracting, based on the timestamping, transition light cones;
extracting, based on the timestamping, audio clips;
associating, based on the timestamping, the transition light cones with the audio clips;
parsing the audio clips into sentences;
tagging the sentences with parts of speech;
training a modulation model on the tagged sentences;
training a light cone model for viseme transitions, and
training an audio model on phoneme transitions.