ADAPTIVE DIGITAL ASSISTANT AND SPOKEN GENOME
Embodiments of the invention include a context sensitive adaptive digital assistant for personalized interaction. Embodiments of the invention also include a spoken genome for characterization and analysis of human voice. Aspects of the invention include selecting a starter vocabulary, receiving voice communications from a user, and modifying the starter vocabulary to generate a personalized lexicon. Aspects of the invention also include analyzing and categorizing human voice according to a plurality of characteristics, and creating a spoken genome database.
1 . A computer-implemented method for personalized digital interaction comprising:
selecting a starter vocabulary from a starter vocabulary set;
receiving a plurality of user voice communications from a user;
generating a frequent word list based at least in part on the plurality of user voice communications;
modifying the starter vocabulary with a plurality of words from the frequent word list to generate a personalized lexicon; and
generating a personalized verbal output based at least in part on the personalized lexicon.
2 . The method of claim 1 , further comprising receiving a secondary learning input from a smart device and modifying the personalized vocabulary based at least in part on the secondary learning input.
3 . The method of claim 1 , wherein selecting the starter vocabulary comprises:
receiving a user demographic data set,
comparing the user demographic data set to a plurality of characteristics of the starter vocabulary set, and
determining, based at least in part on the comparison, a preferred starter vocabulary.
4 . The method of claim 1 , further comprising receiving a voice command from the user.
5 . The method of claim 4 , further comprising:
analyzing the voice command to determine a mood; and
adjusting a characteristic of the personalized verbal output based at least in part on the mood.
6 . The method of claim 1 , further comprising selecting a specialized vocabulary and augmenting the personalized lexicon with the specialized vocabulary.
7 . The method of claim 1 , further comprising selecting a set of accent features and adjusting a characteristic of the personalized verbal output based at least in part on the accent features.
8 . The method of claim 1 , wherein modifying the starter vocabulary comprises replacing a word from the starter vocabulary with a frequent word from the frequent word list or augmenting the starter vocabulary with a new word from the frequent word list.
9 . The method of claim 1 , wherein the personalized verbal output comprises a personalized gender, pitch, cadence, tonality, rhythm, accent, timing, slurring, elocution, or combination thereof.
10 . The method of claim 1 , further comprising communicating with an external classification system.
11 . A computer program product for characterization and analysis of human voice, wherein the computer program product comprises:
a computer readable storage medium readable by a processing circuit and storing program instructions for execution by the processing circuit for performing a method comprising:
receiving a plurality of media files, wherein the plurality of media files comprise spoken words;
categorizing the plurality of media files according to spoken genome properties to create a categorized spoken genome database;
receiving a user media preference;
determining a user media profile based at least in part on the user media preference, wherein the media profile comprises a spoken genome property; and
providing a media recommendation based at least in part on the user media profile and the categorized spoken genome database.
12 . The computer program product of claim 11 , wherein the spoken genome property is selected from the group consisting of gender, pitch, cadence, tonality, rhythm, accent, timing, slurring, and elocution.
13 . The computer program product of claim 11 , wherein the media recommendation comprises a recommended television program, a recommended movie, or a recommended audio book.
14 . A processing system for characterization and analysis of human voice comprising:
a processor in communication with one or more types of memory, the processor configured to:
receive a first reference voice sampling, wherein the first reference voice sampling comprises a plurality of reference voices corresponding to a first known reference quality; and
analyze the first voice sampling to determine a spoken genome property corresponding to the first known reference quality.
15 . The processing system of claim 14 , wherein the processor is configured to output the spoken genome property corresponding to the first known reference quality.
16 . The processing system of claim 14 , wherein the processor is configured to:
receive a second reference voice sampling, wherein the second reference voice sampling comprises a plurality of reference voices corresponding to a second known reference quality; and
analyze the second voice sampling to determine a spoken genome property corresponding to the second known reference quality.
17 . The processing system of claim 16 , wherein the processor is further configured to generate a spoken genome database, wherein the spoken genome database comprises:
a first bin comprising the first known reference quality and the spoken genome property corresponding to the first known reference quality, and
a second bin comprising the second known reference quality and the spoken genome property corresponding to the second known reference quality.
18 . The processing system of claim 14 , wherein the processor is further configured to:
receive a target voice input corresponding to a candidate;
determine a spoken genome property corresponding to the target voice input;
compare the target voice input spoken genome property to the spoken genome property corresponding to the first known reference quality; and
determine whether the candidate corresponds to the first known reference quality.
19 . The processing system of claim 18 , wherein the processor is configured to output the determination of whether the candidate corresponds the first known reference quality to a display.
20 . The processing system of claim 14 , wherein the spoken genome property is selected from the group consisting of pitch, cadence, tonality, rhythm, and timing.