System and method for capturing, preserving, and representing human experiences and personality through a digital interface
A system and method to capture and interact with a comprehensive digital record of an individual's biographical history and produce a synthetic model of their personality. The captured biographical history is a detailed record of this individual's actions, interactions, and experiences over a period which may span decades of their lifetime. The biographical history is indexed by areas of data variability and neural network confidence variability to identify points of likely human interest. A synthetic personality model is generated as a representation of the individual's personality structure, biases, sentiments, and traits. The synthetic personality can be interacted with through a digital interface and demonstrates the interaction patterns, triggers, and habits of the original individual. The functioning and the performance of the system over an individual's lifespan are optimized through data synthesis and disposition.
1 . A method for capturing, preserving, and representing human experiences and personality through a digital interface, comprising:
digital activity recording (DAR) software installed on a personal computing device to capture and transmit digital device usage patterns;
a multimodal sensor device array (MSDA) which captures human activities and interactions within an environment;
an information storage device (ISD) which captures and cryptographically signs sensor information and cooperates with an interface to distribute the sensor information to at least one computer on a network that shares data and resources via wired or wireless technologies;
a methods control interface (MCI) that configures a plurality of modules;
a pattern recognition engine (PRE) analyzes data and categorizes information with metadata based on detected interactions and events; and
a synthetic data generator (SDG) utilizes a classification and categorization metadata created by the PRE to produce new data which is synthetic in nature but based on representative features present in end user data;
the SDG is configured to produce synthetic data and data generation models, which are leveraged by the DAR, the MSDA, the ISD, the MCI, and the PRE;
the SDG uses a Generative Adversarial Network (GAN) process to develop data creation and evaluation methods,
wherein the plurality of modules comprises the DAR, the MSDA, the ISD, the PRE, the SDG and the GAN.
2 . The method of claim 1 , further comprising
digital activity recording (DAR) software installed on at least one of the end user digital devices capturing, recording, and transmitting communications and usage patterns and activities of the end user.
3 . The method of claim 2 , wherein
the DAR software is equipped with a graphical user interface to configure the data from the device;
the DAR software is configured to connect to at least one of the end user devices or software applications and capture a variety of information;
the DAR software integrates with available operating system and an application program interface (APIs) to glean user data including but not limited to user inputs, application usage, system configuration settings and changes, file system changes, displayed data; and
the DAR-enabled devices gather essential data to form inferences into biographical history, the personality, and the psychological state of the end user through specific periods of time.
4 . The method of claim 3 , wherein
the data collected from the DAR, transmitted to the ISD and analyzed by the PRE creates key inferences that collectively illustrates aspects of the biographical and psychological state of the end user, including but not limited to the nature of the end user tasks, human interactions, focus, mood, attitude, level of agitation, and periodic device usage habits;
these inferences can be identified within a variety of data types;
inference metadata is generated within the PRE and augments all relevant data within the ISD; and
inference information can similarly be extracted from the data derived from the multimodal sensor device array (MSDA).
5 . The method of claim 4 , wherein
the DAR provides supporting data and metadata to enable supervised or semi-supervised learning within the pattern recognition engine (PRE);
the DAR-derived information assists significantly in the creation of accurately labeled datasets which supports the development of PRE models because much of the identified data is highly structured and pre-labeled by source, type, and context; and
the DAR software provides a mechanism for the end user to consolidate and centrally store a vast amount of biographical information.
6 . The method of claim 1 , wherein
a multimodal sensor device (MSD) is a digital hardware device possessing at least one microprocessor to assist in the gathering and digitization of sensory information;
each MSD contains at least one integrated or peripheral devices, typically sensors, which are responsible for collecting environment information and converting the collected environment information to either analogue or binary signals; and
a plurality of MSDs is configured in a multimodal sensor device array (MSDA) and function in parallel to capture data of at least one type or within at least one environment or geographical area.
7 . The method of claim 6 wherein,
the MSD applies numerous sensor-based data collection protocols;
audio data is gathered using a high-fidelity microphone;
motion data is gathered using an accelerometer, magnetometer, compass, gyroscope, or a combination of these motion sensors which can determine rotation, motion vectors, vibration, acceleration, or gross or fine movements;
time data is gathered using at least one internal clock within a microprocessor unit or externally functioning as a peripheral device;
touch sensor data is gathered using at least one capacitive touch sensor made available for human interaction;
proximity data is gathered using at least one optical time of flight sensor;
environmental data is gathered from a temperature sensor of both the multimodal device as well as the multimodal device environment; and
the temperature data can be used to determine the environment where the MSD is located and provides an indicator of the proper functioning of core features of the device.
8 . The method of claim 7 wherein,
the data captured by the MSDA is transmitted over a wired USB connection or wirelessly to the information storage device to be stored, indexed, and analyzed;
the data that is transmitted from the multimodal sensor may be encoded in a variety of industry standard or custom/proprietary formats;
the MSDA may adopt at least one standard or custom transmission protocols;
the MSDA contains software or firmware which analyzes the data on the device microprocessor prior to transmission providing an initial analysis of the gross or fine features that is recorded as metadata which is transmitted to the ISD; and
the operation of the MSDA is optimized by utilizing data variability analysis of the datasets collected;
wherein the multimodal sensor device captures several data streams, selected from the group consisting of audio, motion, image, network, and environmental data, each of these streams of data is analyzed for the relative variability of maximum and minimum ranges of the data within a period; and
the MSDA applies this analysis to optimize the flow of information that the MSDA transmits across the network, reduce a sample rate of the MSDA from sensors of the MSDA, or resume transmission when the variability of the data exceeds a specific threshold.
9 . The method of claim 1 , wherein
the information storage device (ISD) which is a physical device or virtual device connected to a local private network or a communication network;
a primary function of the ISD serves as a method of authenticating, receiving, storing, and retrieving data that is sent via the multimodal sensor device array and the digital activity recording software;
the ISD is configured via the methods control interface to work with a collection of multimodal sensors and to receive data from a variety of external computing devices that are equipped with the DAR software; and
through the MCI, the network is configured to add data from the end user from at least one different source.
10 . The method of claim 9 , wherein
the ISD stores the information the ISD receives using a combination of temporary volatile storage and long-term persistent storage;
the ISD utilizes at least one database software application, is configured to capture data relationally, as documents, or graphs;
the ISD implements a time-based index to all data received;
the machine learning models generated by the pattern recognition engine (PRE) are stored in the ISD;
the pattern recognition engine (PRE) is configured to use the ISD for storage of data that the PRE generates, including but not limited to metadata, entities, and networks of relationships derived by the PRE; and
the synthetic data and the data generation models generated by the synthetic data generator (SDG) are also stored in the ISD.
11 . The method of claim 1 , wherein
the methods control interface (MCI) is the primary means for the end user to visualize, interact with, and control the functioning of the plurality of modules, comprising the DAR, the MSDA, ISD, the PRE, the SDG, and the GAN;
the MCI interfaces with the APIs for each module to transmit configuration data to tune a function of each module based on the end user requirements;
the MCI is a virtual interface which may be accessible to any external digital device via a wired or wireless connection;
the MCI uploads and manages external datasets to support the pattern recognition including but not limited to reference data, pre-compiled neural network training datasets, and digital assets; and
the MCI configures and controls the modules within the method by the end user.
12 . The method of claim 1 , further comprising
a pattern recognition engine (PRE) that translates raw data within the information storage device (ISD) into structured metadata using a variety of algorithms and machine learning methods including neural networks;
the metadata generated by the PRE forms the basis for the categorization of data within the ISD for searching and interaction; the metadata being configured to serve as a support function for the synthetic data generator (SDG) and the generative adversarial network (GAN) processes contained therein;
the PRE provides essential metadata for configuring the outputs for the synthetic human voice and likeness engines by classifying data which can be used to tune the engine outputs;
the PRE processes data through the following steps:
the PRE generates and utilizes a large and diverse set of neural network models which are generated, trained, and evaluated regularly against the data in the ISD;
the PRE has access to numerous data types, including audio, motion, sensor, and end user-provided data;
the PRE applies multiple methods to evaluate data within the ISD and identify gross features within the data;
the PRE applies variable duration data sampling (VDDS) to identify features within specific amplitude, frequency, and duration subsets of sensor data; and
fine feature analysis is applied to further classify gross features down to specific events and translate the data within those designated time-ranges into relatively accurate and complete metadata to support biographical history narratives.
13 . The method of claim 12 , wherein the multiple methods used by the PRE to evaluate data within the ISD and identify gross features within the data include,
gross feature categorization of the PRE applies both fixed algorithms as well as pre-trained neural networks to perform this categorization;
gross feature analysis is applied to extract multiple features within the same dataset by applying different algorithms or trained neural networks; and
algorithms establish logical boundaries which can accurately determine logically analyzed parameters.
14 . The method of claim 12 , wherein
the PRE is configured to analyze data again after PRE models have been updated or enhanced through retraining or adjustments of algorithmic threshold bounds of the PRE.
15 . The method of claim 12 , further comprising
the cyclical analysis of data using standard time cycles (days, months years) to develop biological narratives and identify noteworthy or novel events or periods of significant divergence from patterns of predicted events; and
in addition to recognizing features directly within the data from the MSDAs, the DAR software, or any additional data added via the MCI interface, the PRE also detects features and patterns within the metadata that the PRE generates for the purpose of establishing biological rhythms and patterns of the end user.
16 . The method of claim 12 , wherein
the PRE is configured to train and operationalize unsupervised, semi-supervised, and fully supervised neural networks;
with supervised learning, training data is pre-classified into datasets which are used to develop a model to recognize data and to classify the data into one of these known categories;
with unsupervised learning, all data exists within the same dataset, and the model develops classification methods based on the features that the model identifies within the data; and
the classifications that are created with the model can then be identified when compared to another labeled dataset or by a human capable of recognizing and naming the feature that has been classified;
semi-supervised learning provides a small amount of labeled training data with a large amount of unlabeled training data;
in the case of unsupervised or semi-supervised learning, the PRE may refer data samples to the end user for adjudication via the MCI interface; and,
pushing a set of data to the end user, the PRE may ask the user to enter the label for the features identified within the dataset, or the PRE may compare this dataset with a labeled dataset and ask the end user to confirm the end user alignment;
wherein a user-confirmed alignment provides validation that the model is correct, and the interactions with the end user through the MCI significantly enhance the learning outcomes for the PRE.
17 . The method of claim 12 , further comprising
feature recognition and the labelling of such features form the basis of indexing but also form the linguistic and textual basis for forming language-based narratives and enabling a personality simulation engine to produce language-based descriptions of biographical events.
18 . The method of claim 12 , further comprising
the PRE applies predictive modelling to determine the most likely event to happen next within a dataset or anticipate the likelihood of features that have not yet been recorded;
if the PRE predicts the end user will perform an action, and the end user performs the action, confidence in the models increases; and
if the PRE predicts an event which does not occur, and the predictions trend poorly despite the frequent retraining of the predictive models, this prediction may indicate that the method does not have enough data yet to make predictions or that the end user actions are relatively unpredictable by nature.
19 . The method of claim 17 , further comprising
trend analysis of preference and bias over time based on frequency of matched or related events;
the PRE analyzes the actions and habit change based on the frequency or amplitude of related feature occurrence within at least one dataset within a specific time period;
by analyzing the patterns of experiences of the end user, and the duration of time that these experiences repeat and persist, the method is configured to develop a bias model which demonstrates persistent verses transitory interest and actions.
20 . The method of claim 1 , wherein
after the GAN is trained to discern real data from generated data, the SDG is configured to evaluate data within the ISD which is generated by the SDG;
based on the classification of training data of the GAN, the SDG searches the ISD for additional data whose metadata aligns to the training dataset;
when additional data is found, the Discriminator checks the additional data against the Generator to see how similar the datasets are; and
if an output of the Generator exceeds a threshold of confidence compared with the original data, the SDG may trigger the ISD to discard the original data.
21 . The method of claim 20 , wherein
the experientiality of historic data can be recreated through synthetic processes of the SDG, even if the original data is no longer present;
if the original data has been disposed, the synthetic data is used to recreate the specific experience;
the method is configured to use the MCI with the end user input to identify the degrees of disposition of data;
the degree of disposition is related to the confidence threshold of the Discriminator, where only data with a high degree of similarity to the synthesized data, above a high threshold of 0.99 confidence is disposed; and
if those thresholds are lowered, the amount of data disposed will be increased and storage optimized.
22 . The method of claim 21 , wherein
the end user selects a sloped threshold, wherein data that has been recorded most recently, has a high threshold but older data may have a lower threshold for disposition;
the slopes can take any geometric form, linear, sinusoidal, curved, stepped, or user defined forms.