Systems and methods for dynamic passphrases
A technical validation mechanism is described that includes the use of facial feature recognition and tokenization technology operating in combination with machine learning models can be used such that specific facial or auditory characteristics of how an originating script is effectuated can be used to train the machine learning models, which can then be used to validate a video or a particular dynamically generated passphrase by comparing overlapping phonemes or phoneme transitions between the originating script and the dynamically generated passphrase.
1. A computer system for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the computer system comprising a non-transitory computer readable storage device, computer memory, and a processor configured to:
receive a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;
extract, from the script-reading video data set, a data subset representing the one or more facial or lip features of the individual corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;
train, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;
receive an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and
process, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;
wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.
2. The system of claim 1 , wherein the script data structure includes words corresponding to a phonetic pangram or a holo-alphabetic sentence.
3. The system of claim 2 , wherein the phonetic pangram or the holo-alphabetic sentence include repeated phoneme or phoneme transition portions to provide additional data points for training the one or more baseline machine learning data model architectures.
4. The system of claim 1 , wherein the dynamic passphrase challenge is conducted on a graphical user interface where a statement string portion and a question string portion are displayed as textual display elements on a computer display.
5. The system of claim 1 , wherein the secure electronic resource is a secure webpage.
6. The system of claim 5 , wherein the secure webpage is an online banking website.
7. The system of claim 1 , wherein the correct response string is not directly stated in the words corresponding to the script data structure.
8. The system of claim 1 , wherein the words corresponding to the script data structure is provided in the form of a contextual question to be answered.
9. The system of claim 1 , wherein the one or more facial or lip features are extracted from a video having a time-stamped audio and video track.
10. The system of claim 1 , wherein the correct response string is selected to include dictionary words based on the trained machine learning models trained above a threshold confidence level.
11. A method for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the method comprising:
receiving a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;
extracting, from the script-reading video data set, a data subset representing the one or more facial or lip features of the individual corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;
training, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;
receiving an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and
processing, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;
wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.
12. The method of claim 11 , wherein the script data structure includes words corresponding to a phonetic pangram or a holo-alphabetic sentence.
13. The method of claim 12 , wherein the phonetic pangram or the holo-alphabetic sentence include repeated phoneme or phoneme transition portions to provide additional data points for training the one or more baseline machine learning data model architectures.
14. The method of claim 11 , wherein the dynamic passphrase challenge is conducted on a graphical user interface where a statement string portion and a question string portion are displayed as textual display elements on a computer display.
15. The method of claim 11 , wherein the secure electronic resource is a secure webpage.
16. The method of claim 15 , wherein the secure webpage is an online banking website.
17. The method of claim 11 , wherein the correct response string is not directly stated in the words corresponding to the script data structure.
18. The method of claim 11 , wherein the words corresponding to the script data structure is provided in the form of a contextual question to be answered.
19. The method of claim 11 , wherein the one or more facial or lip features are extracted from a video having a time-stamped audio and video track.
20. A non-transitory computer readable medium storing computer interpretable instructions, which when executed by a processor, cause the processor to perform a method for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the method comprising:
receiving a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;
extracting, from the script-reading video data set, a data subset representing the one or more facial or lip features of the corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;
training, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;
receiving an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and
processing, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;
wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.