IP Library Granted Patent US 11,893,099
Granted Patent B2
US 11,893,099 · App. 17/898,463 · Granted Feb 6, 2024

Systems and methods for dynamic passphrases

Inventors: Edison U. Ortiz (Toronto, CA); Mohammad Abuzar Shaikh (Toronto, CA); Margaret Inez Salter (Toronto, CA); Sarah Rachel Waigh Yean Wilkinson (Toronto, CA); Arya Pourtabatabaie (Toronto, CA); Iustina-Miruna Vintila (Bucharest, RO); Steven Fernandes (Omaha, NE); Sumit Kumar Jha (San Antonio, TX)
Assignee: ROYAL BANK OF CANADA
G06F21/32G06F18/22G06F18/24137G06F21/46G06N5/04G06N20/00G06Q20/108G06V10/17G06V10/454G06V10/764G06V10/771G06V10/776G06V10/803G06V10/82G06V40/165G06V40/171G06V40/172G06V40/20G06V40/40G10L15/02G10L15/25G06F2221/2103G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,099
App. No.
17/898,463
Granted
Feb 6, 2024
Kind
B2
Abstract

A technical validation mechanism is described that includes the use of facial feature recognition and tokenization technology operating in combination with machine learning models can be used such that specific facial or auditory characteristics of how an originating script is effectuated can be used to train the machine learning models, which can then be used to validate a video or a particular dynamically generated passphrase by comparing overlapping phonemes or phoneme transitions between the originating script and the dynamically generated passphrase.

Claims (38)

1. A computer system for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the computer system comprising a non-transitory computer readable storage device, computer memory, and a processor configured to:

receive a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;

extract, from the script-reading video data set, a data subset representing the one or more facial or lip features of the individual corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;

train, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;

receive an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and

process, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;

wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.

2. The system of claim 1 , wherein the script data structure includes words corresponding to a phonetic pangram or a holo-alphabetic sentence.

3. The system of claim 2 , wherein the phonetic pangram or the holo-alphabetic sentence include repeated phoneme or phoneme transition portions to provide additional data points for training the one or more baseline machine learning data model architectures.

4. The system of claim 1 , wherein the dynamic passphrase challenge is conducted on a graphical user interface where a statement string portion and a question string portion are displayed as textual display elements on a computer display.

5. The system of claim 1 , wherein the secure electronic resource is a secure webpage.

6. The system of claim 5 , wherein the secure webpage is an online banking website.

7. The system of claim 1 , wherein the correct response string is not directly stated in the words corresponding to the script data structure.

8. The system of claim 1 , wherein the words corresponding to the script data structure is provided in the form of a contextual question to be answered.

9. The system of claim 1 , wherein the one or more facial or lip features are extracted from a video having a time-stamped audio and video track.

10. The system of claim 1 , wherein the correct response string is selected to include dictionary words based on the trained machine learning models trained above a threshold confidence level.

11. A method for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the method comprising:

receiving a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;

extracting, from the script-reading video data set, a data subset representing the one or more facial or lip features of the individual corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;

training, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;

receiving an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and

processing, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;

wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.

12. The method of claim 11 , wherein the script data structure includes words corresponding to a phonetic pangram or a holo-alphabetic sentence.

13. The method of claim 12 , wherein the phonetic pangram or the holo-alphabetic sentence include repeated phoneme or phoneme transition portions to provide additional data points for training the one or more baseline machine learning data model architectures.

14. The method of claim 11 , wherein the dynamic passphrase challenge is conducted on a graphical user interface where a statement string portion and a question string portion are displayed as textual display elements on a computer display.

15. The method of claim 11 , wherein the secure electronic resource is a secure webpage.

16. The method of claim 15 , wherein the secure webpage is an online banking website.

17. The method of claim 11 , wherein the correct response string is not directly stated in the words corresponding to the script data structure.

18. The method of claim 11 , wherein the words corresponding to the script data structure is provided in the form of a contextual question to be answered.

19. The method of claim 11 , wherein the one or more facial or lip features are extracted from a video having a time-stamped audio and video track.

20. A non-transitory computer readable medium storing computer interpretable instructions, which when executed by a processor, cause the processor to perform a method for conducting a dynamic passphrase challenge to control access to a secure electronic resource, the method comprising:

receiving a script-reading video data set capturing a portion of or an entirety of an individual's face while the individual is speaking words corresponding to a script data structure, the script data structure having a sequence of pre-identified phonemes or phoneme transitions, the pre-identified phonemes or phoneme transitions including at least one overlapping phoneme or phoneme transition required to be spoken when speaking words of a correct response string;

extracting, from the script-reading video data set, a data subset representing the one or more facial or lip features of the corresponding to each phoneme or phoneme transition corresponding to the sequence of pre-identified phonemes or phoneme transitions;

training, one or more baseline machine learning data model architectures, each baseline machine learning data model architecture of the one or more baseline machine learning data model architectures corresponding to a corresponding pre-identified phoneme or phoneme transition of the script data structure such that parameters of the baseline machine learning data model architectures are tuned based on the corresponding one or more facial or lip features;

receiving an answer-reading video data set capturing a portion of or an entirety of the individual's face while the individual is speaking the words corresponding to the correct response string; and

processing, the answer-reading video data set, using the one or more baseline machine learning data model architectures corresponding to the at least one overlapping phoneme or phoneme transition to determine an overall classification similarity score;

wherein a provisioning of access to the secure electronic resource only occurs if the overall classification similarity score is greater than a pre-defined threshold similarity score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: ORTIZ, EDISON U.; SHAIKH, MOHAMMAD ABUZAR; SALTER, MARGARET INEZ; WILKINSON, SARAH RACHEL WAIGH YEAN; POURTABATABAIE, ARYA; VINTILA, IUSTINA-MIRUNA; FERNANDES, STEVEN; JHA, SUMIT KUMAR
To: ROYAL BANK OF CANADA
Reel/Frame 066156/0022 →
Continuity (9)
Continuation 17129631 · Dec 21, 2020
Continuation In Part 16521238 · Jul 24, 2019
Provisional Application 62951528 · Dec 20, 2019
Provisional Application 62839384 · Apr 26, 2019
Provisional Application 62775695 · Dec 5, 2018
Provisional Application 62774130 · Nov 30, 2018
Provisional Application 62751369 · Oct 26, 2018
Provisional Application 62702635 · Jul 24, 2018
Related Publication 20220405381A1 · Dec 22, 2022
Cited By (1)
US 12,563,053