IP Library › Granted Patent US 12,536,260
Granted Patent B2
US 12,536,260 · App. 17/938,431 · Granted Jan 27, 2026

System, apparatus, and method for automatically generating negative keystroke examples and training user identification models based on keystroke dynamics

Inventors: Ionut Dumitran (Bucharest, RO); Radu Tudor Ionescu (Bucharest, RO); Florinel-Alin Croitoru (Bucharest, RO); Cristina Mǎdǎlina Noaica (Bucharest, RO)
Assignee: VERIDIUM IP LIMITED
G06F21/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,260
App. No.
17/938,431
Granted
Jan 27, 2026
Kind
B2
Abstract

An apparatus adapted to identify a user based on keystroke dynamics of an input by the user, the apparatus adapted to: execute a first training phase of training a keystroke sample generator to generate negative keystroke samples; execute a second training phase of training a user identification model based at least in part on a plurality of negative keystroke samples generated using the keystroke sample generator; and execute a deployment of the user identification model to authenticate an input sample associated with the user using the trained user identification model.

Claims (103)

1 . An apparatus adapted to identify a user based on keystroke dynamics of an input by the user, comprising:

a communication interface to one or more networks;

one or more processing devices operatively connected to the computer network interface; and

one or more memory storage devices operatively connected to the one or more processing devices and having stored thereon machine-readable instructions that, when executed, cause the one or more processing devices to:

in a first training phase of training a keystroke sample generator,

receive, via the communication interface, a plurality of first text input samples by a plurality of first users; and

for each received first text input sample:

generate a user identification representation, a noise representation, and a text sequence representation of the received first text input sample;

generate, using the keystroke sample generator, a keystroke sequence sample based on a combination of the generated user identification representation, noise representation, and text sequence representation;

input the generated keystroke sequence sample and an actual keystroke sequence of the received first text input sample to a classifier for a user identification classification on the generated keystroke sequence sample;

input the generated keystroke sequence sample and the actual keystroke sequence of the received first text input sample to a regressor for a text character length regression on the generated keystroke sequence sample; and

train the keystroke sample generator based on the user identification classification of the classifier and the text character length classification of the regressor;

in a second training phase of training a user identification model,

receive, via the communication interface, one or more second text input samples by a second user, said second user being different from the plurality of first users;

generate, using the keystroke sample generator, a plurality of negative keystroke samples based on the one or more second text input samples; and

train the user identification model on a user classification of the second user based on the one or more second text input samples and the generated plurality of negative keystroke samples; and

in a deployment of the user identification model,

receive, via the communication interface, a third text input sample in association with the second user; and

authenticate the third text input sample using the trained user identification model.

2 . The apparatus of claim 1 , wherein the user identification representation and the text sequence representation are generated using respective embedding neural layers.

3 . The apparatus of claim 1 , wherein the user identification representation and the noise representation are generated to conform to a format of the text sequence representation using respective neural layers.

4 . The apparatus of claim 1 , wherein the classifier is a multi-class discriminator embodied by a neural network and the regressor is a regression neural network.

5 . The apparatus of claim 4 , wherein the training of the keystroke sample generator, the user identification classification by the classifier, and the text character length regression by the regressor are based on

L ( G,D,R,x,u,t )= L GAN ( G,D,x,u,t )+λ 1 ·L time ( G,u,t )+λ 2 ·L MSE ( G,R,x,t )

where:

G represents the keystroke sample generator,

D represents the multi-class discriminator,

R represents the regressor,

x represents an array of press and release timestamps for a typed text sequence of the received first text input sample,

t represents the typed text sequence of the received first text input sample,

u represents a one-hot encoded vector representing an ID associated with one of the plurality of first users that typed the received first text input sample or a label indicating that an input sample to the classifier is generated,

λ 1 is a hyperparameter that controls an importance of a temporal consistency loss (denoted as L time ),

λ 2 is a hyperparameter that controls an importance of a mean square error (MSE) loss (denoted as L MSE ) with respect to the text character length of the generated keystroke sequence sample,

L GAN ( G,D,x,u,t )= E x˜p data (x) [−Σ i=1 k u i ·log( D ( x ))]+ E z˜p z (z) [−Σ i=1 m u i ·log( D ( G ( z|u,t )))]

E represents an expected value ((average) of each formula included in square brackets),

p data represents a probability distribution of data,

x is a data point sampled from the distribution p data ,

E x˜p data (x) is the expected value over all data points,

p z represents a noise distribution,

z is a noise vector sampled from the distribution p z ,

E z˜p z (z) is the expected value over all noise vectors,

m represents a number of the plurality of first users,

k represents a number of the plurality of second users,

L time ( G,u,t )= E z˜p z ( z )[Σ i=1 n max(0, y i,0 −y i,1 )+Σ i=1 n−1 max(0, y i,0 −y i+1,0 )]

y=G(z|u,t), therefore y i,0 and y i,1 represent press and release timestamps of an i-th key in the sequence t,

n represents a sequence length, and

L MSE ( G,R,x,t )= E x˜p data (x) [( R ( x )− len ( t )) 2 ]+E z˜p z (z) [( R ( G ( z|u,t ))− len ( t )) 2 ]

len is a function that returns a length of a sequence t.

6 . The apparatus of claim 1 , wherein the plurality of negative keystroke samples are generated by the generator based on one or more of the plurality of first text input samples in association with one or more of the plurality of first users, which are different from the second user.

7 . The apparatus of claim 6 , wherein at least one of the negative keystroke samples is generated based on one of the plurality of first text input samples that comprises a same character sequence as the one or more second text input samples.

8 . The apparatus of claim 1 , wherein the user identification model is a binary classifier that determines whether a keystroke sequence of the third text input sample corresponds to the second user based on the user classification training.

9 . The apparatus of claim 1 , wherein the plurality of first text input samples comprise free text inputs by the plurality of first users.

10 . The apparatus of claim 1 , wherein the one or more second text input samples comprise a fixed text input by the second user.

11 . A method for identifying a user based on keystroke dynamics of an input by the user, comprising:

in a first training phase of training a keystroke sample generator,

receiving, by a processing apparatus via a communication interface, a plurality of first text input samples by a plurality of first users; and

for each received first text input sample:

generating, by the processing apparatus, a user identification representation, a noise representation, and a text sequence representation of the received first text input sample;

generating, by the processing apparatus using the keystroke sample generator, a keystroke sequence sample based on a combination of the generated user identification representation, noise representation, and text sequence representation;

inputting, by the processing apparatus, the generated keystroke sequence sample and an actual keystroke sequence of the received first text input sample to a classifier for a user identification classification on the generated keystroke sequence sample;

inputting, by the processing apparatus, the generated keystroke sequence sample and the actual keystroke sequence of the received first text input sample to a regressor for a text character length regression on the generated keystroke sequence sample; and

training, by the processing apparatus, the keystroke sample generator based on the user identification classification of the classifier and the text character length classification of the regressor;

in a second training phase of training a user identification model,

receiving, by the processing apparatus via the communication interface, one or more second text input samples by a second user, said second user being different from the plurality of first users;

generating, by the processing apparatus using the keystroke sample generator, a plurality of negative keystroke samples based on the one or more second text input samples; and

training, by the processing apparatus, the user identification model on a user classification of the second user based on the one or more second text input samples and the generated plurality of negative keystroke samples; and

in a deployment of the user identification model,

receiving, by the processing apparatus via the communication interface, a third text input sample in association with the second user; and

authenticating, by the processing apparatus, the third text input sample using the trained user identification model.

12 . The method of claim 11 , wherein the user identification representation and the text sequence representation are generated using respective embedding neural layers.

13 . The method of claim 11 , wherein the user identification representation and the noise representation are generated to conform to a format of the text sequence representation using respective neural layers.

14 . The method of claim 11 , wherein the classifier is a multi-class discriminator embodied by a neural network and the regressor is a regression neural network.

15 . The method of claim 14 , wherein the training of the keystroke sample generator, the user identification classification by the classifier, and the text character length regression by the regressor are based on

L ( G,D,R,x,u,t )= L GAN ( G,D,x,u,t )+λ 1 ·L time ( G,u,t )+λ 2 ·L MSE ( G,R,x,t )

where:

G represents the keystroke sample generator,

D represents the multi-class discriminator,

R represents the regressor,

x represents an array of press and release timestamps for a typed text sequence of the received first text input sample,

t represents the typed text sequence of the received first text input sample,

u represents a one-hot encoded vector representing an ID associated with one of the plurality of first users that typed the received first text input sample or a label indicating that an input sample to the classifier is generated,

λ 1 is a hyperparameter that controls an importance of a temporal consistency loss (denoted as L time ),

λ 2 is a hyperparameter that controls an importance of a mean square error (MSE) loss (denoted as L MSE ) with respect to the text character length of the generated keystroke sequence sample,

L GAN ( G,D,x,u,t )= E x˜p data (x) [−Σ i=1 k u i ·log( D ( x ))]+ E z˜p z (z) [−Σ i=1 m u i ·log( D ( G ( z|u,t )))]

E represents an expected value ((average) of each formula included in square brackets),

p data represents a probability distribution of data,

x is a data point sampled from the distribution p data ,

E x˜p data (x) is the expected value over all data points,

p z represents a noise distribution,

z is a noise vector sampled from the distribution p z ,

E z˜p z (z) is the expected value over all noise vectors,

m represents a number of the plurality of first users,

k represents a number of the plurality of second users,

L time ( G,u,t )= E z˜p z (z) [Σ i=1 n max(0, y i,0 −y i,1 )+Σ i=1 n−1 max(0, y i,0 −y i+1,0 )]

y=G(z|u,t), therefore y i,0 and y i,1 represent press and release timestamps of an i-th key in the sequence t,

n represents a sequence length, and

L MSE ( G,R,x,t )= E x˜p data (x) [( R ( x )− len ( t )) 2 ]+E z˜p z (z) [( R ( G ( z|u,t ))− len ( t ) 2 ]

len is a function that returns a length of a sequence t.

16 . The method of claim 11 , wherein the plurality of negative keystroke samples are generated by the generator based on one or more of the plurality of first text input samples in association with one or more of the plurality of first users, which are different from the second user.

17 . The method of claim 16 , wherein at least one of the negative keystroke samples is generated based on one of the plurality of first text input samples that comprises a same character sequence as the one or more second text input samples.

18 . The method of claim 11 , wherein the user identification model is a binary classifier that determines whether a keystroke sequence of the third text input sample corresponds to the second user based on the user classification training.

19 . The method of claim 11 , wherein the plurality of first text input samples comprise free text inputs by the plurality of first users.

20 . The method of claim 11 , wherein the one or more second text input samples comprise a fixed text input by the second user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2022
From: DUMITRAN, IONUT; IONESCU, RADU TUDOR; CROITORU, FLORINEL-ALIN; NOAICA, CRISTINA MADALINA
To: VERIDIUM IP LIMITED
Reel/Frame 061336/0595 →
Continuity (1)
Related Publication 20240134949A1 · Apr 25, 2024
References Cited (22)
US 10693661B1 · Hamlet · 2020 [cited by examiner]
US 11449746B2 · Baldwin · 2022 [cited by examiner]
US 20190332876A1 · Khitrov · 2019 [cited by examiner]
US 20210216845A1 · Walters · 2021 [cited by examiner]
Fan, Zijian. “Applying generative adversarial networks for the generation of adversarial attacks against continuous authentication.” (Year: 2020). [cited by examiner]
Jinyin Chen, Yangyang Wu, Chengyu Jia, Haibin Zheng, Guohan Huang, Customizable text generation via conditional text generative adversarial network, Neurocomputing, vol. 416, 2020, pp. 125-135, ISSN 0925-2312, https://d… [cited by examiner]
Conijn, R., Cook, C., van Zaanen, M et al. Early prediction of writing quality using keystroke logging https://doi.org/10.1007/s40593-021-00268-w (Year: 2021). [cited by examiner]
Zhang, Yizhe, Zhe Gan, and Lawrence Carin. “Generating text via adversarial training.” NIPS workshop on Adversarial Training. vol. 21. Academia. edu, 2016. (Year: 2016). [cited by examiner]
Chang et al. “Machine Learning-Based Analysis of Free-Text Keystroke Dynamics” [retrieved on Nov. 8, 2025] Retrieved from the Internet <URL:https://arxiv.org/abs/2107.07409> (Year: 2021). [cited by examiner]
Bitvai, Zsolt, and Trevor Cohn. “Non-linear text regression with a deep convolutional neural network.” Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Jo… [cited by examiner]
D. Deb and M. M. Guirguis, “Use of Auxiliary Classifier Generative Adversarial Network in Touchstroke Authentication,” 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA), Miami, FL, USA… [cited by examiner]
Antal, Margit and Lehel Nemes. “The MOBIKEY Keystroke Dynamics Password Database: Benchmark Results.” Advances in Intelligent Systems and Computing, Computer Science On-line Conference (2016), vol. 465, pp. 35-46. [cited by applicant]
Migdal, Denis and Christophe Rosenberger. “Statistical modeling of keystroke dynamics samples for the generation of synthetic datasets.” Future Gener. Comput. Syst. 100 (2019): 907-920. [cited by applicant]
González, Nahuel et al. “Towards liveness detection in keystroke dynamics: Revealing synthetic forgeries.” Systems and Soft Computing (2022): vol. 4, 200037: pp. 1-12. [cited by applicant]
Monaco, John V. et al. “Spoofing key-press latencies with a generative keystroke dynamics model.” 2015 IEEE 7th International Conference on Biometrics Theory, Applications and Systems (BTAS) (2015): 1-8. [cited by applicant]
Huster, Todd P. et al. “Pareto GAN: Extending the Representational Power of GANs to Heavy-Tailed Distributions.” International Conference on Machine Learning (2021), pp. 4523-4532. [cited by applicant]
Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. and Bengio, Y. (2014) Generative Adversarial Nets. Proceedings of the 27th International Conference on Neural Informatio… [cited by applicant]
Bengio, Yoshua, et al. “Curriculum learning.” Proceedings of the 26th Annual International Conference on Machine earning. 2009, pp. 41-48. [cited by applicant]
Park, Youngja, et al. “Learning from Others: User Anomaly Detection Using Anomalous Samples from Other Users”, 18th International Conference, Austin, TX, USA, Lecture Notes in Computer Science, pp. 396-414. Sep. 21, 201… [cited by applicant]
Buriro, Attaullah et al. “SWIPEGAN: Swiping Data Augmentation Using Generative Adversarial Networks for Smartphone User Authentication.” Proceedings of the 3rd ACM Workshop on Wireless Security and Machine Learning (202… [cited by applicant]
Acien, Alejandro et al. “TypeNet: Deep Learning Keystroke Biometrics.” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 4 No. 1 (2022): 57-70. [cited by applicant]
International Search Report and Written Opinion in PCT Application No. PCT/US2023/073965, mailed Dec. 4, 2023. [cited by applicant]