IP Library › Granted Patent US 12,462,025
Granted Patent B1
US 12,462,025 · App. 19/013,980 · Granted Nov 4, 2025

Multimodal data fusion for cybersecurity applications

Inventor: Kubashen Jerome Naidoo (Heath, TX)
Assignee: 4MindsAI Inc.
G06F21/554G06F2221/034G06F2221/2101
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,025
App. No.
19/013,980
Filed
Jan 8, 2025
Granted
Nov 4, 2025
Kind
B1
Art Unit
2494
USPC
726/23
Abstract

A method for detecting unauthorized access attempts to a digital system using multi-modal data and neural networks. The method involves receiving multi-modal data indicative of an access attempt and processing different portions of this data using modality-specific layers of a neural network to generate corresponding embedding vectors. Custom embedding vectors are then generated for each data portion using respective neural network layers. These custom embedding vectors are combined into a single embedding vector through a fusion layer, with all layers being jointly trained based on a common loss function. The combined embedding vector is processed by a trained model to determine whether the access attempt is unauthorized.

Claims (44)

1 . A computer-implemented method for detecting cybersecurity threats, the method comprising:

receiving multi-modal data, each modality of the multi-modal data representing an attempt to access a digital system;

processing a first portion of the multi-modal data of a first modality using a first modality-specific layer of a neural network to generate a first embedding vector representing features of the first portion of the multi-modal data;

processing at least a second portion of the multi-modal data of a second modality—different from the first modality—using a second modality-specific layer of the neural network to generate a second embedding vector representing features of the second portion of the multi-modal data;

generating a custom embedding vector for each of the first and second portions of the multi-modal data from the first embedding vector and the second embedding vector, respectively, using corresponding neural network layers;

generating, from the multiple custom embedding vectors, a combined embedding vector using a fusion layer of the neural network,

wherein the corresponding neural network layers that generate the custom embedding vectors and the fusion layer are jointly trained based on a common loss function; and

processing the combined embedding vector using a trained model to generate an indication whether or not the access attempt to the digital system is unauthorized.

2 . The method of claim 1 , wherein each of the first modality and the second modality comprises one of: text, images, behavioral data, or biometrics data.

3 . The method of claim 2 , wherein the behavioral data comprises location data corresponding to the access attempt.

4 . The method of claim 2 , wherein the behavioral data comprises data representing user-interactions with one or more input devices corresponding to the access attempt.

5 . The method of claim 1 , wherein the first modality comprises one of: text or images, and the first modality-specific layer of the neural network comprises a layer of a convolutional neural network (CNN).

6 . The method of claim 1 , wherein the second modality comprises time series data, and the second modality-specific layer of the neural network is a layer of a recurrent neural network (RNN).

7 . The method of claim 1 , wherein each of the corresponding neural network layers that are jointly trained based on the common loss function is configured to transform the corresponding one of the first embedding vector or the second embedding vector into a shared embedding space.

8 . The method of claim 1 , wherein generating the combined embedding vector comprises:

combining the custom embedding vectors in a weighted combination within the fusion layer of the neural network.

9 . A system comprising:

one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving multi-modal data, each modality of the multi-modal data used in an access-attempt to a digital system;

processing a first portion of the multi-modal data of a first modality using a first modality-specific layer of a neural network to generate a first embedding vector representing features of the first portion of the multi-modal data;

processing at least a second portion of the multi-modal data of a second modality—different from the first modality—using a second modality-specific layer of the neural network to generate a second embedding vector representing features of the second portion of the multi-modal data;

generating a custom embedding vector for each of the multi-modal data from the first embedding vector and the second embedding vector using corresponding neural network layers;

generating, from the multiple custom embedding vectors, a combined embedding vector using a fusion layer of the neural network,

wherein the corresponding neural network layers that generate the custom embedding vectors and the fusion layer are jointly trained based on a common loss function; and

processing the combined embedding vector using a trained model to generate an indication whether or not the access attempt to the digital system is unauthorized.

10 . The system of claim 9 , wherein each of the first modality and the second modality comprises one of: text, images, biometrics data, or behavioral data comprising one or more of location data corresponding to the access attempt and data representing user-interactions with one or more input devices corresponding to the access attempt.

11 . The system of claim 9 , wherein the first modality comprises one of: text or images, and the first modality-specific layer of the neural network comprises a layer of a convolutional neural network (CNN).

12 . The system of claim 9 , wherein the second modality comprises time series data, and the second modality-specific layer of the neural network is a layer of a recurrent neural network (RNN).

13 . The system of claim 9 , wherein each of the corresponding neural network layers that are jointly trained based on the common loss function is configured to transform the corresponding one of the first embedding vector or the second embedding vector into a shared embedding space.

14 . The system of claim 9 , wherein generating the combined embedding vector comprises:

combining the custom embedding vectors in a weighted combination within the fusion layer of the neural network.

15 . One or more non-transitory computer-readable storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising: receiving multi-modal data, each modality of the multi-modal data used in an access-attempt to a digital system;

processing a first portion of the multi-modal data of a first modality using a first modality-specific layer of a neural network to generate a first embedding vector representing features of the first portion of the multi-modal data;

processing at least a second portion of the multi-modal data of a second modality—different from the first modality—using a second modality-specific layer of the neural network to generate a second embedding vector representing features of the second portion of the multi-modal data;

generating a custom embedding vector for each of the multi-modal data from the first embedding vector and the second embedding vector using corresponding neural network layers;

generating, from the multiple custom embedding vectors, a combined embedding vector using a fusion layer of the neural network,

wherein the corresponding neural network layers that generate the custom embedding vectors and the fusion layer are jointly trained based on a common loss function; and

processing the combined embedding vector using a trained model to generate an indication whether or not the access-attempt to the digital system is unauthorized.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each of the first modality and the second modality comprises one of: text, images, biometrics data, or behavioral data comprising one or more of location data corresponding to the access-attempt and data representing user-interactions with one or more input devices corresponding to the access-attempt.

17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the first modality comprises one of: text or images, and the first modality-specific layer of the neural network comprises a layer of a convolutional neural network (CNN).

18 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the second modality comprises time series data, and the second modality-specific layer of the neural network is a layer of a recurrent neural network (RNN).

19 . The one or more non-transitory computer-readable storage media of claim 15 , wherein each of the corresponding neural network layers that are jointly trained based on the common loss function is configured to transform the corresponding one of the first embedding vector or the second embedding vector into a shared embedding space.

20 . The one or more non-transitory computer-readable storage media of claim 15 , wherein generating the combined embedding vector comprises:

combining the custom embedding vectors in a weighted combination within the fusion layer of the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2025
From: NAIDOO, KUBASHEN JEROME
To: 4MINDSAI INC.
Reel/Frame 069856/0602 →
Continuity (1)
Provisional Application 63724030 · Nov 22, 2024
References Cited (15)
US 10521587B1 · Agranonik · 2019 [cited by examiner]
US 10699002B1 · Walters · 2020 [cited by examiner]
US 11550804B1 · Satuluri · 2023 [cited by examiner]
US 12028374B2 · Soryal et al. · 2024 [cited by applicant]
US 20190258807A1 · DiMaggio · 2019 [cited by examiner]
US 20200097653A1 · Mehta · 2020 [cited by examiner]
US 20210157945A1 · Cobb · 2021 [cited by examiner]
US 20230139161A1 · Tormasov · 2023 [cited by examiner]
US 20230281298A1 · Molloy · 2023 [cited by examiner]
US 20240143744A1 · Tiwari et al. · 2024 [cited by applicant]
US 20240330446A1 · Bulut et al. · 2024 [cited by applicant]
US 20240427879A1 · Bulut et al. · 2024 [cited by applicant]
US 20250005149A1 · Nandi et al. · 2025 [cited by applicant]
WO WO2024145209A1 · 2024 [cited by applicant]
WO WO2024226801A2 · 2024 [cited by applicant]