Network security for machine learning models including large language models (LLMS)
There is provided a system for securing an AI-application guided by a LLM, comprising: a processor executing a code for: receiving a session of an interaction by an entity with the LLM-guided AI-application, the session including a plurality of output tokens generated by the AI-application in response to a plurality of input tokens inputted into the AI-application by the entity and/or from at least one external resource, dynamically during processing of the session by the LLM, tracing control and/or data-flow within attention-heads of transformer layers of the LLM with respect to the session, to identify at least one output token of the plurality of output tokens logically stemming from at least one input token of the plurality of input tokens originating from the at least one external resource, and triggering a security action for preventing the at least one external resource from feeding adversarial input into the AI-application.
1 . A system for securing an AI-application guided by a large language model (LLM), comprising:
a memory storing a code;
at least one hardware processor operatively coupled to the data interface and to the memory, for executing the code comprising instructions for:
generating monitored logs of system calls and library calls captured from the LLM and/or AI-application during inference using a preselected input,
wherein the LLM processes input tokens, serves as a decision-making center that decides what actions to take next, and generates output tokens, and wherein the AI-application interfaces with an environment and executes the LLM's decisions;
filtering the monitored logs with respect to clean logs generated by monitoring a clean version of the LLM and/or AI-application, to identify a subset of logs that do not exist in the clean logs;
triggering a security action in response to detecting the subset of logs;
receiving a session of an interaction by an entity with the LLM-guided AI-application, the session including a plurality of output tokens generated by the AI-application in response to a plurality of input tokens inputted into the AI-application by the entity and/or from at least one external resource;
dynamically during processing of the session by the LLM, tracing control and/or data-flow within attention-heads of transformer layers of the LLM with respect to the session, to identify at least one output token of the plurality of output tokens logically stemming from at least one input token of the plurality of input tokens originating from the at least one external resource; and
triggering the security action for preventing the at least one external resource from feeding additional adversarial input into the AI-application.
2 . The system of claim 1 , wherein the security action is selected from: blocking the external resource from interacting with the AI-application, and terminating the session with the AI-application.
3 . The system of claim 1 , wherein the control and/or data-flow is traced within the attention-heads of a simulated LLM that is different than the AI-application, the simulated LLM is predicted to generate activation patterns within its attention-heads in response to being fed the simulated session, that are correlated to the activation patterns generated within the LLM guiding the AI-application.
4 . The system of claim 3 , wherein receiving the session comprises at least one of: (i) extracting the session from packets sent over a network by sniffing and/or intercepting, and (ii) monitoring a user interface of the entity for extracting the session, and further comprising feeding the extracted session into the simulated LLM, wherein the tracing is performed during processing of the session by the simulated LLM.
5 . The system of claim 1 , wherein the at least one external resource represents an untrusted resource and the security action is triggered in response to detecting that the untrusted resource triggered the at least one output token.
6 . The system of claim 1 , further comprising code for:
identifying dataflow heads from the attention-heads,
wherein the dataflow heads exhibit correlation patterns between the plurality of input tokens and the plurality of output tokens that indicate causal relationships,
wherein the control and/or data-flow is traced within the dataflow heads.
7 . The system of claim 6 , wherein identifying the dataflow heads from the attention-heads comprises:
generating an annotated dataset of a plurality of sample sessions of at least one sample entity interacting with the AI-application, comprising labels for each segment of each sample session selected from: an internal resource, an external resource and manipulation attempt, and an output of the manipulation attempt;
executing the plurality of sample sessions on the AI-application;
during the execution, computing a score for each attention-head of the LLM based on an amount of a correlation between an activation pattern in the attention head and corresponding causal relationships defined by the annotated dataset;
selecting a subset of attention-heads with scores above a threshold; and
defining the dataflow heads as the subset of attention-heads.
8 . The system of claim 7 , wherein each sample session of the plurality of sample sessions comprises and/or is defined for at least one of:
(i) the AI-application has access to internal resources;
(ii) during the sample session, while the AI-application processes an external resource, the AI-application is manipulated into processing at least one of the internal resources;
(iii) during the sample session, while the AI-application processes an external resource, the AI-application is manipulated into generating a misaligned output.
9 . The system of claim 7 , wherein the plurality of sample sessions include at least one of:
(i) the at least one output token is embedded,
(ii) the at least one output token does not appear in a plain-text format,
(iii) the manipulation attempt includes a plurality of chained tool calls that include the manipulation instructions as a combination.
10 . The system of claim 7 , further comprising code for:
generating a heatmap of the attention-heads, the heatmap comprising a matrix of elements, wherein each element is defined by a certain row and a certain column, wherein each row of the matrix denotes a layer representing a single transformer of the LLM, and each column of the matrix denotes a specific attention head, wherein a pixel intensity and/or color of each respective element is according to the score computed for the attention-head corresponding to the respective element.
11 . The system of claim 6 , further comprising code for:
computing a token-graph comprising a plurality of nodes connected by a plurality of edges,
wherein the token-graph is computed based on the control and/or data-flow within the dataflow heads traced during the session,
wherein each node denotes an input token or an output token of the session,
wherein an edge from a first node to a second node is defined if and only if an attention pattern of the dataflow heads denoting a control and/or data-flow indicates that a first token denoted by the first node is related to a second token denoted by the second node.
12 . The system of claim 11 , further comprising code for:
identifying the at least one output token logically stemming from the at least one input token originating from the external resource, by tracing a path within the token graph, from a third node corresponding to the at least one output token to a fourth node corresponding to the at least one input token originating from the external resource.
13 . The system of claim 1 , wherein the tracing control and/or data-flow is performed to identify that the at least one output token is mapped to at least one first input token originating from the external resource and to at least one second input token originating from an internal resource.
14 . The system of claim 1 , wherein the session comprises a graphical user interface (GUI) presented on a display of a client terminal, configured for a user to enter input into the AI-application, the input including human-readable text, and for presenting output generated by the AI-application guided by the LLM, the output including human-readable text.
15 . The system of claim 1 , wherein the AI-application and/or the LLM is selected from a plurality of open-source models with at least one of: different purposes, different architectures, for different domains, trained using different datasets.
16 . The system of claim 1 , further comprising code for:
prior to loading of at least one file of the LLM and/or AI-application, starting a loading process configured to load the at least one file of the LLM and/or AI-application;
triggering a monitoring process that generates the monitored logs by monitoring system calls and library calls made by the loading process;
selecting monitoring logs recorded by the monitoring process during monitoring between starting of loading by the loading process and completion of the loading by the loading process; and
after the LLM and/or AI-application is loaded, calling the inference on the loaded LLM and/or AI-application using the preselected input.
17 . The system of claim 16 , further comprising code for:
detecting in the subset of logs that do not exist in the clean logs, a certain system call and/or certain library call triggered by the at least one external resource embedded in the LLM,
wherein the security action is triggered in response to detecting the certain system call and/or certain library call triggered by the at least one external resource embedded in the LLM.
18 . The system of claim 16 , further comprising code for:
performing a pre-processing procedure for generating the clean logs, comprising:
identifying a clean LLM comprising a clean instance of the LLM excluding inference-related adversarial payloads,
prior to loading of at least one file of the clean LLM, starting the loading process configured to load the at least one file of the clean instance of the LLM,
triggering the monitoring process that monitors system calls and library calls made by the loading process,
after the clean LLM is loaded, calling inference on the loaded clean LLM, and
generating the clean logs by the monitoring process during the inference session.
19 . The system of claim 16 , wherein the monitoring process is implemented as bpftrace implementing eBPF (extended Berkeley Packet Filter) tracing.
20 . The system of claim 16 , wherein the security action is selected from: quarantining the LLM, deleting the LLM, blocking the LLM, blocking certain system call and/or certain library call triggered by the at least one external resource embedded in the LLM.
21 . The system of claim 1 , further comprising code for:
in response triggering the security action, at least one of: generating a security alert indicating that the external resource is suspicious, logging a security event, and generating a pop-up window with a user interface for presentation on a display for requesting user confirmation to proceed.
22 . The system of claim 1 , wherein the generating monitored logs, the filtering the monitored logs, and the triggering the security action in response to detecting the subset of logs, are implemented prior to the session of the interaction by the entity with the LLM-guided AI-application.
23 . A method for securing an AI-application guided by a large language model (LLM), comprising:
using at least one processor for:
generating monitored logs of system calls and library calls captured from the LLM and/or AI-application during inference using a preselected input,
wherein the LLM processes input tokens, serves as a decision-making center that decides what actions to take next, and generates output tokens, and wherein the AI-application interfaces with an environment and executes the LLM's decisions;
filtering the monitored logs with respect to clean logs generated by monitoring a clean version of the LLM and/or AI-application, to identify a subset of logs that do not exist in the clean logs;
triggering a security action in response to detecting the subset of logs;
receiving a session of an interaction by an entity with the LLM-guided AI-application, the session including a plurality of output tokens generated by the AI-application in response to a plurality of input tokens inputted into the AI-application by the entity and/or from at least one external resource;
dynamically during processing of the session by the LLM, tracing control and/or data-flow within attention-heads of transformer layers of the LLM with respect to the session, to identify at least one output token of the plurality of output tokens logically stemming from at least one input token of the plurality of input tokens originating from the at least one external resource; and
triggering the security action for preventing the at least one external resource from feeding additional adversarial input into the AI-application.
24 . A non-transitory medium storing program instructions for securing an AI-application guided by a large language model (LLM), comprising program instructions which when executed by at least one processor, cause the at least one processor to:
generate monitored logs of system calls and library calls captured from the LLM and/or AI-application during inference using a preselected input,
wherein the LLM processes input tokens, serves as a decision-making center that decides what actions to take next, and generates output tokens, and wherein the AI-application interfaces with an environment and executes the LLM's decisions;
filter the monitored logs with respect to clean logs generated by monitoring a clean version of the LLM and/or AI-application, to identify a subset of logs that do not exist in the clean logs;
trigger a security action in response to detecting the subset of logs;
receive a session of an interaction by an entity with the LLM-guided AI-application, the session including a plurality of output tokens generated by the AI-application in response to a plurality of input tokens inputted into the AI-application by the entity and/or from at least one external resource;
dynamically during processing of the session by the LLM, trace control and/or data-flow within attention-heads of transformer layers of the LLM with respect to the session, to identify at least one output token of the plurality of output tokens logically stemming from at least one input token of the plurality of input tokens originating from the at least one external resource; and
trigger the security action for preventing the at least one external resource from feeding additional adversarial input into the AI-application.