Leakage detection for large language models
A method includes receiving, at a server from a user device, a user query to a large language model (LLM), creating an LLM query from the user query and an application context, gathering confidential information from the LLM query, and sending the LLM query to the LLM. The method includes receiving, from the LLM, an LLM response to the LLM query, comparing the LLM response to the confidential information to generate comparison result, and setting a leakage detection signal based on comparison result.
1 . A method comprising:
receiving, at a server from a user device, a user query to a large language model (LLM);
creating an LLM query from the user query and an application context;
gathering confidential information from the LLM query;
sending the LLM query to the LLM;
receiving, from the LLM, an LLM response to the LLM query;
detecting an overlap in the LLM response and the confidential information; and
blocking the LLM response responsive to detecting the overlap.
2 . The method of claim 1 , further comprising:
partitioning the confidential information into a plurality of segments; and
executing a string matching algorithm on the LLM response and the plurality of segments.
3 . The method of claim 2 , wherein the plurality of segments is overlapping.
4 . The method of claim 2 , further comprising:
storing confidential information with a query identifier in storage; and
responsive to receiving the LLM response, obtaining the confidential information matching the query identifier in the LLM response.
5 . The method of claim 2 , further comprising:
performing an Aho-Corasick algorithm on the LLM response and the confidential information.
6 . The method of claim 1 , wherein the confidential information comprises at least one selected from a group consisting of application context, a set of previous queries, a set of previous responses, and the LLM query.