IP Library › Granted Patent US 12,737,462
Granted Patent B2
US 12,737,462 · App. 18/656,650 · Granted Sep 15, 2026

Detection of indirect prompt injection attacks with malicious instructions detection models

Inventors: Chien-Hua Lu (San Jose, CA); Bo Qu (Saratoga, CA); Xu Zou (Saratoga, CA); Sergey Sviridov (Santa Clara, CA)
Assignee: Palo Alto Networks, Inc.
G06F21/56G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,462
App. No.
18/656,650
Filed
May 7, 2024
Granted
Sep 15, 2026
Kind
B2
Art Unit
2431
USPC
726/23
Abstract

A malicious instructions detection model (“detector”) intercepts augmented prompts destined for a large language model (“LLM”). Each augmented prompt was augmented with data from potentially compromised data sources susceptible to indirect prompt injection attacks. The detector tokenizes/preprocesses sentences in the augmented prompts and is invoked on the tokenized/preprocessed sentences to obtain confidence scores that each sentence comprises malicious instructions. If one or more of the confidence scores is above a threshold, the detector blocks the augmented prompt and generates an alert indicating the blocking and the malicious instructions. Otherwise, the detector communicates the augmented prompt to its intended LLM.

Claims (54)

1 . A method comprising:

based on receiving a user query, retrieving knowledge related to the user query from a knowledge base, wherein the knowledge base accesses and stores data for retrieval-augmented generation, without security precautions, from one or more data sources that are potentially compromised with poisoned data;

augmenting one or more prompt templates for prompts intended for a language model to respond to user queries with the retrieved knowledge to generate augmented prompts; and

for each augmented prompt of the augmented prompts,

generating feature vectors for each sentence in the augmented prompt, wherein

generating the feature vectors comprises, for each sentence in the augmented prompt,

tokenizing the sentence; and

generating a natural language processing embedding of the tokenized sentence, wherein the natural language processing embedding comprises one of the feature vectors corresponding to the sentence;

invoking a machine learning model on the feature vectors to obtain confidence scores indicating confidence that corresponding sentences in the augmented prompt comprise malicious task instructions for the language model; and

based on determining from the confidence scores that at least a subset of the sentences of the augmented prompt comprises malicious task instructions,

identifying one or more of the sentences in the augmented prompt corresponding to highest of the confidence scores as a source of malicious task instructions from indirect prompt injection in the augmented prompt via the one or more potentially compromised data sources; and

filtering the augmented prompt.

2 . The method of claim 1 , wherein the malicious task instructions comprise task instructions to the language model to ignore a conversational history for the language model.

3 . The method of claim 1 , wherein the machine learning model comprises at least one of a Bidirectional Encoder Representations from Transformers model and a one-dimensional convolutional neural network.

4 . The method of claim 1 , wherein generating the natural language processing embedding comprises applying sentence2vec to the tokenized sentence.

5 . The method of claim 1 , wherein retrieving the knowledge from the knowledge base comprises querying a vector database corresponding to the knowledge base with a semantic embedding of the user query.

6 . The method of claim 1 , wherein each prompt template of the one or more prompt templates comprises task instructions to the language model to restrict at least one of actions taken by the language model in response to the user query, types of responses to return in response to the user query, and data sources accessed in response to the user query.

7 . The method of claim 1 , further comprising determining from the confidence scores that at least a subset of the sentences of the augmented prompt comprises malicious task instructions, wherein determining that at least the subset of the sentences of the augmented prompt comprises malicious task instructions comprises determining that one or more of the confidence scores exceeds a threshold confidence score.

8 . The method of claim 1 , further comprising sanitizing training prompts for a chatbot based on detecting sentences comprising malicious task instructions in the training prompts with the machine learning model.

9 . A non-transitory machine-readable medium having program code stored thereon, the program code comprising instructions to:

based on receiving user queries, retrieve knowledge related to the user queries from a knowledge base, wherein the knowledge base accesses and stores data for retrieval-augmented generation, without security precautions, from one or more data sources that are potentially compromised with poisoned data;

augment one or more prompt templates for prompts intended for a language model to respond to the user queries with the retrieved knowledge to generate augmented prompts; and

for each augmented prompt of the augmented prompts,

generate feature vectors for each sentence in the augmented prompt, wherein the instructions to generate feature vectors for each sentence in the augmented prompt comprise instructions to, for each sentence in the augmented prompt,

tokenize the sentence; and

generate a natural language processing embedding of the tokenized sentence, wherein the natural language processing embedding comprises one of the feature vectors corresponding to the sentence;

invoke a machine learning model on the feature vectors to obtain confidence scores of whether each of the sentences in the augmented prompt comprises malicious task instructions; and

based on a determination from the confidence scores that at least a subset of the sentences of the augmented prompt comprises malicious instructions,

identify one or more of the sentences in the augmented prompt corresponding to highest of the confidence scores as a source of malicious task instructions from indirect prompt injection in the augmented prompt via the one or more potentially compromised data sources; and

filter the augmented prompt.

10 . The non-transitory machine-readable medium of claim 9 , wherein the instructions to, for each augmented prompt, determine whether the augmented prompt is malicious based on classifications by the machine learning model on the feature vectors comprise instructions to determine that the confidence scores satisfy a criterion for maliciousness.

11 . The non-transitory machine-readable medium of claim 10 , wherein the criterion for maliciousness comprises that one or more of the confidence scores exceed a threshold confidence score.

12 . The non-transitory machine-readable medium of claim 11 , wherein the program code further comprises instructions to generate an alert indicating the one or more of the sentences in the augmented prompt corresponding to the highest of the confidence scores as the source of malicious task instructions.

13 . The non-transitory machine-readable medium of claim 9 , wherein the program code further comprises instructions to, for each augmented prompt, based on a determination by the machine learning model that the augmented prompt is benign, communicate the augmented prompt to the language model.

14 . The non-transitory machine-readable medium of claim 9 , wherein the instructions to generate the natural language processing embedding comprise instructions to apply sentence2vec to the tokenized sentence.

15 . An apparatus comprising:

a processor; and

a machine-readable medium having instructions stored thereon that are executable by the processor to cause the apparatus to,

based on receiving user queries, retrieve knowledge related to the user queries from a knowledge base, wherein the knowledge base accesses and stores data for retrieval-augmented generation, without security precautions, from one or more data sources that are potentially compromised with poisoned data;

augment one or more prompt templates for prompts intended for a language model to respond to the user queries with the retrieved knowledge to generate augmented prompts; and

for each augmented prompt of the augmented prompts,

generate feature vectors of sentences in the augmented prompt, wherein the instructions to generate feature vectors of sentences in the augmented prompt comprise instructions executable by the processor to cause the apparatus to, for each of the sentences in the augmented prompt,

tokenize the sentence; and

generate a natural language processing embedding of the tokenized sentence, wherein the natural language processing embedding comprises one of the feature vectors corresponding to the sentence;

invoke a machine learning model on the feature vectors to determine confidence scores of whether corresponding ones of the sentences of the augmented prompt comprise malicious task instructions; and

based on a determination from the confidence scores that at least a subset of the sentences of the augmented prompt comprises malicious instructions,

identify one or more of the sentences in the augmented prompt corresponding to highest of the confidence scores as a source of malicious task instructions from indirect prompt injection in the augmented prompt via the one or more potentially compromised data sources; and

filter the augmented prompt from the augmented prompts.

16 . The apparatus of claim 15 , wherein the instructions to, for each augmented prompt in the augmented prompts, invoke the machine learning model on the feature vectors to determine whether the augmented prompt comprises malicious task instructions comprise instructions executable by the processor to cause the apparatus to:

determine that the confidence scores satisfy a criterion for maliciousness.

17 . The apparatus of claim 16 , wherein the criterion for maliciousness comprises that one or more of the confidence scores exceed a threshold confidence score.

18 . The apparatus of claim 17 , wherein the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to generate an alert indicating the identified one or more of the sentences in the augmented prompt corresponding to highest of the confidence scores as the source of the malicious task instructions.

19 . The apparatus of claim 15 , the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to, for each augmented prompt of the augmented prompts, based on a determination that the augmented prompt does not comprise malicious instructions, communicate the augmented prompt to the language model.

20 . The apparatus of claim 15 , wherein the instructions to generate the natural language processing embedding comprise instructions executable by the processor to cause the apparatus to apply sentence2vec to the tokenized sentence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2024
From: LU, CHIEN-HUA; QU, BO; ZOU, XU; SVIRIDOV, SERGEY
To: PALO ALTO NETWORKS, INC.
Reel/Frame 067336/0125 →
Continuity (1)
Related Publication 20250348583A1 · Nov 13, 2025
References Cited (13)
US 12147513B1 · Jain · 2024 [cited by examiner]
US 12248883B1 · Rideout · 2025 [cited by examiner]
US 20210042662A1 · Pu · 2021 [cited by examiner]
US 20250103715A1 · Jackson · 2025 [cited by examiner]
US 20250110711A1 · Palanki · 2025 [cited by examiner]
US 20250117414A1 · McCurdy · 2025 [cited by examiner]
US 20250156527A1 · Palanki · 2025 [cited by examiner]
US 20250190801A1 · Lucas · 2025 [cited by examiner]
US 20250209208A1 · Vaknin · 2025 [cited by examiner]
US 20250245315A1 · Palanki · 2025 [cited by examiner]
US 20250252320A1 · Mayande · 2025 [cited by examiner]
US 20250342822A1 · Shepherd · 2025 [cited by examiner]
Greshake, et al., “Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”, https://arxiv.org/pdf/2302.12173, May 5, 2023, 33 pages. [cited by applicant]