IP Library › Granted Patent US 12,248,883
Granted Patent B1
US 12,248,883 · App. 18/605,337 · Granted Mar 11, 2025

Generative artificial intelligence model prompt injection classifier

Inventors: Jacob Rideout (Raleigh, NC); Tanner Burns (Austin, TX); Kwesi Cappel (Austin, TX); Kenneth Yeung (Ottawa, CA)
Assignee: HiddenLayer, Inc.
G06N3/094G06F21/55G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,883
App. No.
18/605,337
Filed
Mar 14, 2024
Granted
Mar 11, 2025
Kind
B1
Examiner
KIM, SEHWAN
Art Unit
2129
USPC
706/15
Abstract

An analysis engine receives data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI) model. The analysis engine, using a prompt injection classifier determines whether the prompt comprises or is indicative of malicious content or otherwise elicits malicious actions. Data characterizing the determination is provided to a consuming application or process. Related apparatus, systems, techniques and articles are also described.

Claims (50)

1. A computer-implemented method comprising:

receiving, by an analysis engine from a model environment executing a generative artificial intelligence (GenAI) model, data characterizing a prompt for ingestion by the GenAI model, the analysis engine executing in a monitoring environment remote from the model environment;

determining, by the analysis engine using an ensemble of machine learning-based prompt injection classifiers, whether the prompt comprises malicious content or elicits malicious actions, a first of the prompt injection classifiers being trained to identify a first type of prompt injection attack and a second of the prompt injection classifiers being trained to identify a second, different type of prompt injection attack; and

providing data characterizing the determination to a consuming application or process to (i) initiate a remediation action to ensure that the GenAI model does not operate in an undesired manner when it is determined that the prompt comprises malicious content or elicits malicious actions and (ii) allow the prompt to be ingested when it is determined that the prompt does not comprise malicious content or elicit malicious actions;

wherein the initiated remediation action prevents the prompt from being ingested by the GenAI model to ensure that the GenAI model does not operate in an undesired manner and is based on whether the ensemble of machine learning-based prompt injection classifiers identifies the first type of prompt injection attack or the second type of prompt injection attack.

2. The method of claim 1 further comprising:

vectorizing the data characterizing the prompt to result in one or more vectors; and

generating one or more embeddings based on the one or more vectors, the embeddings having a lower dimensionality than the one or more vectors;

wherein at least one of the prompt injection classifiers uses the generated one or more embeddings when making the determination.

3. The method of claim 1 , wherein the GenAI model comprises a large language model.

4. The method of claim 1 , wherein the consuming application or process allows the prompt to be input into the GenAI model upon a determination that the prompt does not comprise or elicit malicious content.

5. The method of claim 1 , wherein the consuming application or process flags the prompt as being malicious for quality assurance upon a determination that the prompt comprises or elicits malicious content.

6. The method of claim 1 , wherein the consuming application or process modifies the prompt to be benign upon a determination that the prompt comprises or elicits malicious content and causes the modified prompt to be ingested by the GenAI model.

7. The method of claim 1 , wherein the consuming application or process blocks an internet protocol (IP) address of a requester of the prompt upon a determination that the prompt comprises or elicits malicious content.

8. The method of claim 1 , wherein the consuming application or process causes a subsequent prompt from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of a requester of the prompt to be modified upon a determination that the prompt comprises or elicits malicious content to result in a modified prompt and causes the modified prompt to be ingested by the GenAI model.

9. A system comprising:

at least one data processor; and

memory storing instructions which, when executed by the at least one data processor, result in operations comprising:

receiving, by an analysis engine from a model environment executing a generative artificial intelligence (GenAI) model, data characterizing a prompt for ingestion by the GenAI model, the analysis engine executing in a monitoring environment remote from the model environment;

determining, by the analysis engine using an ensemble of machine learning-based prompt injection classifiers, whether the prompt comprises malicious content or elicits malicious actions, a first of the prompt injection classifiers being trained to identify a first type of prompt injection attack and a second of the prompt injection classifiers being trained to identify a second, different type of prompt injection attack; and

providing data characterizing the determination to a consuming application or process to (i) initiate a remediation action to ensure that the GenAI model does not operate in an undesired manner when it is determined that the prompt comprises malicious content or elicits malicious actions and (ii) allow the prompt to be ingested when it is determined that the prompt does not comprise malicious content or elicit malicious actions;

wherein the initiated remediation action prevents the prompt from being ingested by the GenAI model to ensure that the GenAI model does not operate in an undesired manner and is based on whether the ensemble of machine learning-based prompt injection classifiers identifies the first type of prompt injection attack or the second type of prompt injection attack.

10. The system of claim 9 , wherein the operations further comprise:

vectorizing the data characterizing the prompt to result in one or more vectors; and

generating one or more embeddings based on the one or more vectors, the embeddings having a lower dimensionality than the one or more vectors;

wherein at least one of the prompt injection classifiers uses the generated one or more embeddings when making the determination.

11. The system of claim 9 , wherein the GenAI model comprises a large language model.

12. The system of claim 9 , wherein the consuming application or process allows the prompt to be input into the GenAI model upon a determination that the prompt does not comprise or elicit malicious content.

13. The system of claim 9 , wherein the consuming application or process flags the prompt as being malicious for quality assurance upon a determination that the prompt comprises or elicits malicious content.

14. The system of claim 9 , wherein the consuming application or process modifies the prompt to be benign upon a determination that the prompt comprises or elicits malicious content and causes the modified prompt to be ingested by the GenAI model.

15. The system of claim 9 , wherein the consuming application or process blocks an internet protocol (IP) address of a requester of the prompt upon a determination that the prompt comprises or elicits malicious content.

16. The system of claim 9 , wherein the consuming application or process causes a subsequent prompt from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of a requester of the prompt to be modified upon a determination that the prompt comprises or elicits malicious content to result in a modified prompt and causes the modified prompt to be ingested by the GenAI model.

17. A system comprising:

a model environment comprising memory and a plurality of data processors executing a generative artificial intelligence (GenAI) model; and

a monitoring environment comprising memory and a plurality of data processors executing an analysis engine, the monitoring environment being remote from the model environment;

wherein:

the analysis engine receives, from the model environment, data characterizing a prompt for ingestion by the GenAI model;

the analysis engine determines, using an ensemble of machine learning-based prompt injection classifiers, whether the prompt comprises malicious content or elicits malicious actions, a first of the prompt injection classifiers being trained to identify a first type of prompt injection attack and a second of the prompt injection classifiers being trained to identify a second, different type of prompt injection attack; and

data characterizing the determination is provided to a consuming application or process to (i) initiate a remediation action to ensure that the GenAI model does not operate in an undesired manner when it is determined that the prompt comprises malicious content or elicits malicious actions and (ii) allow the prompt to be ingested when it is determined that the prompt does not comprise malicious content or elicit malicious actions;

wherein the initiated remediation action prevents the prompt from being ingested by the GenAI model to ensure that the GenAI model does not operate in an undesired manner and is based on whether the ensemble of machine learning-based prompt injection classifiers identifies the first type of prompt injection attack or the second type of prompt injection attack.

18. The system of claim 17 , wherein the analysis engine:

vectorizes the data characterizing the prompt to result in one or more vectors; and

generates one or more embeddings based on the one or more vectors, the embeddings having a lower dimensionality than the one or more vectors;

wherein at least one of the prompt injection classifiers uses the generated one or more embeddings when making the determination.

19. The system of claim 17 , wherein the GenAI model comprises a large language model.

20. The system of claim 17 , wherein the consuming application or process allows the prompt to be input into the GenAI model upon a determination that the prompt does not comprise or elicit malicious content.

21. The system of claim 17 , wherein the consuming application or process flags the prompt as being malicious for quality assurance upon a determination that the prompt comprises or elicits malicious content.

22. The system of claim 17 , wherein the consuming application or process modifies the prompt to be benign upon a determination that the prompt comprises or elicits malicious content and causes the modified prompt to be ingested by the GenAI model.

23. The system of claim 17 , wherein the consuming application or process blocks an internet protocol (IP) address of a requester of the prompt upon a determination that the prompt comprises or elicits malicious content.

24. The system of claim 17 , wherein the consuming application or process causes a subsequent prompt from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of a requester of the prompt to be modified upon a determination that the prompt comprises or elicits malicious content to result in a modified prompt and causes the modified prompt to be ingested by the GenAI model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2024
From: RIDEOUT, JACOB; BURNS, TANNER; CAPPEL, KWESI; YEUNG, KENNETH
To: HIDDENLAYER, INC.
Reel/Frame 066784/0317 →
References Cited (107)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh et al. · 2024 [cited by applicant]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 20100082811A1 · Van Der Merwe · 2010 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140157415A1 · Abercrombie · 2014 [cited by examiner]
US 20150074392A1 · Boivie et al. · 2015 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170154021A1 · Vidhani et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa et al. · 2017 [cited by applicant]
US 20170331841A1 · Ilu et al. · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180063190A1 · Wright et al. · 2018 [cited by applicant]
US 20180205734A1 · Wing · 2018 [cited by examiner]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20180324193A1 · Ronen · 2018 [cited by examiner]
US 20190050564A1 · Pogorelik et al. · 2019 [cited by applicant]
US 20190238572A1 · Manadhata et al. · 2019 [cited by applicant]
US 20190260784A1 · Stockdale et al. · 2019 [cited by applicant]
US 20190311118A1 · Grafi · 2019 [cited by examiner]
US 20190392176A1 · Taron et al. · 2019 [cited by applicant]
US 20200019721A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200076771A1 · Maier et al. · 2020 [cited by applicant]
US 20200175094A1 · Palmer et al. · 2020 [cited by applicant]
US 20200219009A1 · Dao et al. · 2020 [cited by applicant]
US 20200233979A1 · Maraghoosh et al. · 2020 [cited by applicant]
US 20200285737A1 · Kraus et al. · 2020 [cited by applicant]
US 20200409323A1 · Spalt et al. · 2020 [cited by applicant]
US 20210110062A1 · Oliner et al. · 2021 [cited by applicant]
US 20210141897A1 · Seifert et al. · 2021 [cited by applicant]
US 20210209464A1 · Bala et al. · 2021 [cited by applicant]
US 20210224425A1 · Nasr-Azadani et al. · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik et al. · 2021 [cited by applicant]
US 20210319784A1 · Le Roux et al. · 2021 [cited by applicant]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210374247A1 · Sultana et al. · 2021 [cited by applicant]
US 20210407051A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220030009A1 · Hasan · 2022 [cited by applicant]
US 20220058444A1 · Olabiyi et al. · 2022 [cited by applicant]
US 20220070195A1 · Sern · 2022 [cited by examiner]
US 20220083658A1 · Shah et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147597A1 · Bhide et al. · 2022 [cited by applicant]
US 20220164444A1 · Prudkovskij · 2022 [cited by examiner]
US 20220166795A1 · Simioni et al. · 2022 [cited by applicant]
US 20220182410A1 · Tupsamudre et al. · 2022 [cited by applicant]
US 20220253464A1 · Sloane et al. · 2022 [cited by applicant]
US 20220269796A1 · Chase et al. · 2022 [cited by applicant]
US 20220309179A1 · Payne et al. · 2022 [cited by applicant]
US 20230008037A1 · Venugopal et al. · 2023 [cited by applicant]
US 20230027149A1 · Kuan · 2023 [cited by examiner]
US 20230049479A1 · Mozo Velasco et al. · 2023 [cited by applicant]
US 20230109426A1 · Hashimoto et al. · 2023 [cited by applicant]
US 20230148116A1 · Stokes et al. · 2023 [cited by applicant]
US 20230169397A1 · Smith et al. · 2023 [cited by applicant]
US 20230185912A1 · Sinn et al. · 2023 [cited by applicant]
US 20230185915A1 · Rao et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230229960A1 · Zhu et al. · 2023 [cited by applicant]
US 20230252178A1 · Ruelke et al. · 2023 [cited by applicant]
US 20230259787A1 · David et al. · 2023 [cited by applicant]
US 20230269263A1 · Yarabolu · 2023 [cited by applicant]
US 20230274003A1 · Liu et al. · 2023 [cited by applicant]
US 20230289604A1 · Chan et al. · 2023 [cited by applicant]
US 20230351143A1 · Kutt · 2023 [cited by examiner]
US 20230359903A1 · Cefalu · 2023 [cited by examiner]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20230388324A1 · Thompson · 2023 [cited by applicant]
US 20240022585A1 · Burns et al. · 2024 [cited by applicant]
US 20240039948A1 · Koc et al. · 2024 [cited by applicant]
US 20240045959A1 · Marson et al. · 2024 [cited by applicant]
US 20240078337A1 · Kamyshenko et al. · 2024 [cited by applicant]
US 20240080333A1 · Burns et al. · 2024 [cited by applicant]
US 20240126611A1 · Phanishayee et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240289628A1 · Parmar et al. · 2024 [cited by applicant]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240386103A1 · Clement et al. · 2024 [cited by applicant]
Arora, Rakhi, Rishi Gupta, and Pradeep Yadav. “Utilizing Ensemble Learning to enhance the detection of Malicious URLs in the Twitter dataset.” 2024 4th International Conference on Innovative Practices in Technology and … [cited by examiner]
Gupta, Maanak, et al. “From chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy.” IEEE Access (2023). (Year: 2023). [cited by examiner]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10. [cited by applicant]
Cited By (15)
US 12,328,331 US 12,475,215 US 12,505,648 US 12,549,598 US 12,554,855 US 12,566,846 US 12,572,777 US 12,596,839 US 12,608,861 US 12,632,545 US 12,657,297 US 12,717,909 US 12,724,883 US 12,724,894 US 12,737,462