IP Library › Granted Patent US 12,657,297
Granted Patent B2
US 12,657,297 · App. 18/888,093 · Granted Jun 16, 2026

GenAI prompt injection classifier training using prompt attack structures

Inventors: Kenneth Yeung (Ottawa, CA); Tanner Burns (Austin, TX); Kwesi Cappel (Austin, TX)
Assignee: HiddenLayer, Inc.
G06F21/56G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,297
App. No.
18/888,093
Filed
Sep 17, 2024
Granted
Jun 16, 2026
Kind
B2
Art Unit
2433
USPC
726/23
Abstract

An analysis engine receives data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI) model. The analysis engine, using a prompt injection classifier determines whether the prompt comprises or is indicative of malicious content or otherwise elicits malicious actions. The prompt injection classifier can be trained using a dataset generated by populating benign content and malicious content into a plurality of different prompt attack structures at pre-defined locations. Data characterizing the determination is provided to a consuming application or process. Related apparatus, systems, techniques and articles are also described.

Claims (41)

1 . A computer-implemented method comprising:

receiving data characterizing a prompt for ingestion by a generative artificial intelligence (GenAI) model;

determining, using a prompt injection classifier, whether the prompt comprises malicious content or elicits malicious actions; and

providing data characterizing the determination to a consuming application or process;

wherein the prompt injection classifier is trained using a dataset generated by at least one large language model populating benign content and malicious content into a plurality of different prompt attack structures.

2 . The method of claim 1 , wherein at least a portion of the malicious content is generated by instructing a misaligned large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the misaligned large language model having been fine-tuned to output malicious strings.

3 . The method of claim 1 , wherein at least a portion of the malicious content is generated by instructing a jailbroken large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the jailbroken large language model having been prompted with a specific input that allows it to respond with malicious strings.

4 . The method of claim 1 , wherein at least a portion of the malicious content is generated by instructing a large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the large language model having not been aligned in a way that restricts its output.

5 . The method of claim 1 , wherein at least a portion of the benign content is generated using a large language model with guardrails to generate a plurality of strings known to be benign.

6 . The method of claim 1 , wherein at least a portion of the benign content comprises human-generated content.

7 . The method of claim 1 , wherein at least a portion of the malicious content comprises human-generated content.

8 . The method of claim 1 , wherein the GenAI model comprises a large language model.

9 . The method of claim 1 , wherein the consuming application or process prevents the prompt from being input into the GenAI model upon a determination that the prompt comprises malicious content or elicits malicious actions content.

10 . The method of claim 1 , wherein the consuming application or process allows the prompt to be input into the GenAI model upon a determination that the prompt does not comprise malicious content or elicit malicious actions content.

11 . The method of claim 1 , wherein the consuming application or process flags the prompt as being malicious for quality assurance upon a determination that the prompt comprises malicious content or elicits malicious actions content.

12 . The method of claim 1 , wherein the consuming application or process modifies the prompt to be benign upon a determination that the prompt comprises malicious content or elicits malicious actions content and causes the modified prompt to be ingested by the GenAI model.

13 . The method of claim 1 , wherein the consuming application or process blocks an internet protocol (IP) address of a requester of the prompt upon a determination that the prompt comprises malicious content or elicits malicious actions content.

14 . The method of claim 1 , wherein the consuming application or process causes subsequent prompts from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of a requester of the prompt to be modified upon a determination that the prompt comprises malicious content or elicits malicious actions content and causes the modified prompt to be ingested by the GenAI model.

15 . The method of claim 1 , wherein the data characterizing the prompt is vectorized to generate sentence embeddings, and the prompt injection classifier determines whether the prompt comprises malicious content or elicits malicious actions based at least in part on features derived from the embeddings.

16 . The method of claim 1 , wherein the prompt injection classifier outputs a confidence score indicative of whether the prompt comprises malicious content or elicits malicious actions, and providing data characterizing the determination comprises providing the confidence score to the consuming application or process.

17 . The method of claim 1 , wherein determining uses an ensemble comprising a binary prompt injection classifier executed prior to a multi-class prompt injection classifier that, responsive to the binary classifier indicating the prompt is malicious, identifies a type of prompt injection attack for the prompt.

18 . The method of claim 17 , wherein the type of prompt injection attack is selected from a direct task deflection attack, a special case attack, a context continuation attack, a context termination attack, a syntactic transformation attack, an encryption attack, or a text redirection attack.

19 . The method of claim 1 , wherein the plurality of different prompt attack structures comprises templates selected from defined dictionary templates, context switching templates, and jailbreak templates.

20 . The method of claim 1 , wherein the prompt injection classifier is executed by an analysis engine in a monitoring environment communicatively coupled to a proxy interposed between client devices and a model environment hosting the GenAI model, the proxy relaying prompts or features characterizing prompts for analysis prior to ingestion by the GenAI model.

21 . The method of claim 20 , wherein the proxy intercepts prompts from the client devices and relays the prompts or features characterizing the prompts to the analysis engine in the monitoring environment prior to the prompts being input into the GenAI model.

22 . The method of claim 20 , wherein the proxy further relays outputs of the GenAI model to the monitoring environment and the determining is based on at least one of the prompt and the output.

23 . The method of claim 1 , wherein providing data characterizing the determination to the consuming application or process causes generation of an alert and capture of system or process behavior associated with the prompt.

24 . The method of claim 1 , wherein the prompt injection classifier is executed by a local analysis engine within a model environment and the determination is transmitted to a monitoring environment or external remediation resources for remediation.

25 . A computer-implemented method comprising:

training a prompt injection classifier using a dataset generated by at least one large language model to populate benign content and malicious content into a plurality of different prompt attack structures at pre-defined locations; and

causing the trained prompt injection classifier to be deployed to determine whether prompts to be ingested by a generative artificial intelligence (GenAI) model comprise malicious content or elicits malicious actions.

26 . The method of claim 25 , wherein at least a portion of the malicious content is generated by instructing a misaligned large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the misaligned large language model having been fine-tuned to output malicious strings.

27 . The method of claim 25 , wherein at least a portion of the malicious content is generated by instructing a jailbroken large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the jailbroken large language model having been prompted with a specific input that allows it to respond with malicious strings.

28 . The method of claim 25 , wherein at least a portion of the malicious content is generated by instructing a large language model to generate a plurality of strings having malicious content or eliciting malicious actions, the large language model having not been aligned in a way that restricts its output.

29 . A computer-implemented method comprising:

generating malicious content by instructing a misaligned or jailbroken first large language model to generate malicious strings having malicious content or eliciting malicious actions;

generating benign content by instructing a second large language model to generate benign strings having benign content; and

generating at least a portion of a training dataset by using a large language model to populate a plurality of different prompt attack structures with generated benign strings at locations tagged in the structures as being benign and with generated malicious strings at locations tagged in the structures as being malicious.

30 . The method of claim 29 further comprising:

training a prompt injection classifier using the training dataset; and

deploying the trained prompt injection classifier to determine whether prompts to be ingested by a generative artificial intelligence (GenAI) model comprise malicious content or elicits malicious actions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2024
From: YEUNG, KENNETH; BURNS, TANNER; CAPPEL, KWESI
To: HIDDENLAYER, INC.
Reel/Frame 068635/0920 →
Continuity (2)
Continuation 18676190 · May 28, 2024
Related Publication 20250371148A1 · Dec 4, 2025
References Cited (159)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9350748B1 · McClintock et al. · 2016 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10121104B1 · Hu et al. · 2018 [cited by applicant]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10205735B2 · Apostolopoulos · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 10824721B2 · Kesarwani et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710045B2 · Lee et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930030B1 · Burns et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11954199B1 · Burns et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh et al. · 2024 [cited by applicant]
US 11995180B1 · Cappel · 2024 [cited by examiner]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12026255B1 · Burns et al. · 2024 [cited by applicant]
US 12052206B1 · Lai · 2024 [cited by examiner]
US 12105844B1 · Burns et al. · 2024 [cited by applicant]
US 12107885B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12111926B1 · Beveridge et al. · 2024 [cited by applicant]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 12130943B1 · Burns et al. · 2024 [cited by applicant]
US 12137118B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12174954B1 · Yeung et al. · 2024 [cited by applicant]
US 12182264B2 · Sinha et al. · 2024 [cited by applicant]
US 12197859B1 · Malviya et al. · 2025 [cited by applicant]
US 12204323B1 · Malviya et al. · 2025 [cited by applicant]
US 12229265B1 · Yeung et al. · 2025 [cited by applicant]
US 12248883B1 · Rideout et al. · 2025 [cited by applicant]
US 12293277B1 · Yeung et al. · 2025 [cited by applicant]
US 20100082811A1 · Van Der Merwe · 2010 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140157415A1 · Abercrombie et al. · 2014 [cited by applicant]
US 20150074392A1 · Boivie et al. · 2015 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170154021A1 · Vidhani et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa et al. · 2017 [cited by applicant]
US 20170331841A1 · Hu et al. · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180063190A1 · Wright et al. · 2018 [cited by applicant]
US 20180205734A1 · Wing et al. · 2018 [cited by applicant]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20180324193A1 · Ronen et al. · 2018 [cited by applicant]
US 20190050564A1 · Pogorelik et al. · 2019 [cited by applicant]
US 20190238568A1 · Goswami et al. · 2019 [cited by applicant]
US 20190238572A1 · Manadhata et al. · 2019 [cited by applicant]
US 20190260784A1 · Stockdale et al. · 2019 [cited by applicant]
US 20190311118A1 · Grafi et al. · 2019 [cited by applicant]
US 20190392176A1 · Taron et al. · 2019 [cited by applicant]
US 20200019721A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200076771A1 · Maier et al. · 2020 [cited by applicant]
US 20200092299A1 · Srinivasan et al. · 2020 [cited by applicant]
US 20200175094A1 · Palmer et al. · 2020 [cited by applicant]
US 20200219009A1 · Dao et al. · 2020 [cited by applicant]
US 20200233979A1 · Maraghoosh et al. · 2020 [cited by applicant]
US 20200279192A1 · Godfrey et al. · 2020 [cited by applicant]
US 20200285737A1 · Kraus et al. · 2020 [cited by applicant]
US 20200313849A1 · Kar et al. · 2020 [cited by applicant]
US 20200364333A1 · Derks et al. · 2020 [cited by applicant]
US 20200403826A1 · Dawani et al. · 2020 [cited by applicant]
US 20200409323A1 · Spalt et al. · 2020 [cited by applicant]
US 20210110062A1 · Oliner et al. · 2021 [cited by applicant]
US 20210141897A1 · Seifert et al. · 2021 [cited by applicant]
US 20210209464A1 · Bala et al. · 2021 [cited by applicant]
US 20210218673A1 · Ma et al. · 2021 [cited by applicant]
US 20210224425A1 · Nasr-Azadani et al. · 2021 [cited by applicant]
US 20210303695A1 · Grosse et al. · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik et al. · 2021 [cited by applicant]
US 20210319784A1 · Le Roux et al. · 2021 [cited by applicant]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210374247A1 · Sultana et al. · 2021 [cited by applicant]
US 20210407051A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220030009A1 · Hasan · 2022 [cited by applicant]
US 20220058444A1 · Olabiyi et al. · 2022 [cited by applicant]
US 20220070195A1 · Sern et al. · 2022 [cited by applicant]
US 20220083658A1 · Shah et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147597A1 · Bhide et al. · 2022 [cited by applicant]
US 20220164444A1 · Prudkovskij et al. · 2022 [cited by applicant]
US 20220166795A1 · Simioni et al. · 2022 [cited by applicant]
US 20220182410A1 · Tupsamudre et al. · 2022 [cited by applicant]
US 20220253464A1 · Sloane et al. · 2022 [cited by applicant]
US 20220269796A1 · Chase et al. · 2022 [cited by applicant]
US 20220284283A1 · Yin et al. · 2022 [cited by applicant]
US 20220309179A1 · Payne et al. · 2022 [cited by applicant]
US 20230008037A1 · Venugopal et al. · 2023 [cited by applicant]
US 20230027149A1 · Kuan et al. · 2023 [cited by applicant]
US 20230049479A1 · Mozo Velasco et al. · 2023 [cited by applicant]
US 20230109426A1 · Hashimoto et al. · 2023 [cited by applicant]
US 20230111744A1 · Chandrasekaran et al. · 2023 [cited by applicant]
US 20230128947A1 · Bhaskar et al. · 2023 [cited by applicant]
US 20230148116A1 · Stokes et al. · 2023 [cited by applicant]
US 20230169397A1 · Smith et al. · 2023 [cited by applicant]
US 20230185912A1 · Sinn et al. · 2023 [cited by applicant]
US 20230185915A1 · Rao et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230229960A1 · Zhu et al. · 2023 [cited by applicant]
US 20230252178A1 · Ruelke et al. · 2023 [cited by applicant]
US 20230259787A1 · David et al. · 2023 [cited by applicant]
US 20230269263A1 · Yarabolu · 2023 [cited by applicant]
US 20230274003A1 · Liu et al. · 2023 [cited by applicant]
US 20230289604A1 · Chan et al. · 2023 [cited by applicant]
US 20230351143A1 · Kutt et al. · 2023 [cited by applicant]
US 20230359903A1 · Cefalu et al. · 2023 [cited by applicant]
US 20230359924A1 · Maman et al. · 2023 [cited by applicant]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20230388324A1 · Thompson · 2023 [cited by applicant]
US 20240005690A1 · Brodie et al. · 2024 [cited by applicant]
US 20240007469A1 · Wang et al. · 2024 [cited by applicant]
US 20240022585A1 · Burns et al. · 2024 [cited by applicant]
US 20240031026A1 · Fujisawa et al. · 2024 [cited by applicant]
US 20240039948A1 · Koc et al. · 2024 [cited by applicant]
US 20240045959A1 · Marson et al. · 2024 [cited by applicant]
US 20240054233A1 · Ohayon et al. · 2024 [cited by applicant]
US 20240078337A1 · Kamyshenko et al. · 2024 [cited by applicant]
US 20240080333A1 · Burns et al. · 2024 [cited by applicant]
US 20240126611A1 · Phanishayee et al. · 2024 [cited by applicant]
US 20240127065A1 · Ren et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240289628A1 · Parmar et al. · 2024 [cited by applicant]
US 20240289863A1 · Smith Lewis et al. · 2024 [cited by applicant]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240386103A1 · Clement et al. · 2024 [cited by applicant]
US 20240414177A1 · Lal et al. · 2024 [cited by applicant]
US 20240427986A1 · Shakarian et al. · 2024 [cited by applicant]
US 20250086455A1 · Yang et al. · 2025 [cited by applicant]
Wang et al., 2023, “Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models,” arXiv:2308.11521v1 [cs.CL] Aug. 16, 2023 (15 pages). [cited by applicant]
Shayegani et al., 2023, “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks,” arXiv:2310.10844v1 [cs.CL] Oct. 16, 2023 (54 pages). [cited by applicant]
Bezymiannyi et al., 2023, “Filter for confidential information,” Electronics and Control Systems 4(78):21-25 (5 pages). [cited by applicant]
Wang et al., 2023, “Self-Guard: Empower the LLM to Safeguard Itself,” ACL Anthology, NAACL (21 pages). [cited by applicant]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10. [cited by applicant]
International Search Report and Written Opinion mailed Aug. 7, 2025 for PCT/US2025/030944 filed May 27, 2025 (14 pages). [cited by applicant]
Automorphic.ai, 2024, “Github—automorphic-ai/aegis: Self-hardening firewall for large language models,” XP093278213, Available online at https://web.archive.org/web/20240222171700/https://github.com/automorphic-ai/aegis… [cited by applicant]
Sun et al., “Conscendi: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants,” arXiv:2304.14364v1 [cs.CL] Apr. 27, 2023 (20 pages). [cited by applicant]
Hu et al., “Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes,” arXiv:2403.00867v1 [cs.CR] Mar. 1, 2024 (19 pages). [cited by applicant]
Robey et al., “SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks,” arXiv:2310.03684v3 [cs.LG] Nov. 29, 2023 (42 pages). [cited by applicant]
Mohtashami et al., “Social Learning: Towards Collaborative Learning with Large Language Models,” arXiv:2312.11441v2 [cs.LG] Feb. 8, 2024 (19 pages). [cited by applicant]
Dinan et al., 2021, “Anticipating safety issues in e2e conversational ai: Framework and tooling,” arXiv preprint arXiv:2107.03451v3 (43 pages). [cited by applicant]