IP Library Granted Patent US 12,632,545
Granted Patent B2
US 12,632,545 · App. 18/935,336 · Granted May 19, 2026

Generative AI model information leakage prevention

Inventors: Kenneth Yeung (Ottawa, CA); Tanner Burns (Austin, TX); Kwesi Cappel (Austin, TX)
Assignee: HiddenLayer, Inc.
G06F21/554G06F21/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,545
App. No.
18/935,336
Filed
Nov 1, 2024
Granted
May 19, 2026
Kind
B2
Art Unit
2435
USPC
726/22
Abstract

An output of a GenAI model responsive to a prompt is received. The GenAI model is configured using one or more system prompts including one or more Easter eggs. The output is scanned to confirm whether an Easter egg is present. In cases in which at least one Easter egg is present, one or more remediation actions can be initiated to thwart an information leak by the GenAI model. Related apparatus, systems, techniques and articles are also described.

Claims (54)

1 . A computer-implemented method for thwarting an information leak by a generative artificial intelligence (GenAI) model comprising:

intercepting, by a proxy executing in a computing environment of the GenAI model, an output of the GenAI model responsive to a prompt, the GenAI model being configured using one or more system prompts including an Easter egg, the system prompts causing the Easter egg to be part of select outputs of the GenAI model;

scanning the output to confirm that the Easter egg is present; and

initiating, in response to the confirmation that the Easter egg is present, one or more remediation actions to thwarting the information leak by the GenAI model.

2 . The method of claim 1 , wherein the GenAI model comprises a large language model.

3 . The method of claim 1 , wherein the one or more remediation actions sanitize the output to remove undesired content prior to returning the output to a requesting user.

4 . The method of claim 1 , wherein the one or more remediation actions modify the output to remove the Easter egg prior to returning the output to a requesting user.

5 . The method of claim 1 , wherein the Easter egg comprises a pseudorandom sequence of characters.

6 . The method of claim 1 , wherein the Easter egg comprises Unicode.

7 . The method of claim 1 , wherein the Easter egg comprises pictographic languages.

8 . The method of claim 1 , wherein the Easter egg comprises a word that is unlikely to form part of the output.

9 . The method of claim 1 further comprising:

transmitting, by the proxy, the output to a monitoring environment;

wherein the scanning is performed by an analysis engine executing in the monitoring environment.

10 . The method of claim 9 , wherein the one or more remediation actions are initiated by a remediation engine forming part of the monitoring environment.

11 . The method of claim 1 further comprising:

transmitting, by the proxy, to an analysis engine, the analysis engine residing with a model environment executing the GenAI model;

wherein the scanning is performed by the analysis engine.

12 . The method of claim 11 , wherein the one or more remediation actions are initiated by a remediation engine forming part of the model environment.

13 . A computer-implemented method for preventing an information leak of a generative artificial intelligence (GenAI) model comprising:

intercepting, by a proxy executing in a computing environment of the GenAI model, an output of the GenAI model responsive to a prompt;

scanning the output to confirm that at least one tag used to configure the GenAI model is present, the GenAI model being configured using system prompts which cause the at least one tag to be part of select outputs of the GenAI model; and

providing data characterizing the scanning to a consuming application or process to initiate one or more actions preventing the information leak.

14 . The method of claim 13 , wherein the at least one tag comprises an Easter egg.

15 . The method of claim 14 , wherein the one or more actions modify the output to remove the Easter egg prior to returning the output to a requesting user.

16 . A system comprising:

a model environment comprising a plurality of computing devices executing a generative artificial intelligence (GenAI) model and a proxy, the GenAI model being configured using one or more system prompts including an Easter egg, the system prompts causing the Easter egg to be part of select outputs of the GenAI model; and

a monitoring environment separate and distinct from the model environment comprising a plurality of computing devices executing an analysis engine and a remediation engine;

the monitoring environment:

receiving an output of the GenAI model responsive to a prompt;

scanning, by the analysis engine, the output to confirm that the Easter egg is present; and

initiating, by the remediation engine in response to the confirmation that the Easter egg is present, one or more remediation actions thwarting an information leak by the GenAI model.

17 . The system of claim 16 , wherein the GenAI model comprises a large language model.

18 . The system of claim 16 , wherein the one or more remediation actions sanitize the output to remove undesired content prior to returning the output to a requesting user.

19 . The system of claim 16 , wherein the one or more remediation actions modify the output to remove the Easter egg prior to returning the output to a requesting user.

20 . A system for thwarting an information leak of a generative artificial intelligence (GenAI) model comprising:

at least one data processor; and

memory storing instructions which, when executed by the at least one data processor, result in operations comprising:

intercepting, by a proxy executing in a computing environment of the GenAI model, an output of the GenAI model responsive to a prompt, the GenAI model being configured using one or more system prompts including at least one Easter egg, the system prompts causing the Easter egg to be part of select outputs of the GenAI model;

scanning the output to confirm that one or more Easter eggs are present; and

initiating, in response to the confirmation that one or more Easter eggs are present, one or more remediation actions thwarting the information leak by the GenAI model.

21 . The system of claim 20 , wherein the GenAI model comprises a large language model.

22 . The system of claim 20 , wherein the one or more remediation actions sanitize the output to remove undesired content prior to returning the output to a requesting user.

23 . The system of claim 20 , wherein the one or more remediation actions modify the output to remove the Easter egg prior to returning the output to a requesting user.

24 . The system of claim 20 , wherein the Easter egg comprises a pseudorandom sequence of characters, Unicode, or pictographic languages.

25 . The system of claim 20 , wherein the Easter egg comprises a word that is unlikely to form part of the output.

26 . The system of claim 20 , wherein the operations further comprise: intercepting, by the proxy, the output of the GenAI model prior to the scanning.

27 . The system of claim 26 , wherein the operations further comprise: transmitting, by the proxy, to an analysis engine, the analysis engine residing with a model environment executing the GenAI model;

wherein the scanning is performed by the analysis engine.

28 . The system of claim 27 , wherein the one or more remediation actions are initiated by a remediation engine forming part of the model environment.

29 . The system of claim 20 , wherein the operations further comprise:

transmitting, by the proxy, the output to a monitoring environment;

wherein the scanning is performed by an analysis engine executing in the monitoring environment.

30 . The system of claim 29 , wherein the one or more remediation actions are initiated by a remediation engine forming part of the monitoring environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2024
From: YEUNG, KENNETH; BURNS, TANNER; CAPPEL, KWESI
To: HIDDENLAYER, INC.
Reel/Frame 069132/0005 →
Continuity (2)
Continuation 18673033 · May 23, 2024
Related Publication 20250363210A1 · Nov 27, 2025
References Cited (164)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9350748B1 · McClintock et al. · 2016 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10121104B1 · Hu et al. · 2018 [cited by applicant]
US 10169315B1 · Heckel · 2019 [cited by examiner]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10205735B2 · Apostolopoulos · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 10824721B2 · Kesarwani et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710045B2 · Lee et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930030B1 · Burns et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11954199B1 · Burns et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh et al. · 2024 [cited by applicant]
US 11995180B1 · Cappel et al. · 2024 [cited by applicant]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12026255B1 · Burns et al. · 2024 [cited by applicant]
US 12105844B1 · Burns et al. · 2024 [cited by applicant]
US 12107885B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12111926B1 · Beveridge et al. · 2024 [cited by applicant]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 12130917B1 · Yeung et al. · 2024 [cited by applicant]
US 12130943B1 · Burns et al. · 2024 [cited by applicant]
US 12137118B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12182264B2 · Sinha et al. · 2024 [cited by applicant]
US 12197859B1 · Malviya et al. · 2025 [cited by applicant]
US 12204323B1 · Malviya et al. · 2025 [cited by applicant]
US 12229265B1 · Yeung et al. · 2025 [cited by applicant]
US 12248883B1 · Rideout et al. · 2025 [cited by applicant]
US 12293277B1 · Yeung et al. · 2025 [cited by applicant]
US 20100082811A1 · Van Der Merwe · 2010 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140157415A1 · Abercrombie et al. · 2014 [cited by applicant]
US 20150074392A1 · Boivie et al. · 2015 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170154021A1 · Vidhani et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa et al. · 2017 [cited by applicant]
US 20170331841A1 · Hu et al. · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180063190A1 · Wright et al. · 2018 [cited by applicant]
US 20180205734A1 · Wing et al. · 2018 [cited by applicant]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20180324193A1 · Ronen et al. · 2018 [cited by applicant]
US 20190050564A1 · Pogorelik et al. · 2019 [cited by applicant]
US 20190104124A1 · Buford · 2019 [cited by examiner]
US 20190238568A1 · Goswami et al. · 2019 [cited by applicant]
US 20190238572A1 · Manadhata et al. · 2019 [cited by applicant]
US 20190260784A1 · Stockdale et al. · 2019 [cited by applicant]
US 20190311118A1 · Grafi et al. · 2019 [cited by applicant]
US 20190392176A1 · Taron et al. · 2019 [cited by applicant]
US 20200019721A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200076771A1 · Maier et al. · 2020 [cited by applicant]
US 20200092299A1 · Srinivasan et al. · 2020 [cited by applicant]
US 20200175094A1 · Palmer et al. · 2020 [cited by applicant]
US 20200219009A1 · Dao et al. · 2020 [cited by applicant]
US 20200233977A1 · Chickerur · 2020 [cited by examiner]
US 20200233979A1 · Maraghoosh et al. · 2020 [cited by applicant]
US 20200279192A1 · Godfrey et al. · 2020 [cited by applicant]
US 20200285737A1 · Kraus et al. · 2020 [cited by applicant]
US 20200313849A1 · Kar et al. · 2020 [cited by applicant]
US 20200349298A1 · Nagpal · 2020 [cited by examiner]
US 20200364333A1 · Derks et al. · 2020 [cited by applicant]
US 20200403826A1 · Dawani et al. · 2020 [cited by applicant]
US 20200409323A1 · Spalt et al. · 2020 [cited by applicant]
US 20210110062A1 · Oliner et al. · 2021 [cited by applicant]
US 20210141897A1 · Seifert et al. · 2021 [cited by applicant]
US 20210209464A1 · Bala et al. · 2021 [cited by applicant]
US 20210218673A1 · Ma et al. · 2021 [cited by applicant]
US 20210224425A1 · Nasr-Azadani et al. · 2021 [cited by applicant]
US 20210303695A1 · Grosse et al. · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik et al. · 2021 [cited by applicant]
US 20210319784A1 · Le Roux et al. · 2021 [cited by applicant]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210374247A1 · Sultana et al. · 2021 [cited by applicant]
US 20210383014A1 · Yang · 2021 [cited by examiner]
US 20210407051A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220030009A1 · Hasan · 2022 [cited by applicant]
US 20220058444A1 · Olabiyi et al. · 2022 [cited by applicant]
US 20220070195A1 · Sern et al. · 2022 [cited by applicant]
US 20220083658A1 · Shah et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147597A1 · Bhide et al. · 2022 [cited by applicant]
US 20220164444A1 · Prudkovskij et al. · 2022 [cited by applicant]
US 20220166795A1 · Simioni et al. · 2022 [cited by applicant]
US 20220182410A1 · Tupsamudre et al. · 2022 [cited by applicant]
US 20220253464A1 · Sloane et al. · 2022 [cited by applicant]
US 20220269796A1 · Chase et al. · 2022 [cited by applicant]
US 20220284283A1 · Yin et al. · 2022 [cited by applicant]
US 20220309179A1 · Payne et al. · 2022 [cited by applicant]
US 20230008037A1 · Venugopal et al. · 2023 [cited by applicant]
US 20230027149A1 · Kuan et al. · 2023 [cited by applicant]
US 20230049479A1 · Mozo Velasco et al. · 2023 [cited by applicant]
US 20230109426A1 · Hashimoto et al. · 2023 [cited by applicant]
US 20230111744A1 · Chandrasekaran et al. · 2023 [cited by applicant]
US 20230128947A1 · Bhaskar et al. · 2023 [cited by applicant]
US 20230148116A1 · Stokes et al. · 2023 [cited by applicant]
US 20230169397A1 · Smith et al. · 2023 [cited by applicant]
US 20230185912A1 · Sinn et al. · 2023 [cited by applicant]
US 20230185915A1 · Rao et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230229803A1 · Mozer · 2023 [cited by examiner]
US 20230229960A1 · Zhu et al. · 2023 [cited by applicant]
US 20230252178A1 · Ruelke et al. · 2023 [cited by applicant]
US 20230259787A1 · David et al. · 2023 [cited by applicant]
US 20230269263A1 · Yarabolu · 2023 [cited by applicant]
US 20230274003A1 · Liu et al. · 2023 [cited by applicant]
US 20230289604A1 · Chan et al. · 2023 [cited by applicant]
US 20230351143A1 · Kutt et al. · 2023 [cited by applicant]
US 20230359903A1 · Cefalu et al. · 2023 [cited by applicant]
US 20230359924A1 · Maman et al. · 2023 [cited by applicant]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20230388324A1 · Thompson · 2023 [cited by applicant]
US 20240005690A1 · Brodie et al. · 2024 [cited by applicant]
US 20240007469A1 · Wang et al. · 2024 [cited by applicant]
US 20240022585A1 · Burns et al. · 2024 [cited by applicant]
US 20240031026A1 · Fujisawa et al. · 2024 [cited by applicant]
US 20240039948A1 · Koc et al. · 2024 [cited by applicant]
US 20240045959A1 · Marson et al. · 2024 [cited by applicant]
US 20240054233A1 · Ohayon et al. · 2024 [cited by applicant]
US 20240078337A1 · Kamyshenko et al. · 2024 [cited by applicant]
US 20240080333A1 · Burns et al. · 2024 [cited by applicant]
US 20240126611A1 · Phanishayee et al. · 2024 [cited by applicant]
US 20240127065A1 · Ren et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240289628A1 · Parmar et al. · 2024 [cited by applicant]
US 20240289863A1 · Smith Lewis et al. · 2024 [cited by applicant]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240386103A1 · Clement et al. · 2024 [cited by applicant]
US 20240414177A1 · Lal et al. · 2024 [cited by applicant]
US 20240427986A1 · Shakarian et al. · 2024 [cited by applicant]
US 20250086455A1 · Yang et al. · 2025 [cited by applicant]
Automorphic.ai, 2024, “Github automorphic-ai/aegis: Self-hardening firewall for large language models,” XP093278213, Available online at https://web.archive.org/web/20240222171700/https://github.com/automorphic-ai/aegis… [cited by applicant]
Sun et al., “CONSCENDI: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants,” arXiv:2304.14364v1 [cs.CL] Apr. 27, 2023 (20 pages). [cited by applicant]
International Search Reportand Written Opinion mailed Sep. 8, 2025 for PCT/US2025/030125 filed May 20, 2025 (7 pages). [cited by applicant]
Hu et al., “Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes,” arXiv:2403.00867v1 [cs.CR] Mar. 1, 2024 (19 pages). [cited by applicant]
Robey et al., “SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks,” arXiv:2310.03684v3 [cs.LG] Nov. 29, 2023 (42 pages). [cited by applicant]
Mohtashami et al., “SocialLearning: Towards Collaborative Learning with Large Language Models,” arXiv:2312.11441v2 [cs.LG] Feb. 8, 2024 (19 pages). [cited by applicant]
Wang et al., 2023, “Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models,” arXiv:2308.11521v1 [cs.CL] Aug. 16, 2023 (15 pages). [cited by applicant]
Shayegani et al., 2023, “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks,” arXiv:2310.10844v1 [cs.CL ] Oct. 16, 2023 (54 pages). [cited by applicant]
Bezymiannyi et al., 2023, “Filter for confidential information,” Electronics and Control Systems 4(78):21-25 (5 pages). [cited by applicant]
Wang et al., 2023, “Self-Guard: Empower the LLM to Safeguard Itself,” ACL Anthology, NAACL (21 pages). [cited by applicant]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10. [cited by applicant]
Dinan et al., 2021, “Anticipating safety issues in e2e conversational ai: Framework and tooling,” arXiv preprint arXiv:2107.03451v3 (43 pages). [cited by applicant]