IP Library › Granted Patent US 12,293,277
Granted Patent B1
US 12,293,277 · App. 18/792,455 · Granted May 6, 2025

Multimodal generative AI model protection using sequential sidecars

Inventors: Kenneth Yeung (Ottawa, CA); Jason Martin (Beaverton, OR)
Assignee: HiddenLayer, Inc.
G06N3/0475G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,277
App. No.
18/792,455
Granted
May 6, 2025
Kind
B1
Abstract

Data is received which includes multimodal input for ingestion by a first generative AI (GenAI) model is received. This received data is input into the first GenAI model to result in a first output. The first output along with the received data is input into a second GenAI model to result in a second output. The first GenAI model is a modified (e.g., fine-tuned, etc.) version of the second GenAI model. When the second output indicates that guardrails associated with the second GenAI model have been triggered, one or more remediation actions are initiated. Otherwise, the first output is returned to the requestor. Related apparatus, systems, techniques and articles are also described.

Claims (61)

1. A computer-implemented method comprising:

receiving, over a computer network from a requestor, data comprising multimodal input for ingestion by a first generative artificial intelligence (GenAI) model, the first GenAI model comprising one or more machine learning models, the multimodal input comprising input having two or more modalities;

inputting the received data into the first GenAI model to result in a first output;

inputting both of the received data and the first output into a second GenAI model to result in a second output, the second GenAI model comprising one or more machine learning models;

determining, based on content of the second output, whether the second output indicates that guardrails associated with the second GenAI model have been triggered;

returning, over the computer network, the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have not been triggered; and

initiating one or more remediation actions in lieu of returning the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have been triggered, the one or more remediation actions preventing the first GenAI model from behaving in an undesired manner.

2. The method of claim 1 , wherein the first GenAI model is a modified version of the second GenAI model.

3. The method of claim 2 , wherein the first GenAI model is a fine-tuned version of the second GenAI model.

4. The method of claim 1 , wherein the first GenAI model is of a different type than the second GenAI model.

5. The method of claim 1 , wherein the second GenAI model is a different, aligned model relative to the first GenAI model.

6. The method of claim 1 , wherein the data is received from a proxy intercepting inputs to the first GenAI model, the proxy being executed in a model environment of the first GenAI model.

7. The method of claim 1 , wherein the initiated one or more remediation actions comprise:

returning the second output to the requestor.

8. The method of claim 1 , wherein the initiated one or more remediation actions comprise:

flagging the input as being malicious for quality assurance.

9. The method of claim 1 , wherein the initiated one or more remediation actions comprise:

modifying the input to be benign;

causing the modified input to be ingested by the first GenAI model to result in a third output; and

returning the third output to the requestor.

10. The method of claim 1 , wherein the initiated one or more remediation actions comprise:

blocking one or more of an internet protocol (IP) address of the requestor, a media access control (MAC) address associated with the requestor, or a session identifier of the requestor.

11. The method of claim 1 , wherein the initiated one or more remediation actions comprise:

causing subsequent inputs from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of the requester to be modified prior to input into the first GenAI model.

12. The method of claim 1 , wherein both of the first GenAI model and the second GenAI model comprise a large language model.

13. A system comprising:

at least one data processor; and

memory storing instructions which, when executed by the at least one data processor, result in operations comprising:

receiving, over a computer network from a requestor, data comprising multimodal input for ingestion by a first generative artificial intelligence (GenAI) model, the first GenAI model comprising one or more machine learning models, the multimodal input comprising input having two or more modalities;

inputting the received data into the first GenAI model to result in a first output;

inputting both of the received data and the first output into a second GenAI model to result in a second output, the second GenAI model comprising one or more machine learning models;

determining, based on content of the second output, whether the second output indicates that guardrails associated with the second GenAI model have been triggered;

returning, over the computer network, the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have not been triggered; and

initiating one or more remediation actions in lieu of returning the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have been triggered, the one or more remediation actions preventing the first GenAI model from behaving in an undesired manner.

14. The system of claim 13 , wherein the first GenAI model is a modified version of the second GenAI model.

15. The system of claim 14 , wherein the first GenAI model is a fine-tuned version of the second GenAI model.

16. The system of claim 13 , wherein the first GenAI model is of a different type than the second GenAI model.

17. The system of claim 13 wherein the second GenAI model is a different, aligned model relative to the first GenAI model.

18. The system of claim 13 , wherein the data is received from a proxy intercepting inputs to the first GenAI model, the proxy being executed in a model environment of the first GenAI model.

19. The system of claim 13 , wherein the initiated one or more remediation actions comprise:

returning the second output to the requestor.

20. The system of claim 13 , wherein the initiated one or more remediation actions comprise:

flagging the input as being malicious for quality assurance.

21. The system of claim 13 , wherein the initiated one or more remediation actions comprise:

modifying the input to be benign;

causing the modified input to be ingested by the first GenAI model to result in a third output; and

returning the third output to the requestor.

22. The system of claim 13 , wherein the initiated one or more remediation actions comprise:

blocking one or more of an internet protocol (IP) address of the requestor, a media access control (MAC) address associated with the requestor, or a session identifier of the requestor.

23. The system of claim 13 , wherein the initiated one or more remediation actions comprise:

causing subsequent inputs from an entity identified by one or more of an internet protocol (IP) address, a media access control (MAC) address, or a session identifier of the requester to be modified prior to input into the first GenAI model.

24. The system of claim 13 , wherein both of the first GenAI model and the second GenAI model each comprise a large language model.

25. A computer-implemented method comprising:

receiving, over a computer network from a requestor, data comprising multimodal input for ingestion by a first generative artificial intelligence (GenAI) model, the first GenAI model comprising one or more machine learning models, the multimodal input comprising input having two or more modalities;

inputting the received data into the first GenAI model to result in a first output;

intercepting, by a proxy in a model computing environment executing the first GenAI model, the received data and the first output;

transmitting, over the computer network, the intercepted received data and the first output to a monitoring computing environment executing a second GenAI model;

inputting both of the received data and the first output into the second GenAI model to result in a second output, the second GenAI model comprising one or more machine learning models;

determining, based on content of the second output, whether the second output indicates that guardrails associated with the second GenAI model have been triggered;

returning, over the computer network, the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have not been triggered; and

initiating one or more remediation actions in lieu of returning the first output to the requestor when it is determined that the second output indicates that guardrails associated with the second GenAI model have been triggered, the one or more remediation actions preventing the first GenAI model from behaving in an undesired manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2024
From: YEUNG, KENNETH; MARTIN, JASON
To: HIDDENLAYER, INC.
Reel/Frame 068201/0286 →
References Cited (124)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh · 2024 [cited by examiner]
US 11978437B1 · Thattai · 2024 [cited by examiner]
US 11995180B1 · Cappel et al. · 2024 [cited by applicant]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12099781B1 · Malladi · 2024 [cited by examiner]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 12197859B1 · Malviya · 2025 [cited by examiner]
US 12204323B1 · Malviya · 2025 [cited by examiner]
US 12229265B1 · Yeung · 2025 [cited by examiner]
US 20100082811A1 · Van Der Merwe · 2010 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140157415A1 · Abercrombie et al. · 2014 [cited by applicant]
US 20150074392A1 · Boivie et al. · 2015 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170154021A1 · Vidhani et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa et al. · 2017 [cited by applicant]
US 20170331841A1 · Hu et al. · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180063190A1 · Wright et al. · 2018 [cited by applicant]
US 20180205734A1 · Wing et al. · 2018 [cited by applicant]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20180324193A1 · Ronen et al. · 2018 [cited by applicant]
US 20190050564A1 · Pogorelik et al. · 2019 [cited by applicant]
US 20190147357A1 · Erlandson · 2019 [cited by examiner]
US 20190238572A1 · Manadhata et al. · 2019 [cited by applicant]
US 20190260784A1 · Stockdale et al. · 2019 [cited by applicant]
US 20190311118A1 · Grafi et al. · 2019 [cited by applicant]
US 20190392176A1 · Taron et al. · 2019 [cited by applicant]
US 20200019721A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200076771A1 · Maier et al. · 2020 [cited by applicant]
US 20200175094A1 · Palmer et al. · 2020 [cited by applicant]
US 20200219009A1 · Dao et al. · 2020 [cited by applicant]
US 20200233979A1 · Maraghoosh et al. · 2020 [cited by applicant]
US 20200285737A1 · Kraus et al. · 2020 [cited by applicant]
US 20200409323A1 · Spalt et al. · 2020 [cited by applicant]
US 20210110062A1 · Oliner et al. · 2021 [cited by applicant]
US 20210141897A1 · Seifert et al. · 2021 [cited by applicant]
US 20210209464A1 · Bala et al. · 2021 [cited by applicant]
US 20210224425A1 · Nasr-Azadani et al. · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik et al. · 2021 [cited by applicant]
US 20210319784A1 · Le Roux et al. · 2021 [cited by applicant]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210374247A1 · Sultana et al. · 2021 [cited by applicant]
US 20210407051A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220030009A1 · Hasan · 2022 [cited by applicant]
US 20220058444A1 · Olabiyi et al. · 2022 [cited by applicant]
US 20220070195A1 · Sern et al. · 2022 [cited by applicant]
US 20220083658A1 · Shah et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147597A1 · Bhide et al. · 2022 [cited by applicant]
US 20220164444A1 · Prudkovskij et al. · 2022 [cited by applicant]
US 20220166795A1 · Simioni et al. · 2022 [cited by applicant]
US 20220182410A1 · Tupsamudre et al. · 2022 [cited by applicant]
US 20220253464A1 · Sloane et al. · 2022 [cited by applicant]
US 20220269796A1 · Chase et al. · 2022 [cited by applicant]
US 20220309179A1 · Payne et al. · 2022 [cited by applicant]
US 20230008037A1 · Venugopal et al. · 2023 [cited by applicant]
US 20230027149A1 · Kuan et al. · 2023 [cited by applicant]
US 20230049479A1 · Mozo Velasco et al. · 2023 [cited by applicant]
US 20230109426A1 · Hashimoto et al. · 2023 [cited by applicant]
US 20230148116A1 · Stokes et al. · 2023 [cited by applicant]
US 20230169397A1 · Smith et al. · 2023 [cited by applicant]
US 20230185912A1 · Sinn et al. · 2023 [cited by applicant]
US 20230185915A1 · Rao et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230229960A1 · Zhu et al. · 2023 [cited by applicant]
US 20230252178A1 · Ruelke et al. · 2023 [cited by applicant]
US 20230259787A1 · David et al. · 2023 [cited by applicant]
US 20230269263A1 · Yarabolu · 2023 [cited by applicant]
US 20230274003A1 · Liu et al. · 2023 [cited by applicant]
US 20230289604A1 · Chan et al. · 2023 [cited by applicant]
US 20230315856A1 · Lee · 2023 [cited by examiner]
US 20230351143A1 · Kutt et al. · 2023 [cited by applicant]
US 20230359903A1 · Cefalu · 2023 [cited by examiner]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20230388324A1 · Thompson · 2023 [cited by applicant]
US 20240022585A1 · Burns et al. · 2024 [cited by applicant]
US 20240039948A1 · Koc et al. · 2024 [cited by applicant]
US 20240045959A1 · Marson et al. · 2024 [cited by applicant]
US 20240054233A1 · Ohayon · 2024 [cited by examiner]
US 20240078337A1 · Kamyshenko et al. · 2024 [cited by applicant]
US 20240080333A1 · Burns et al. · 2024 [cited by applicant]
US 20240126611A1 · Phanishayee et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240289628A1 · Parmar et al. · 2024 [cited by applicant]
US 20240289863A1 · Smith Lewis · 2024 [cited by examiner]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240354379A1 · Sun · 2024 [cited by examiner]
US 20240386103A1 · Clement et al. · 2024 [cited by applicant]
US 20240414177A1 · Lal · 2024 [cited by examiner]
WO WO2024145209A1 · 2024 [cited by examiner]
WO WO2024166446A1 · 2024 [cited by examiner]
Dong, Yi, et al. “Safeguarding Large Language Models: A Survey.” arXiv preprint arXiv:2406.02622 (Jun. 2024). (Year: 2024). [cited by examiner]
Lu, Qinghua, et al. “Towards responsible generative ai: A reference architecture for designing foundation model based agents.” 2024 IEEE 21st International Conference on Software Architecture Companion (ICSA-C). IEEE, J… [cited by examiner]
Lu, Qinghua, et al. “A taxonomy of foundation model based systems through the lens of software architecture.” Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI. Jun. … [cited by examiner]
Deng, Zehang, et al. “AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways.” arXiv preprint arXiv:2406.02630 v1 (Jun. 2024). (Year: 2024). [cited by examiner]
Dinan, Emily, et al. “Anticipating safety issues in e2e conversational ai: Framework and tooling.” arXiv preprint arXiv:2107.03451v3 (2021). (Year: 2021). [cited by examiner]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10. [cited by applicant]
Cited By (12)
US 12,475,215 US 12,505,648 US 12,549,598 US 12,554,855 US 12,572,777 US 12,596,839 US 12,608,861 US 12,632,545 US 12,657,297 US 12,717,909 US 12,724,883 US 12,724,894