IP Library › Granted Patent US 12,505,648
Granted Patent B1
US 12,505,648 · App. 19/261,899 · Granted Dec 23, 2025

Multimodal AI model protection using embeddings

Inventors: Ravikumar Balakrishnan (Beaverton, OR); Jason Martin (Beaverton, OR); Andrew Davis (Portland, OR)
Assignee: HiddenLayer, Inc.
G06V10/761G06V10/7747G06V10/95
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,648
App. No.
19/261,899
Filed
Jul 7, 2025
Granted
Dec 23, 2025
Kind
B1
Examiner
ABDI, AMARA
Art Unit
2668
USPC
382/159
Abstract

Techniques for assessing multi-modal inputs to a machine learning model involve receiving a multimodal input containing an image, producing several transformed versions of that image, and generating embeddings for both the original and transformed images. A pairwise similarity analysis among all embeddings is conducted to determine distance values. Two dissimilarity metrics can then be calculated: one reflecting the differences among the transformed images, and another comparing the original image to its transformed versions. If the dissimilarity among the transformed images is greater than that between the original and transformed images plus a threshold, the system triggers a remediation action. This action either blocks the input from being processed by the machine learning model or prevents the model's output from being returned to the requester, thereby enhancing the reliability and security of the model.

Claims (59)

1 . A method for implementation by one or more computing devices comprising:

receiving a multi-modal input to a machine learning model, the multi-modal input comprising an image;

generating a plurality of transformed versions of the image;

generating an embedding for the image and an embedding for each of the transformed versions of the image;

performing a pairwise similarity analysis among the generated embeddings to generate a distance value amongst each pair of embeddings;

generating a first dissimilarity metric based on the distance values corresponding to the transformed versions of the image excluding the image;

generating a second dissimilarity metric based on the distance values corresponding to the image relative to the transformed versions of the image; and

initiating at least one remediation action (i) preventing the multi-modal input to be ingested by the machine learning model or (ii) blocking an output of the machine learning model after ingesting the input in response to the first dissimilarity metric being greater than the second dissimilarity metric.

2 . The method of claim 1 , wherein the plurality of transformed versions of the image are generated by adding one or more of: Gaussian noise, impulse noise, shot noise, periodic noise, speckle noise, or quantization noise, to the image.

3 . The method of claim 1 , wherein the multi-modal input is allowed to be ingested by the machine learning model in response to a difference of the first dissimilarity metric and the second dissimilarity metric being equal to or greater than a first threshold.

4 . The method of claim 3 , wherein the at least one remediation action is initiated in response to a difference of the first dissimilarity metric and the second dissimilarity metric being below the first threshold.

5 . The method of claim 1 , wherein the similarity analysis comprises a vector similarity analysis using a vector corresponding to each generated embedding.

6 . The method of claim 1 , wherein the similarity analysis comprises: a cosine similarity analysis, a Euclidean distance, a Jaccard similarity analysis, a Manhattan distance, a Minkowski distance, a Chebyshev distance, a dot product, a Mahalanobis distance, or a Word Movers's distance.

7 . The method of claim 1 , wherein the first dissimilarity metric and the second dissimilarity metric comprises average or mean values for the respective groups of embeddings.

8 . The method of claim 1 , wherein the embeddings are generated using an image embedding model forming part of the machine learning model.

9 . The method of claim 1 , wherein the embeddings are generated using a second machine learning model external to the machine learning model.

10 . The method of claim 1 , wherein there are two or more embeddings generated for each of the transformed versions of the image.

11 . The method of claim 10 , wherein a plurality of different machine learning models external to the machine learning model are used to generated the two or more embeddings of the transformed versions of each image.

12 . The method of claim 1 , wherein the input comprises audio, video, and/or text in addition to the image.

13 . The method of claim 1 , wherein the input comprises two or more images and the analysis is performed for each image individually.

14 . A method for implementation by one or more computing devices comprising:

receiving a multi-modal input to a machine learning model, the multi-modal input comprising a first portion in a first modality and a second portion in a second, different modality;

generating a plurality of transformed versions of the first portion;

generating an embedding for the first portion and an embedding for each of the transformed versions of the first portion;

performing a pairwise similarity analysis among the generated embeddings to generate a distance value amongst each pair of embeddings;

generating a first dissimilarity metric based on the distance values corresponding to the transformed versions of the image excluding the image;

generating a second dissimilarity metric based on the distance values corresponding to the image relative to the transformed versions of the image; and

initiating at least one remediation action (i) preventing the multi-modal input to be ingested by the machine learning model or (ii) blocking an output of the machine learning model after ingesting the input in response to the first dissimilarity metric being greater than the second dissimilarity metric.

15 . The method of claim 14 , wherein the plurality of transformed versions of the first portion are generated by adding one or more of: Gaussian noise, impulse noise, shot noise, periodic noise, speckle noise, or quantization noise, to the image.

16 . The method of claim 14 , wherein the multi-modal input is allowed to be ingested by the machine learning model in response to a difference of the first dissimilarity metric and the second dissimilarity metric being equal to or greater than a first threshold.

17 . The method of claim 16 , wherein the at least one remediation action is initiated in response to a difference of the first dissimilarity metric and the second dissimilarity metric being below the first threshold.

18 . The method of claim 14 , wherein the similarity analysis comprises a vector similarity analysis using a vector corresponding to each generated embedding.

19 . The method of claim 14 , wherein the similarity analysis comprises: a cosine similarity analysis, a Euclidean distance, a Jaccard similarity analysis, a Manhattan distance, a Minkowski distance, a Chebyshev distance, a dot product, a Mahalanobis distance, or a Word Movers's distance.

20 . The method of claim 14 , wherein the first dissimilarity metric and the second dissimilarity metric comprises average or mean values for the respective groups of embeddings.

21 . The method of claim 14 , wherein the embeddings are generated using an image embedding model forming part of the machine learning model.

22 . The method of claim 14 , wherein the embeddings are generated using a second machine learning model external to the machine learning model.

23 . The method of claim 14 , wherein the first modality comprises an image.

24 . The method of claim 14 , wherein the first modality comprises audio.

25 . The method of claim 14 , wherein the first modality comprises video.

26 . A method for implementation by one or more computing devices comprising:

receiving a multi-modal input to a machine learning model, the multi-modal input comprising an image;

generating a plurality of transformed versions of the image;

generating an embedding for the image and an embedding for each of the transformed versions of the image;

performing a pairwise similarity analysis among the generated embeddings to generate a distance value amongst each pair of embeddings;

generating a first dissimilarity metric based on the distance values corresponding to the transformed versions of the image excluding the image;

generating a second dissimilarity metric based on the distance values corresponding to the image relative to the transformed versions of the image; and

initiating at least one remediation action (i) preventing the multi-modal input to be ingested by the machine learning model or (ii) blocking an output of the machine learning model after ingesting the input in response to the first dissimilarity metric being greater than the second dissimilarity metric plus a threshold.

27 . A method for implementation by one or more computing devices comprising:

receiving a multi-modal input to a machine learning model, the multi-modal input comprising two or more image;

for each image:

generating a plurality of transformed versions of the image;

generating an embedding for the image and an embedding for each of the transformed versions of the image;

performing a pairwise similarity analysis among the generated embeddings to generate a distance value amongst each pair of embeddings;

generating a first dissimilarity metric based on the distance values corresponding to the transformed versions of the image excluding the image;

generating a second dissimilarity metric based on the distance values corresponding to the image relative to the transformed versions of the image; and

initiating, based on the first similarity metrics and the second similarity metrics generated for the images, at least one remediation action (i) preventing the multi-modal input to be ingested by the machine learning model or (ii) blocking an output of the machine learning model after ingesting the input.

28 . The method of claim 1 , wherein the multi-modal input is received by a proxy in a model environment hosting the machine learning model, the proxy transmitting data characterizing each image to a monitoring environment comprising an analysis engine that performs the pairwise similarity analysis and a remediation engine that, responsive to the first and second dissimilarity metrics for at least one image, returns an instruction to the proxy to block forwarding of the flagged image prior to ingestion by the machine learning model.

29 . The method of claim 1 , wherein the machine learning model is an adapter-based vision-language multimodal model, embeddings are generated by an image embedding model forming part of the multimodal model on a graphics processing unit within the model environment, the transformed versions are produced by adding pixel-wise uniform random noise to normalized images, the pairwise similarity analysis comprises cosine similarity, and the first and second dissimilarity metrics are computed as means.

30 . The method of claim 1 , wherein the monitoring environment further comprises a proxy of the machine learning model configured to emulate the model for analysis, and initiating the remediation action comprises quarantining only images whose second dissimilarity metric relative to the first dissimilarity metric violates a per-image threshold while allowing remaining images and non-image portions of the multi-modal input to proceed to the machine learning model, and generating an alert with event metadata stored in a monitoring data store.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2025
From: BALAKRISHNAN,, RAVIKUMAR; MARTIN, JASON; DAVIS, ANDREW
To: HIDDENLAYER, INC.
Reel/Frame 071700/0965 →
References Cited (103)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10121104B1 · Hu et al. · 2018 [cited by applicant]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 10824721B2 · Kesarwani et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710045B2 · Lee et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930030B1 · Burns et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11954199B1 · Burns et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh et al. · 2024 [cited by applicant]
US 11995180B1 · Cappel · 2024 [cited by examiner]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12026255B1 · Burns · 2024 [cited by examiner]
US 12105844B1 · Burns et al. · 2024 [cited by applicant]
US 12107885B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12111926B1 · Beveridge et al. · 2024 [cited by applicant]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 12130917B1 · Yeung et al. · 2024 [cited by applicant]
US 12130943B1 · Burns et al. · 2024 [cited by applicant]
US 12137118B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12160409B2 · Kuta · 2024 [cited by examiner]
US 12174954B1 · Yeung et al. · 2024 [cited by applicant]
US 12182264B2 · Sinha et al. · 2024 [cited by applicant]
US 12197859B1 · Malviya et al. · 2025 [cited by applicant]
US 12204323B1 · Malviya et al. · 2025 [cited by applicant]
US 12229265B1 · Yeung et al. · 2025 [cited by applicant]
US 12248883B1 · Rideout et al. · 2025 [cited by applicant]
US 12293277B1 · Yeung et al. · 2025 [cited by applicant]
US 12314380B2 · Burns et al. · 2025 [cited by applicant]
US 12328331B1 · Jia et al. · 2025 [cited by applicant]
US 20210157912A1 · Kruthiveti Subrahmanyeswara Sai · 2021 [cited by examiner]
US 20220343026A1 · Yang · 2022 [cited by examiner]
US 20240320952A1 · Petitpont · 2024 [cited by examiner]
US 20250005201A1 · Qiao · 2025 [cited by examiner]
CN 117786750A · 2024 [cited by applicant]
Abadi et al., 2016, “Deep Learning with Differential Privacy,” arXiv:1607.00133v2 [stat.ML] Oct. 24, 2016 (14 pages). [cited by applicant]
Carlini et al., 2021, “Extracting Training Data from Large Language Models,” arXiv:2012.07805v2 [cs.CR] Jun. 15, 2021 (19 pages). [cited by applicant]
Chao et al., 2023, “Jailbreaking black box large language models in twenty queries,” University of Pennsylvania, Available online at: https://arxiv.org/abs/2211.09527 (21 pages). [cited by applicant]
Choquette-Choo et al., 2021, “Label-Only Membership Inference Attacks,” arXiv:2007.14321v3 [cs.CR] Dec. 5, 2021 (17 pages). [cited by applicant]
Goodfellow et al., 2015, “Explaining and harnessing adversarial examples,” 3rd International Conference on Learning Representations, ICLR 2015, Available online at: http://arxiv.org/abs/1412.6572 (11 pages). [cited by applicant]
Heusel et al., 2018, “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” arXiv:1706.08500v6 [cs.LG] Jan. 12, 2018 (38 pages). [cited by applicant]
Hu et al., 2021, “LoRA: Low-rank adaptation of large language models,” arXiv:2106.09685v2 [cs.CL] Oct. 16, 2021 (26 pages). [cited by applicant]
Hu et al., 2022, “LoRA: Low-rank adaptation of large language models,” International Conference on Learning Representations, Available online at: https://openreview.net/forum?id=nZeVKeeFYf9 (13 pages). [cited by applicant]
Imoxto, 2024, “prompt injection cleaned data set-v2,” Hugging Face, available online at: https://huggingface.co/datasets/imoxto/prompt_injection_cleaned_datasetv2 (3 pages). [cited by applicant]
Jiang et al., 2023, “Mistral 7b,” Available online at: https://mistral.ai/news/announcing-mistral-7b (9 pages). [cited by applicant]
Kahla et al., 2022, “Label-Only Model Inversion Attacks via Boundary Repulsion,” arXiv:2203.01925v1 [cs.LG] Mar. 3, 2022 (13 pages). [cited by applicant]
Ke et al., 2017, “Lightgbm: a highly efficient gradient boosting decision tree,” Proceedings of the 31st International Conference on Neural Information Processing Systems (9 pages). [cited by applicant]
Ko et al., 2023, “PrivMon: A Stream-Based System for Real-Time Privacy Attack Detection for Machine Learning Models,” Raid 2023 https://doi.org/10.1145/3607199.3607232 (18 pages). [cited by applicant]
Lee et al., 2023, “Wizardvicunalm,” Available online at: https://github.com/melodysdreamj/WizardVicunaLM (6 pages). [cited by applicant]
Lee, 2023, “ChatGPT DAN,” ChatGPT DAN, Jailbreaks prompt, Available online at: https://github.com/0xk1h0/ChatGPT_DAN (3 pages). [cited by applicant]
Li et al., 2021, “Membership Leakage in Label-Only Exposures,” arXiv:2007.155 28v3 [cs.LG] Sep. 17, 2021 (17 pages). [cited by applicant]
Li et al., 2022, “Blacklight: Scalable Defense for Neural Networks a ga inst Query-Based Black-Box Attacks,” Proceedings of the 31st Usenix Security Symposium (19 pages). [cited by applicant]
Lian et al., 2023. “Openorca: An open datasetof gpt augmented flan rea soning traces,” Hugging Face, Available online at: https://huggingface.co/Open-Orca/OpenOrca (8 pages). [cited by applicant]
Liu et al., 2015, “Deep Learning Face Attributes in the Wild,” arXiv:1411.7766v3 [cs.CV] Sep. 24, 2015 (11 pages). [cited by applicant]
Luo et al., 2024, “Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks,” Available online at: https://arxiv.org/abs/2404.03027 (20 pages). [cited by applicant]
MacDiarmid et al., 2024, “Simple probes can catch sleeper agents,” Available online at: https://www.anthropic.com/news/probescatch-sleeper-agents (18 pages). [cited by applicant]
Madry et al., 2019, “Towards Deep Learning Models Resistant to Adversarial Attacks,” arXiv:1706.06083v4 [stat.ML] Sep. 4, 2019 (28 pages). [cited by applicant]
Mattern et al., 2023, “Membership Inference Attacks against Language Models via Neighborhood Comparison,” arXiv:2305.18462v2 [cs.CL] Aug. 7, 2023 (12 pages). [cited by applicant]
Mireshghallah et al., 2022, “Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks,” arXiv:2203.03929v2 [cs.LG] Nov. 4, 2022 (16 pages). [cited by applicant]
Perez et al., 2022, “Ignore previous prompt: Attack techniques for language models,” NeurIPS ML Safety Workshop, 36th Conference on Neural Information Processing System (NeurIPS2022), Available online at: https://openre… [cited by applicant]
Raman et al., 2023, “Model-tuning via prompts makes NLP models adversarially robust,” The 2023 Conference on Empirical Methods in Natural Language Processing, Available online at: https://openreview.net/forum?id=R4yb4m7… [cited by applicant]
Salimans et al., 2016, “Improved Techniques for Training GANs,” arXiv:1606.03498v1 [cs.LG] Jun. 10, 2016 (10 pages). [cited by applicant]
Schulhoff et al., 2023, “Ignore this title andhackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompthacking competition,” The 2023 Conference on Empirical Methods in Natural Language Processing, … [cited by applicant]
Sujet-Ai, 2024, “Sujetfinance dataset,” Huging Face, https://huggingface.co/datasets/sujetai/Sujet-Finance-Instruct-177k (6 pages). [cited by applicant]
Templeton et al., 2024, “Scalingmonosemanticity : Extracting interpretable features from claude 3 sonnet,” Transformer Circuits Thread, Available online at: https://transformer-circuits.pub/2024/scalingmonosemanticity/i… [cited by applicant]
Touvron et al., 2023, “Llama 2: Open foundation and fine-tuned chat models,” Available online at: https://arxiv.org/abs/2307.09288 (77 pages). [cited by applicant]
Zhang et al., 2020, “The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks,” https://arxiv.org/abs/1911.07135 (9 pages). [cited by applicant]
Zhang et al., 2024, “Tinyllama : An open-source small language model,” Available online at: https://arxiv.org/abs/2401.02385 (10 pages). [cited by applicant]
Zheng et al., 2023, “Judging LLM-as-a-judge with MT-bench and chatbot arena,” Thirty -seventh Conference on Neural Information Processing Systems Data sets and Benchmarks Track, Available online at: https://openreview.n… [cited by applicant]
Zou et al., 2023, “Representation engineering: A top-down approach to ai transparency,” Available online at: https://arxiv.org/abs/2310.01405 (55 pages). [cited by applicant]
Zou et al., 2023, “Universal and transferable adversarial attacks on aligned language models,” Available online at: https://arxiv.org/abs/2307.15043 (31 pages). [cited by applicant]
Kim et al., 2023, “Robust Safety Classifier for Large Language Models: Adversarial Prompt Shield,” Available online at: https://arxiv.org/abs/2311.00172 (11 pages). [cited by applicant]
Kim et al., 2023, “Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer,” Available online at https://arxiv.org/abs/2305.14152 (21 pages). [cited by applicant]
Zhou et al., 2024, “A Survey on Efficient Inference for Large Language Models,” Available online at https://arxiv.org/abs/2404.14294 (36 pages). [cited by applicant]
International Search Reportand Written Opinion mailed Jun. 25, 2025 for PCT/US2025/026109 filed Apr. 24, 2025 (11 pages). [cited by applicant]
Dinan et al., 2021, “Anticipating safety issues in e2e conversational ai: Framework andtooling,” arXiv preprint arXiv:2107.03451v3 (43 pages). [cited by applicant]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10 (10 pages). [cited by applicant]
Wang et al., 2023, “Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models,” arXiv:2308.11521v1 [cs.CL] Aug. 16, 2023 (15 pages). [cited by applicant]
Shayegani et al., 2023, “Survey of Vulnerabilities in Large Language Models Revealedby Adversarial Attacks,” arXiv:2310.10844v1 [cs.CL] Oct. 16, 2023 (54 pages). [cited by applicant]
Bezymiannyi et al., 2023, “Filter for confidential information,” Electronics and Control Systems 4(78):21-25 (5 pages). [cited by applicant]
Wang et al., 2023, “Self-Guard: Empower the LLM to Safeguard Itself,” ACL Anthology, NAACL (21 pages). [cited by applicant]