IP Library Granted Patent US 12,596,839
Granted Patent B2
US 12,596,839 · App. 18/787,889 · Granted Apr 7, 2026

Selective redaction of personally identifiable information in generative artificial intelligence model outputs

Inventors: Tanner Burns (Austin, TX); Kwesi Cappel (Austin, TX); Kenneth Yeung (Ottawa, CA)
Assignee: HiddenLayer, Inc.
G06F21/6245G06F40/166G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,839
App. No.
18/787,889
Filed
Jul 29, 2024
Granted
Apr 7, 2026
Kind
B2
Art Unit
2435
USPC
726/26
Abstract

An output of a generative artificial intelligence (GenAI) model is received which is responsive to a prompt by a requestor. The output is tokenized to result in a plurality of tokens. These tokens are then used to determine that the output includes at least one string comprising personally identifiable information (PII). This determined can use pattern recognition to identify tokens and sequence of tokens indicative of PII. Thereafter, a classifier is used to assign a PII type to each string in the output comprising PII. It is then determined that at least one of the PII types in the output requires redaction which results in strings having a PII type determined to require redaction to be redacted which, in turn, results in a modified output for transmission to the requester. Related apparatus, systems, techniques and articles are also described.

Claims (50)

1 . A computer-implemented method comprising:

receiving an output of a generative artificial intelligence (GenAI) model responsive to a prompt by a requestor, the output being intercepted by a proxy and transmitted to a remote computing device;

tokenizing the output to result in a plurality of tokens;

determining, using the tokens, that the output comprises at least one string comprising personally identifiable information (PII), the determining using pattern recognition to identify tokens and sequence of tokens indicative of PII;

initiating at least one remediation action responsive to the determination comprising:

assigning, using a machine learning-based classifier, a PII type to each string in the output comprising PII;

determining that at least one of the PII types in the output requires redaction; and

redacting the strings having a PII type determined to require redaction to result in a modified output; and

causing the modified output to be transmitted to the requestor.

2 . The method of claim 1 , wherein the determination that least one of the PII types in the output requires redaction is based on a policy associated with the requestor.

3 . The method of claim 2 , wherein the policy is unique to the requestor and is one of a plurality of different available policies.

4 . The method of claim 3 , wherein the policy is one of a plurality of different available policies and is based on a class of users which includes the requestor.

5 . The method of claim 1 further comprising:

determining that at least one of the PII types in the output does not requires redaction;

wherein strings corresponding to the PII types not requiring redaction are not modified as part of the modified output.

6 . The method of claim 1 , wherein the GenAI model comprises a large language model.

7 . The method of claim 1 , wherein the proxy is executing further comprising:

intercepting, by a proxy executing in a computing environment in which the GenAI model executes, the output prior to the tokenizing.

8 . A system comprising:

a model computing environment comprising at least one computing device executing a generative artificial intelligence (GenAI) model and a proxy;

a monitoring computing environment comprising at least one computing device for monitoring the GenAI model by way of the proxy;

wherein a combination of the model computing environment and the monitoring environment execute operations comprising:

receiving, by the monitoring computing environment over a network from the proxy, an output of a generative artificial intelligence (GenAI) model responsive to a prompt by a requestor;

tokenizing the output to result in a plurality of tokens;

determining, using the tokens, that the output comprises at least one string comprising personally identifiable information (PII), the determining using pattern recognition to identify tokens and sequence of tokens indicative of PII;

assigning, using a classifier, a PII type to each string in the output comprising PII;

determining that at least one of the PII types in the output requires redaction;

redacting the strings having a PII type determined to require redaction to result in a modified output; and

causing the modified output to be transmitted to the requestor.

9 . The system of claim 8 , wherein the determination that least one of the PII types in the output requires redaction is based on a policy associated with the requestor.

10 . The system of claim 9 , wherein the policy is unique to the requestor and is one of a plurality of different available policies.

11 . The system of claim 10 , wherein the policy is one of a plurality of different available policies and is based on a class of users which includes the requestor.

12 . The system of claim 8 , wherein the operations further comprise:

determining that at least one of the PII types in the output does not requires redaction;

wherein strings corresponding to the PII types not requiring redaction are not modified as part of the modified output.

13 . The system of claim 8 , wherein the classifier comprises at least one machine learning model.

14 . The system of claim 8 , wherein the GenAI model comprises a large language model.

15 . A computer-implemented method comprising:

intercepting receiving a prompt from a requestor for ingestion by an artificial intelligence (AI) model executing in a model environment, the model environment comprising one or more computing devices;

determining, within a monitoring environment separate from the model environment and comprising one or more computing devices, whether the prompt comprises personally identifiable information (PII);

blocking the prompt for ingestion by the AI model when it is determined that the prompt comprises PII;

receiving, an output of the AI model responsive to the prompt by the monitoring environment, when it is determined that the prompt does not comprise PII;

determining, within the monitoring environment, whether the output comprises PII;

allowing the output to be transmitted to the requestor from the model environment if it is determined that the output does not comprise PII; and

selectively redacting the output before transmission from the model environment to the requestor based on a policy associated the requestor, the policy specifying levels of PII that require redaction for the requestor.

16 . The method of claim 15 , wherein a PII type is assigned to each string in the output comprising PII.

17 . The method of claim 16 , wherein the PII type is assigned by a machine learning-based classifier.

18 . The method of claim 15 , wherein the AI model is a large language model.

19 . The method of claim 15 , wherein the selective redaction is based on a policy associated with the requestor.

20 . The method of claim 19 , wherein the policy is one of a plurality of different available policies and is based on a class of users which includes the requestor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2024
From: BURNS, TANNER; CAPPEL, KWESI; YEUNG, KENNETH
To: HIDDENLAYER, INC.
Reel/Frame 068116/0753 →
Continuity (2)
Continuation 18621791 · Mar 29, 2024
Related Publication 20250307456A1 · Oct 2, 2025
References Cited (159)
US 7802298B1 · Hong et al. · 2010 [cited by applicant]
US 9350748B1 · McClintock et al. · 2016 [cited by applicant]
US 9356941B1 · Kislyuk et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu et al. · 2016 [cited by applicant]
US 10121104B1 · Hu et al. · 2018 [cited by applicant]
US 10193902B1 · Caspi et al. · 2019 [cited by applicant]
US 10205735B2 · Apostolopoulos · 2019 [cited by applicant]
US 10210036B2 · Iyer et al. · 2019 [cited by applicant]
US 10462168B2 · Shibahara et al. · 2019 [cited by applicant]
US 10637884B2 · Apple et al. · 2020 [cited by applicant]
US 10673880B1 · Pratt et al. · 2020 [cited by applicant]
US 10764313B1 · Mushtaq · 2020 [cited by applicant]
US 10803188B1 · Rajput et al. · 2020 [cited by applicant]
US 10824721B2 · Kesarwani et al. · 2020 [cited by applicant]
US 11310270B1 · Weber et al. · 2022 [cited by applicant]
US 11483327B2 · Hen et al. · 2022 [cited by applicant]
US 11501101B1 · Ganesan et al. · 2022 [cited by applicant]
US 11551137B1 · Echauz et al. · 2023 [cited by applicant]
US 11601468B2 · Angel et al. · 2023 [cited by applicant]
US 11710045B2 · Lee et al. · 2023 [cited by applicant]
US 11710067B2 · Harris et al. · 2023 [cited by applicant]
US 11762998B2 · Kuta et al. · 2023 [cited by applicant]
US 11777957B2 · Chen et al. · 2023 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11893111B2 · Sai et al. · 2024 [cited by applicant]
US 11893358B1 · Lakshmikanthan et al. · 2024 [cited by applicant]
US 11930030B1 · Burns et al. · 2024 [cited by applicant]
US 11930039B1 · Geethakumar et al. · 2024 [cited by applicant]
US 11954199B1 · Burns et al. · 2024 [cited by applicant]
US 11960514B1 · Taylert et al. · 2024 [cited by applicant]
US 11962546B1 · Hattangady et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11972333B1 · Horesh et al. · 2024 [cited by applicant]
US 11995180B1 · Cappel et al. · 2024 [cited by applicant]
US 11997059B1 · Su et al. · 2024 [cited by applicant]
US 12026255B1 · Burns et al. · 2024 [cited by applicant]
US 12107885B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12111926B1 · Beveridge et al. · 2024 [cited by applicant]
US 12124592B1 · O'Hern et al. · 2024 [cited by applicant]
US 12130917B1 · Yeung et al. · 2024 [cited by applicant]
US 12130943B1 · Burns et al. · 2024 [cited by applicant]
US 12137118B1 · Kawasaki et al. · 2024 [cited by applicant]
US 12174954B1 · Yeung et al. · 2024 [cited by applicant]
US 12182264B2 · Sinha et al. · 2024 [cited by applicant]
US 12197483B1 · Shmukler · 2025 [cited by examiner]
US 12197859B1 · Malviya et al. · 2025 [cited by applicant]
US 12204323B1 · Malviya et al. · 2025 [cited by applicant]
US 12229265B1 · Yeung et al. · 2025 [cited by applicant]
US 12248883B1 · Rideout et al. · 2025 [cited by applicant]
US 12293277B1 · Yeung et al. · 2025 [cited by applicant]
US 20100082811A1 · Van Der Merwe · 2010 [cited by applicant]
US 20140033307A1 · Schmidtler · 2014 [cited by applicant]
US 20140157415A1 · Abercrombie et al. · 2014 [cited by applicant]
US 20150074392A1 · Boivie et al. · 2015 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170154021A1 · Vidhani et al. · 2017 [cited by applicant]
US 20170251006A1 · LaRosa et al. · 2017 [cited by applicant]
US 20170331841A1 · Hu et al. · 2017 [cited by applicant]
US 20180018590A1 · Szeto et al. · 2018 [cited by applicant]
US 20180063190A1 · Wright et al. · 2018 [cited by applicant]
US 20180205734A1 · Wing et al. · 2018 [cited by applicant]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20180324193A1 · Ronen et al. · 2018 [cited by applicant]
US 20190050564A1 · Pogorelik et al. · 2019 [cited by applicant]
US 20190238568A1 · Goswami et al. · 2019 [cited by applicant]
US 20190238572A1 · Manadhata et al. · 2019 [cited by applicant]
US 20190260784A1 · Stockdale et al. · 2019 [cited by applicant]
US 20190311118A1 · Grafi et al. · 2019 [cited by applicant]
US 20190392176A1 · Taron et al. · 2019 [cited by applicant]
US 20200019721A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200076771A1 · Maier et al. · 2020 [cited by applicant]
US 20200092299A1 · Srinivasan et al. · 2020 [cited by applicant]
US 20200175094A1 · Palmer et al. · 2020 [cited by applicant]
US 20200219009A1 · Dao et al. · 2020 [cited by applicant]
US 20200233979A1 · Maraghoosh et al. · 2020 [cited by applicant]
US 20200279192A1 · Godfrey et al. · 2020 [cited by applicant]
US 20200285737A1 · Kraus et al. · 2020 [cited by applicant]
US 20200313849A1 · Kar et al. · 2020 [cited by applicant]
US 20200364333A1 · Derks et al. · 2020 [cited by applicant]
US 20200403826A1 · Dawani et al. · 2020 [cited by applicant]
US 20200409323A1 · Spalt et al. · 2020 [cited by applicant]
US 20210110062A1 · Oliner et al. · 2021 [cited by applicant]
US 20210141897A1 · Seifert et al. · 2021 [cited by applicant]
US 20210209464A1 · Bala et al. · 2021 [cited by applicant]
US 20210218673A1 · Ma et al. · 2021 [cited by applicant]
US 20210224425A1 · Nasr-Azadani et al. · 2021 [cited by applicant]
US 20210303695A1 · Grosse et al. · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik et al. · 2021 [cited by applicant]
US 20210319784A1 · Le Roux et al. · 2021 [cited by applicant]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210374247A1 · Sultana et al. · 2021 [cited by applicant]
US 20210407051A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220030009A1 · Hasan · 2022 [cited by applicant]
US 20220058444A1 · Olabiyi et al. · 2022 [cited by applicant]
US 20220070195A1 · Sern et al. · 2022 [cited by applicant]
US 20220083658A1 · Shah et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147597A1 · Bhide et al. · 2022 [cited by applicant]
US 20220164444A1 · Prudkovskij et al. · 2022 [cited by applicant]
US 20220166795A1 · Simioni et al. · 2022 [cited by applicant]
US 20220182410A1 · Tupsamudre et al. · 2022 [cited by applicant]
US 20220253464A1 · Sloane et al. · 2022 [cited by applicant]
US 20220269796A1 · Chase et al. · 2022 [cited by applicant]
US 20220284283A1 · Yin et al. · 2022 [cited by applicant]
US 20220309179A1 · Payne et al. · 2022 [cited by applicant]
US 20230008037A1 · Venugopal et al. · 2023 [cited by applicant]
US 20230027149A1 · Kuan et al. · 2023 [cited by applicant]
US 20230049479A1 · Mozo Velasco et al. · 2023 [cited by applicant]
US 20230109426A1 · Hashimoto et al. · 2023 [cited by applicant]
US 20230111744A1 · Chandrasekaran et al. · 2023 [cited by applicant]
US 20230128947A1 · Bhaskar et al. · 2023 [cited by applicant]
US 20230148116A1 · Stokes et al. · 2023 [cited by applicant]
US 20230169397A1 · Smith et al. · 2023 [cited by applicant]
US 20230185912A1 · Sinn et al. · 2023 [cited by applicant]
US 20230185915A1 · Rao et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230229960A1 · Zhu et al. · 2023 [cited by applicant]
US 20230252178A1 · Ruelke et al. · 2023 [cited by applicant]
US 20230259787A1 · David et al. · 2023 [cited by applicant]
US 20230269263A1 · Yarabolu · 2023 [cited by applicant]
US 20230274003A1 · Liu et al. · 2023 [cited by applicant]
US 20230289604A1 · Chan et al. · 2023 [cited by applicant]
US 20230351143A1 · Kutt et al. · 2023 [cited by applicant]
US 20230359903A1 · Cefalu et al. · 2023 [cited by applicant]
US 20230359924A1 · Maman et al. · 2023 [cited by applicant]
US 20230385085A1 · Singh · 2023 [cited by applicant]
US 20230388324A1 · Thompson · 2023 [cited by applicant]
US 20240005690A1 · Brodie et al. · 2024 [cited by applicant]
US 20240007469A1 · Wang et al. · 2024 [cited by applicant]
US 20240022585A1 · Burns et al. · 2024 [cited by applicant]
US 20240031026A1 · Fujisawa et al. · 2024 [cited by applicant]
US 20240039948A1 · Koc et al. · 2024 [cited by applicant]
US 20240045959A1 · Marson et al. · 2024 [cited by applicant]
US 20240054233A1 · Ohayon et al. · 2024 [cited by applicant]
US 20240078337A1 · Kamyshenko et al. · 2024 [cited by applicant]
US 20240080333A1 · Burns et al. · 2024 [cited by applicant]
US 20240126611A1 · Phanishayee et al. · 2024 [cited by applicant]
US 20240127065A1 · Ren et al. · 2024 [cited by applicant]
US 20240160902A1 · Padgett · 2024 [cited by examiner]
US 20240289628A1 · Parmar et al. · 2024 [cited by applicant]
US 20240289863A1 · Smith Lewis et al. · 2024 [cited by applicant]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240386103A1 · Clement et al. · 2024 [cited by applicant]
US 20240414177A1 · Lal et al. · 2024 [cited by applicant]
US 20240427986A1 · Shakarian et al. · 2024 [cited by applicant]
US 20250086455A1 · Yang et al. · 2025 [cited by applicant]
US 20250112768A1 · Ho · 2025 [cited by examiner]
Dinan et al., 2021, “Anticipating safety issues in e2e conversational ai: Framework and tooling,” arXiv preprint arXiv:2107.03451v3 (43 pages). [cited by applicant]
Morozov et al., 2019, “Unsupervised Neural Quantization for Compressed-Domain Similarity Search,” International Conference on Computer Vision (ICCV) 2019 (11 pages). [cited by applicant]
Rijthoven et al., 2021, “HookNet: Multi-resolution convulational neural networks for semantic segmentation in histopathology whole-slide images,” Medical Imange Analysis 68:1-10. [cited by applicant]
Wang et al., 2023, “Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models,” arXiv:2308.11521v1 [cs.CL] Aug. 16, 2023 (15 pages). [cited by applicant]
Shayegani et al., 2023, “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks,” arXiv:2310.10844v1 [cs.CL] Oct. 16, 2023 (54 pages). [cited by applicant]
Bezymiannyi et al., 2023, “Filter for confidential information,” Electronics and Control Systems 4(78):21-25 (5 pages). [cited by applicant]
Wang et al., 2023, “Self-Guard: Empower the LLM to Safeguard Itself,” ACL Anthology, NAACL (21 pages). [cited by applicant]
Automorphic.ai, 2024, “Github - automorphic-ai/aegis: Self-hardening firewall for large language models,” XP093278213, Available online at https://web.archive.org/web/20240222171700/https://github.com/automorphic-ai/aeg… [cited by applicant]
Sun et al., “CONSCENDI: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants,” arXiv:2304.14364v1 [cs.CL] Apr. 27, 2023 (20 pages). [cited by applicant]
Hu et al., “Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes,” arXiv:2403.00867v1 [cs.CR] Mar. 1, 2024 (19 pages). [cited by applicant]
Robey et al., “SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks,” arXiv:2310.03684v3 [cs.LG] Nov. 29, 2023 (42 pages). [cited by applicant]
Mohtashami et al., “Social Learning: Towards Collaborative Learning with Large Language Models,” arXiv:2312.11441v2 [cs.LG] Feb. 8, 2024 (19 pages). [cited by applicant]