IP Library › Granted Patent US 12,361,334
Granted Patent B1
US 12,361,334 · App. 19/015,646 · Granted Jul 15, 2025

Validating vector constraints of outputs generated by machine learning models

Inventors: Vishal Mysore (Mississauga, CA); Ramkumar Ayyadurai (Jersey City, NJ); Chamindra Desilva (London, GB)
Assignee: CITIBANK, N.A.
G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,334
App. No.
19/015,646
Granted
Jul 15, 2025
Kind
B1
Abstract

The technology evaluates the compliance of an AI application with predefined vector constraints. The technology employs multiple specialized models trained to identify specific types of non-compliance with the vector constraints within AI-generated responses. One or more models evaluate the existence of certain patterns within responses generated by an AI model by analyzing the representation of the attributes within the responses. Additionally, one or more models can identify vector representations of alphanumeric characters in the AI model's response by assessing the alphanumeric character's proximate locations, frequency, and/or associations with other alphanumeric characters. Moreover, one or more models can determine indicators of vector alignment between the vector representations of the AI model's response and the vector representations of the predetermined characters by measuring differences in the direction or magnitude of the vector representations.

Claims (92)

1. A non-transitory, computer-readable storage medium storing instructions for evaluating responses generated by one or more artificial intelligence (AI) models, wherein the instructions when executed by at least one data processor of a system, cause the system to:

train a machine learning (ML) model on a training dataset including a set of attributes to, in response to an input, generate an output that identifies a presence of one or more certain patterns of the set of attributes in the input,

wherein each certain pattern represents a disproportionate association of one or more attributes of the set of attributes within the input;

for each certain pattern, using the trained ML model, construct a set of validation actions for a set of responses generated by an AI model,

wherein each validation action includes: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result, and

wherein the set of validation actions is constructed by:

(1) determining a set of specific use cases associated with a set of guidelines defining a set of operative boundaries of the AI model, and

(2) mapping the set of specific use cases to the set of validation actions;

using the trained ML model, execute each set of validation actions on the AI model to determine the presence of one or more certain patterns within the set of responses generated by the AI model by:

inputting, into the AI model, the test command set of each validation action to receive a first test output comprising (1) a set of test results and (2) a set of test descriptors associated with a second series of steps to generate the set of test results,

comparing, for each validation action, (1) the set of test results and (2) the set of test descriptors of the AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively, and

aggregating the comparisons of each validation action to generate a result indicating the presence of one or more certain patterns within the first test output generated by the AI model;

using the result, generate a set of corrective actions to remove a portion of the first test output generated by the AI model indicated by the one or more certain patterns;

presenting a representation including one or more of: a graphical user interface component or a set of text via a computing device, wherein the representation indicates at least one of: the result or the set of corrective actions;

responsive to a user input received via the computing device, automatically executing the set of corrective actions to remove the portion of the first test output generated by the AI model; and

transmit the test command set of one or more validation actions into one or more nodes of an input layer of the AI model to validate an absence of the one or more certain patterns within a second test output generated by the AI model.

2. The non-transitory, computer-readable storage medium of claim 1 ,

wherein the output is a first output, and

wherein the ML model is further trained to generate a second output including a confidence score associated with a likelihood of the presence of the one or more certain patterns within the set of responses generated by the AI model.

3. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to:

identify one or more new patterns within at least one of: the set of test results or the set of test descriptors of the AI model, and

iteratively update the set of validation actions based on the one or more new patterns.

4. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to:

receive an indicator of a type of application associated with the AI model, identify a relevant set of attributes associated with the type of the application defining one or more operation boundaries of the AI model, and

obtain the relevant set of attributes, via an Application Programming Interface (API).

5. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to:

wherein the set of corrective actions include one or more of: adjusting parameters of the AI model or updating training data of the AI model to remove the portion of the response from the AI model indicated by the one or more certain patterns.

6. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to:

evaluate the set of attributes by analyzing one or more of:

proximate locations of alphanumeric characters within the set of responses, frequency of alphanumeric characters within the set of responses, or associations between alphanumeric characters within the set of responses.

7. The non-transitory, computer-readable storage medium of claim 6 , wherein the instructions further cause the system to:

segment the alphanumeric characters of the set of responses into a set of tokens;

normalize the set of tokens by removing one or more of: suffixes or prefixes of words within the alphanumeric characters; and

using the normalized set of tokens to generate the one or more certain patterns.

8. A computing system comprising:

at least one processor; and

one or more non-transitory computer-readable media storing instructions, which when executed by the at least one processor, perform operations comprising:

for each certain pattern of a set of certain patterns, constructing a set of validation actions for a set of responses generated by an AI model,

wherein each validation action includes: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result,

wherein the set of validation actions is constructed by:

(1) determining a set of specific use cases associated with a set of guidelines defining a set of operative boundaries of the AI model, and

(2) mapping the set of specific use cases to the set of validation actions, and

wherein each certain pattern represents a disproportionate association of one or more attributes of a set of attributes within the set of responses;

using a trained ML model, executing one or more sets of validation actions on the AI model to determine a presence of one or more certain patterns within the set of responses generated by the AI model by:

inputting, into the AI model, the test command set of the one or more sets of validation actions to receive a first test output comprising (1) a set of test results and (2) a set of test descriptors associated with a second series of steps to generate the set of test results,

comparing, for the one or more sets of validation actions, (1) the set of test results and (2) the set of test descriptors of the AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively, and

aggregating the comparisons of the one or more sets of validation actions to generate a result indicating the presence of the one or more certain patterns within the first test output generated by the AI model; and

responsive to the result exceeding a predefined threshold:

using the result, generating a set of corrective actions to remove a portion of the first test output generated by the AI model indicated by the one or more certain patterns,

displaying, at a graphical user interface (GUI), a graphical layout including (1) a first graphical representation indicating the result and (2) a second graphical representation indicating the set of corrective actions,

responsive to receiving a user input, automatically executing the set of corrective actions to remove the portion of the first test output generated by the AI model, and

transmit the test command set of one or more validation actions into one or more nodes of an input layer of the AI model to validate an absence of the one or more certain patterns within a second test output generated by the AI model.

9. The system of claim 8 , wherein the operations further comprise:

comparing, for one or more sets of validation actions, the expected result of the one or more sets of validation actions to a corresponding set of test results received from the AI model; and

responsive to the expected result of the one or more sets of validation actions satisfying the corresponding set of test results received from the AI model, comparing the expected set of descriptors of the one or more sets of validation actions to a corresponding set of test descriptors of the corresponding set of test results.

10. The system of claim 8 , wherein the operations further comprise:

receiving a set of feedback from a user on the result indicating the presence of one or more certain patterns within the set of responses generated by the AI model; and

adjusting operational parameters of the trained ML model using the set of feedback.

11. The system of claim 8 , wherein the operations further comprise:

classifying the one or more certain patterns into categories using a set of predefined criteria; and

assigning a corresponding set of validation actions to each category of certain patterns.

12. The system of claim 8 , wherein modifying the AI model further comprises:

updating training data of the AI model to remove a portion of the first test output generated by the AI model indicated by the one or more certain patterns.

13. The system of claim 8 ,

wherein the graphical layout includes a third graphical representation indicating the constructed one or more sets set of validation actions.

14. The system of claim 13 , wherein the operations further comprise:

responsive to a received user input, automatically executing the generated set of corrective actions to remove the portion of the set of responses generated by the AI model.

15. A computer-implemented method for evaluating responses generated by one or more AI models, the method comprising:

for each certain pattern of a set of certain patterns, constructing a set of validation actions for a set of responses generated by a second AI model, wherein one or more validation actions include: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result,

wherein each certain pattern represents a disproportionate association of one or more attributes of a set of attributes within the set of responses;

using a trained first AI model, executing one or more sets of validation actions on the second AI model to determine a presence of one or more certain patterns within the set of responses generated by the second AI model by:

inputting, into the second AI model, the test command set of the one or more sets of validation actions to receive a first test output comprising (1) a set of test results and (2) a set of test descriptors associated with a second series of steps to generate the set of test results,

comparing, for the sets of one or more validation actions, (1) the set of test results and (2) the set of test descriptors of the second AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively, and

aggregating the comparisons of the sets of one or more validation actions to generate a result indicating the presence of one or more certain patterns within the first test output generated by the second AI model;

transmitting, via a computing device, a representation indicating the result;

responsive to an input, trigger execution of a set of actions to modify the first test output generated by the second AI model; and

transmit the test command set of the one or more sets of validation actions into one or more nodes of an input layer of an second AI model to validate an absence of the one or more certain patterns within a second test output generated by the second AI model.

16. The computer-implemented method of claim 15 , further comprising:

displaying, on a graphical user interface (GUI), a graphical layout including (1) a first graphical representation indicating the result and (2) a second graphical representation indicating the constructed set of validation actions.

17. The computer-implemented method of claim 16 , further comprising:

responsive to a received user input on a graphical user interface (GUI) the GUI, automatically executing the constructed set of validation actions on the second AI model.

18. The computer-implemented method of claim 15 , further comprising:

modifying the second AI model by adjusting one or more parameters of the second AI model,

wherein the modified second AI model is trained to, in response to the test command set, generate, using the adjusted one or more parameters, a second set of responses without the presence of the one or more certain patterns.

19. The computer-implemented method of claim 15 , wherein the expected series of steps is a first series of steps, further comprising:

evaluating the second AI model against the set of validation actions by applying one or more validation actions in the set of validation actions to the second AI model by, for each particular validation action of the one or more validation actions:

supplying the test command set of the particular validation action into the second AI model,

responsive to inputting the test command set, receiving, from the second AI model, a test result and a corresponding test set of descriptors associated with a second series of steps to generate the test result,

comparing the expected result of the particular validation action to the test result received from the second AI model, and

responsive to the expected result of the particular validation action satisfying the test result received from the second AI model, comparing the expected set of descriptors of the particular validation action to the corresponding test set of descriptors of the test result.

20. The computer-implemented method of claim 15 , further comprising:

modifying the second AI model by updating training data of the second AI model to remove a portion of the set of responses generated by the second AI model indicated by the one or more certain patterns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: MYSORE, VISHAL; AYYADURAI, RAMKUMAR; DESILVA, CHAMINDRA
To: CITIBANK, N.A.
Reel/Frame 069837/0085 →
Continuity (2)
Division 18653858 · May 2, 2024
Continuation In Part 18637362 · Apr 16, 2024
References Cited (97)
US 9842045B2 · Heorhiadi et al. · 2017 [cited by applicant]
US 10324827B2 · Narayanan et al. · 2019 [cited by applicant]
US 10949337B1 · Yalla et al. · 2021 [cited by applicant]
US 11106801B1 · Levine et al. · 2021 [cited by applicant]
US 11227047B1 · Vashisht et al. · 2022 [cited by applicant]
US 11449798B2 · Olgiati et al. · 2022 [cited by applicant]
US 11573848B2 · Linck et al. · 2023 [cited by applicant]
US 11636027B2 · Sloane · 2023 [cited by applicant]
US 11652839B1 · Aloisio et al. · 2023 [cited by applicant]
US 11656852B2 · Mazurskiy · 2023 [cited by applicant]
US 11681811B1 · Dixit · 2023 [cited by applicant]
US 11683333B1 · Dominessy et al. · 2023 [cited by applicant]
US 11750717B2 · Walsh et al. · 2023 [cited by applicant]
US 11875123B1 · Ben David et al. · 2024 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11924027B1 · Mysore et al. · 2024 [cited by applicant]
US 11947435B2 · Boulineau et al. · 2024 [cited by applicant]
US 11960386B2 · Indani et al. · 2024 [cited by applicant]
US 11960515B1 · Pallakonda et al. · 2024 [cited by applicant]
US 11983806B1 · Ramesh et al. · 2024 [cited by applicant]
US 11990139B1 · Sandrew · 2024 [cited by applicant]
US 11995412B1 · Mishra · 2024 [cited by applicant]
US 12001463B1 · Pallakonda et al. · 2024 [cited by applicant]
US 12026599B1 · Lewis et al. · 2024 [cited by applicant]
US 20170262164A1 · Jain et al. · 2017 [cited by applicant]
US 20180089252A1 · Long et al. · 2018 [cited by applicant]
US 20180095866A1 · Narayanan et al. · 2018 [cited by applicant]
US 20190079854A1 · Lassance Oliveira E Silva et al. · 2019 [cited by applicant]
US 20200043164A1 · Fuchs et al. · 2020 [cited by applicant]
US 20200334326A1 · Zhang et al. · 2020 [cited by applicant]
US 20210012486A1 · Huang et al. · 2021 [cited by applicant]
US 20210097433A1 · Olgiati et al. · 2021 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220179906A1 · Desai et al. · 2022 [cited by applicant]
US 20220198304A1 · Szczepanik et al. · 2022 [cited by applicant]
US 20220311681A1 · Palladino et al. · 2022 [cited by applicant]
US 20220318654A1 · Lin et al. · 2022 [cited by applicant]
US 20220358023A1 · Moser et al. · 2022 [cited by applicant]
US 20220366140A1 · Saito et al. · 2022 [cited by applicant]
US 20230009999A1 · Higuchi et al. · 2023 [cited by applicant]
US 20230019072A1 · Okunlola · 2023 [cited by applicant]
US 20230028339A1 · Sloane · 2023 [cited by applicant]
US 20230033317A1 · Lin et al. · 2023 [cited by applicant]
US 20230039855A1 · Greene · 2023 [cited by applicant]
US 20230076795A1 · Indani et al. · 2023 [cited by applicant]
US 20230113621A1 · Griffin et al. · 2023 [cited by applicant]
US 20230164158A1 · Fellows et al. · 2023 [cited by applicant]
US 20230171282A1 · Bollinger · 2023 [cited by applicant]
US 20230177441A1 · Durvasula et al. · 2023 [cited by applicant]
US 20230177613A1 · Crabtree et al. · 2023 [cited by applicant]
US 20230252393A1 · Orzechowski et al. · 2023 [cited by applicant]
US 20230269272A1 · Dambrot et al. · 2023 [cited by applicant]
US 20230359789A1 · Andre et al. · 2023 [cited by applicant]
US 20240020538A1 · Socher et al. · 2024 [cited by applicant]
US 20240095077A1 · Singh et al. · 2024 [cited by applicant]
US 20240129345A1 · Kassam et al. · 2024 [cited by applicant]
US 20240144082A1 · Tarapov et al. · 2024 [cited by applicant]
US 20240202442A1 · Saito et al. · 2024 [cited by applicant]
US 20240346283A1 · Ayachitula et al. · 2024 [cited by applicant]
US 20240370476A1 · Madisetti et al. · 2024 [cited by applicant]
CN 106502890A · 2017 [cited by applicant]
WO 2022125803A1 · 2022 [cited by applicant]
WO 2024020416A1 · 2024 [cited by applicant]
Aka et al., Measuring Model Biases in the Absence of Ground Truth, AIES '21, May 19-21, 2021, Virtual Event, USA.; pp. 327-335 (Year: 2021). [cited by examiner]
AI Risk Management Framework NIST, retrieved on Jun. 17, 2024, https://www.nist.gov/itl/ai-risk-management-framework. [cited by applicant]
Empower Your Team with a Compliance Co-Pilot, Sedric, retrieved on Sep. 25, 2024. https://www.sedric.ai/. [cited by applicant]
Independent analysis of AI language models and API providers. Artificial Analysis, retrieved on Jun. 13, 2024, https://artificialanalysis.ai/, 11 pages. [cited by applicant]
What is AI Verify?, AI Verify Foundation, Jun. 11, 2024, 3 pages, https://aiverifyfoundation.sg/. [cited by applicant]
Brown, D., et al., “The Great AI Challenge: We Test Five Top Bots on Useful, Everyday Skills,” The Wall Street Journal, published May 25, 2024. [cited by applicant]
Cranium, Adopt & Accelerate AI Safely, retrieved on Nov. 7, 2024, from https://cranium.ai/. [cited by applicant]
Dong, Y., et al., “Building Guardrails for Large Language Models,” https://ar5iv.labs.arxiv.org/html/2402.01822v1, published May 29, 2024, 20 pages. [cited by applicant]
Futurism, “Sam Altman Admits That OpenAI Doesn't Actually Understand How Its AI Works”, Jun. 11, 2024, 4 pages, https://futurism.com/sam-altman-admits-openai-understand-ai. [cited by applicant]
Generative machine learning models; IPCCOM000272835D, Aug. 17, 2023. (Year: 2023). [cited by applicant]
Guldimann, P., et al. “COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act,” arXiv:2410.07959v1 [cs.CL] Oct. 10, 2024, 38 pages. [cited by applicant]
Hu, Q., J., et al., “Routerbench: A Benchmark for Multi-LLM Routing System,” arXiv:2403.12031v2 [cs.LG] Mar. 28, 2024, 16 pages. [cited by applicant]
International Search Report and Written Opinion received in Application No. PCT/US24/47571, dated Dec. 9, 2024, 10 pages. [cited by applicant]
Kojima, Takeshi, et al. “Large Language Models are Zero-Shot Reasoners,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2205.11916 [cs.CL], Jan. 29, 2023, 42 pages. [cited by applicant]
Mathews, A. W., “What AI Can Do in Healthcare—and What It Should Never Do,” The Wall Street Journal, published on Aug. 21, 2024, retrieved on Sep. 5, 2024. https://www.wsj.com. [cited by applicant]
Mavrepis, P., et al., “XAI for All: Can Large Language Models Simplify Explainable AI?,” https://arxiv.org/abs/2401.13110, Jan. 23, 2024, 10 pages. [cited by applicant]
Mollick, E., “Latent Expertise: Everyone is in R&D,” One Useful Thing, published on Jun. 20, 2024, https://www.oneusefulthing.org/p/latent-expertise-everyone-is-in-r. [cited by applicant]
Nauta, M., et al., “From Anecdotal Evidence to Quantative Evaluation Methods: A Systematic Review of Evaluating Explainable AI” ACM Computing Surveys, vol. 55 No. 13s Article 295, 2023 [retrieved Jul. 3, 2024]. [cited by applicant]
Peers, M., “What California AI Bill Could Mean,” The Briefing, published and retrieved Aug. 30, 2024, 8 pages, https://www.theinformation.com/articles/what-california-ai-bill-could-mean. [cited by applicant]
Wei, Jason, et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2201.11903 [cs.CL], Jan. 10, 2023, 43 pages. [cited by applicant]
Zhao, H., et al., “Explainability for Large Language Models: A Survey,” https://arxiv.org/abs/2309.01029, Nov. 28, 2024, 38 pages. [cited by applicant]
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T., & Yu, P. S. (2024). Trustworthiness in Retrieval-Augmented Generation Systems: A Survey. ArXiv./abs/2409.10102. [cited by applicant]
Aggarwal, Nitin , “Why measuring your new AI is essential to its succes”, KPIs for gen AI: Why measuring your new AI is essential to its succes, 7 pages. [cited by applicant]
ANTHROP/C , “Mapping the Mind of a Large Language Model”, Mapping the Mind of a Large Language Model, May 21, 2024. [cited by applicant]
Claburn, Thomas , “OpenAI's GPT-4 can exploit real vulnerabilities by reading security advisories”, OpenAI's GPT-4 can exploit real vulnerabilities by reading security advisories, Apr. 17, 2024, 3 pages. [cited by applicant]
Marshall, Andrew , “Threat Modeling AI/ML Systems and Dependencies”, Threat Modeling AI/ML Systems and Dependencies, Nov. 2, 2022, 27 pages. [cited by applicant]
Roose, Kevin , “A.I. Has a Measurement Problem”, A.I. Has a Measurement Problem, Apr. 15, 2024, 5 pages. [cited by applicant]
Roose, Kevin , “A.I.'s Black Boxes Just Got a Little Less Mysterious”, A.I.'s Black Boxes Just Got a Little Less Mysterious, May 21, 2024, 5 pages. [cited by applicant]
Shah, Harshay , “Decomposing and Editing Predictions by Modeling Model Computation”, Decomposing and Editing Predictions by Modeling Model Computation, 5 pages. [cited by applicant]
Shankar, Ram , “Failure Modes in Machine Learning”, , Nov. 2019, 14 pages. [cited by applicant]
Teo, Josephine , “Singapore launches Project Moonshot”, Singapore launches Project Moonshot—a generative Artificial Intelligence testing toolkit to address LLM safety and security challenges, May 31, 2024, 8 pages. [cited by applicant]
Lai et al., Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies, arXiv:2112.11471v1 [cs.AI] Dec. 21, 2021; Total pp. 36 (Year: 2021). [cited by applicant]
Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools, 37th Conference on Neural Information Processing Systems (NeurIPS 2023); Total pp. 13 (Year: 2023). [cited by applicant]
Yuan et al., R-Judge: Benchmarking Safety Risk Awareness for LLM Agents, arXiv:2401.10019v1 [cs.CL] Jan. 18, 2024; Total pp. 23 (Year: 2024). [cited by applicant]