IP Library › Granted Patent US 12,361,335
Granted Patent B1
US 12,361,335 · App. 19/015,660 · Granted Jul 15, 2025

Validating vector constraints of outputs generated by machine learning models

Inventors: Vishal Mysore (Mississauga, CA); Ramkumar Ayyadurai (Jersey City, NJ); Chamindra Desilva (London, GB)
Assignee: CITIBANK, N.A.
G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,335
App. No.
19/015,660
Filed
Jan 10, 2025
Granted
Jul 15, 2025
Kind
B1
Examiner
CHEN, ALAN S
Art Unit
2125
USPC
706/12
Abstract

The technology evaluates the compliance of an AI application with predefined vector constraints. The technology employs multiple specialized models trained to identify specific types of non-compliance with the vector constraints within AI-generated responses. One or more models evaluate the existence of certain patterns within responses generated by an AI model by analyzing the representation of the attributes within the responses. Additionally, one or more models can identify vector representations of alphanumeric characters in the AI model's response by assessing the alphanumeric character's proximate locations, frequency, and/or associations with other alphanumeric characters. Moreover, one or more models can determine indicators of vector alignment between the vector representations of the AI model's response and the vector representations of the predetermined characters by measuring differences in the direction or magnitude of the vector representations.

Claims (85)

1. A non-transitory, computer-readable storage medium storing instructions for evaluating and correcting responses generated by one or more artificial intelligence (AI) models, wherein the instructions when executed by at least one data processor of a system, cause the system to:

train a machine learning (ML) model on a training dataset including predetermined alphanumeric characters to, in response to an input:

measure a set of differences in one or more of: direction or magnitude between one or more vector representations of alphanumeric characters in the input and one or more vector representations of the predetermined alphanumeric characters,

determine whether a volume of the set of differences satisfies a predetermined threshold, and

responsive to the volume of the set of differences satisfying the predetermined threshold, generate an output that indicates a presence of the one or more vector representations of the alphanumeric characters within the input;

for each of the vector representations of the predetermined alphanumeric characters, using the trained ML model, construct a set of validation actions configured to test the presence of the one or more vector representations of the predetermined alphanumeric characters within a set of responses of an AI model,

wherein each validation action includes: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result, and

wherein the set of validation actions is constructed by:

(1) determining a set of specific use cases associated with a set of guidelines defining a set of operative boundaries of the AI model, and

(2) mapping the set of specific use cases to the set of validation actions;

using the trained ML model, execute each set of validation actions on the AI model to determine the presence of the one or more vector representations of predetermined alphanumeric characters within the set of responses of the AI model by:

inputting, into the AI model, the test command set of each validation action to receive a first test output comprising (1) a set of test results and (2) a set of test descriptors associated with a second series of steps to generate the set of test results,

comparing, for each validation action, (1) the set of test results and (2) the set of test descriptors of the AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively, and

aggregating the comparisons of each validation action to generate a result indicating the presence of the one or more vector representations of predetermined alphanumeric characters within the first test output of the AI model;

using the result, generate a set of corrective actions to remove a portion of the first test output generated by the AI model indicated by the presence of the one or more vector representations of predetermined alphanumeric characters;

presenting a representation including one or more of: a graphical user interface component or a set of text via a computing device, wherein the representation indicates at least one of: the result or the set of corrective actions;

responsive to a user input received via the computing device, automatically executing the set of corrective actions to remove the portion of the first test output generated by the AI model; and

transmit the test command set of one or more validation actions into one or more nodes of an input layer of the AI model to validate an absence of the one or more vector representations of predetermined alphanumeric characters within a second test output generated by the AI model.

2. The non-transitory, computer-readable storage medium of claim 1 , wherein the ML model is trained to use (1) an order and (2) a timing of appearance of the alphanumeric characters within the set of responses to determine proximate locations of the alphanumeric characters within the set of responses of the AI model.

3. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to:

weigh the one or more vector representations of the alphanumeric characters within the set of responses of the AI model based on predetermined weights corresponding with each of the one or more vector representations of the alphanumeric characters,

wherein the output includes an overall score aggregating the weighted one or more vector representations of the alphanumeric characters.

4. The non-transitory, computer-readable storage medium of claim 1 ,

wherein the set of validation actions constructed by the trained ML model is ordered based on a complexity of the set of specific use cases determined from the one or more vector representations of the alphanumeric characters, and

wherein subsequently constructed validation actions are progressively more complex than preceding validation actions.

5. The non-transitory, computer-readable storage medium of claim 1 ,

wherein the set of validation actions constructed by the trained ML model is categorized based on an indicator of vector alignment, and

wherein the indicator of the vector alignment includes one or more of: complete alignment, partial alignment, or misalignment.

6. The non-transitory, computer-readable storage medium of claim 1 , wherein the training dataset includes unstructured alphanumeric characters, wherein the instructions further cause the system to:

extract the predetermined alphanumeric characters from the unstructured alphanumeric characters by identifying and isolating the predetermined alphanumeric characters from surrounding unstructured alphanumeric characters.

7. The non-transitory, computer-readable storage medium of claim 1 , wherein the instructions further cause the system to evaluate one or more of:

(1) proximate locations of the alphanumeric characters within the set of responses of the AI model,

(2) a frequency of the alphanumeric characters within the set of responses of the AI model, or

(3) an association between the alphanumeric characters within the set of responses of the AI model.

8. The non-transitory, computer-readable storage medium of claim 1 , wherein the representation is a graphical layout displayed via a graphical user interface (GUI).

9. A computing system comprising:

at least one processor; and

one or more non-transitory computer-readable media storing instructions, which when executed by at least one processor, perform operations comprising:

for one or more vector representations of a set of predetermined alphanumeric characters, using a trained machine learning (ML) model, constructing a set of validation actions configured to test a presence of the one or more vector representations of the set of predetermined alphanumeric characters within a set of responses of an artificial intelligence (AI) model,

wherein the trained ML model is trained to measure a set of differences between one or more vector representations of alphanumeric characters in the set of responses and the one or more vector representations of the set of predetermined alphanumeric characters, and

wherein one or more validation actions include: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result;

using the trained ML model, executing one or more sets of validation actions on the AI model to generate a result indicating the presence of the one or more vector representations of the predetermined alphanumeric characters within the set of responses of the AI model by:

inputting, into the AI model, the test command set of the one or more sets of validation actions to receive a first test output comprising (1) a set of test results and (2) a set of test descriptors associated with a second series of steps to generate the set of test results, and

comparing, for one or more sets of validation actions, (1) the set of test results and (2) the set of test descriptors of the AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively; and

responsive to the result exceeding a predefined threshold:

using the result, generating a set of corrective actions to remove a portion of the first test output generated by the AI model indicated by the one or more certain patterns,

displaying, at a graphical user interface (GUI), a graphical layout including (1) a first graphical representation indicating the result and (2) a second graphical representation indicating the set of corrective actions,

responsive to receiving a user input, automatically executing the set of corrective actions to remove the portion of the first test output generated by the AI model, and

transmit the test command set of one or more validation actions into one or more nodes of an input layer of the AI model to validate an absence of the one or more vector representations of predetermined alphanumeric characters within a second test output generated by the AI model.

10. The system of claim 9 , wherein the ML model is trained to use temporal dependencies between (1) an order and (2) a timing of appearance of the alphanumeric characters within the set of responses to determine proximate locations of the alphanumeric characters within the set of responses.

11. The system of claim 9 , further comprising:

weighing the one or more vector representations of the set of alphanumeric characters within the set of responses based on predetermined weights corresponding with each of the one or more vector representations of the set of alphanumeric characters,

wherein the result includes an overall score aggregating the weighted one or more vector representations of the set of alphanumeric characters.

12. The system of claim 9 , wherein constructing the set of validation actions comprises:

determining a set of validation criteria associated with the AI model using one or more metadata tags of the AI model, and

filtering stored validation actions using the determined set of validation criteria.

13. The system of claim 9 , wherein the trained ML model is trained on a training dataset including one or more of:

a domain-specific set of characters associated with a set of specific use cases of the AI model, or

one or more variations of a predetermined character sequence.

14. The system of claim 9 , further comprising:

categorizing a sequence of characters in a category in a set of categories using an associated semantic attribute, and

using the ML model to identify one or more patterns specific to each category.

15. A computer-implemented method for evaluating and correcting responses generated by one or more artificial intelligence (AI) models, the method comprising:

for one or more vector representations of a set of predetermined alphanumeric characters, using a trained first AI model, constructing a set of validation actions configured to test a presence of the one or more vector representations of the set of predetermined alphanumeric characters within a first set of responses of a second AI model,

wherein the first AI model is trained to measure a set of differences between the one or more vector representations of alphanumeric characters in the first set of responses and the one or more vector representations of the set of predetermined alphanumeric characters, and

wherein one or more validation actions include: (1) a test command set, (2) an expected result, and (3) an expected set of descriptors associated with an expected series of steps to generate the expected result;

using the trained first AI model, inputting one or more test command sets into the second AI model to generate a result indicating the presence of the one or more vector representations of the set of predetermined alphanumeric characters within the first set of responses of the second AI model by comparing (1) a set of test results and (2) a set of test descriptors of the second AI model with (1) the expected result and (2) the expected set of descriptors of a corresponding validation action, respectively; and

transmitting, via a computing device, a representation indicating the result;

responsive to an input, trigger execution of a set of actions to modify the first set of responses generated by the second AI model; and

transmit the test command set of the one or more sets of validation actions into one or more nodes of an input layer of an second AI model to validate an absence of the one or more vector representations within a second set of responses generated by the second AI model.

16. The computer-implemented method of claim 15 , further comprising:

receiving an indicator of a type of application associated with the second AI model,

identifying a relevant set of predetermined alphanumeric characters associated with the type of the application defining one or more operation boundaries of the second AI model, and

obtaining the relevant set of predetermined alphanumeric characters, via an Application Programming Interface (API).

17. The computer-implemented method of claim 15 , further comprising:

maintaining a cache including previous results of the second AI model, and

using the cache to identify a set of patterns associated with the presence of the one or more vector representations of the set of predetermined alphanumeric characters.

18. The computer-implemented method of claim 15 , further comprising:

receiving a set of guidelines defining operational boundaries for the second AI model, and

generating the set of validation actions using the received set of guidelines.

19. The computer-implemented method of claim 15 , further comprising:

mapping the result to a set of categories using a specific use case of the second AI model,

generating an indication of compliance for each category, and

aggregating the generated indicators to generate an overall compliance indicator.

20. The method of claim 15 , wherein one or more of the first AI model or the second AI model is an LLM.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: MYSORE, VISHAL; AYYADURAI, RAMKUMAR; DESILVA, CHAMINDRA
To: CITIBANK, N.A.
Reel/Frame 069840/0001 →
Continuity (2)
Division 18653858 · May 2, 2024
Continuation In Part 18637362 · Apr 16, 2024
References Cited (98)
US 9842045B2 · Heorhiadi et al. · 2017 [cited by applicant]
US 10324827B2 · Narayanan et al. · 2019 [cited by applicant]
US 10949337B1 · Yalla et al. · 2021 [cited by applicant]
US 11106801B1 · Levine et al. · 2021 [cited by applicant]
US 11227047B1 · Vashisht et al. · 2022 [cited by applicant]
US 11449798B2 · Olgiati et al. · 2022 [cited by applicant]
US 11573848B2 · Linck et al. · 2023 [cited by applicant]
US 11636027B2 · Sloane · 2023 [cited by applicant]
US 11652839B1 · Aloisio et al. · 2023 [cited by applicant]
US 11656852B2 · Mazurskiy · 2023 [cited by applicant]
US 11681811B1 · Dixit · 2023 [cited by applicant]
US 11683333B1 · Dominessy et al. · 2023 [cited by applicant]
US 11750717B2 · Walsh et al. · 2023 [cited by applicant]
US 11875123B1 · Ben David et al. · 2024 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11924027B1 · Mysore et al. · 2024 [cited by applicant]
US 11947435B2 · Boulineau et al. · 2024 [cited by applicant]
US 11960386B2 · Indani et al. · 2024 [cited by applicant]
US 11960515B1 · Pallakonda et al. · 2024 [cited by applicant]
US 11983806B1 · Ramesh et al. · 2024 [cited by applicant]
US 11990139B1 · Sandrew · 2024 [cited by applicant]
US 11995412B1 · Mishra · 2024 [cited by applicant]
US 12001463B1 · Pallakonda et al. · 2024 [cited by applicant]
US 12026599B1 · Lewis et al. · 2024 [cited by applicant]
US 20170262164A1 · Jain et al. · 2017 [cited by applicant]
US 20180089252A1 · Long et al. · 2018 [cited by applicant]
US 20180095866A1 · Narayanan et al. · 2018 [cited by applicant]
US 20190079854A1 · Lassance Oliveira E Silva et al. · 2019 [cited by applicant]
US 20200043164A1 · Fuchs et al. · 2020 [cited by applicant]
US 20200334326A1 · Zhang et al. · 2020 [cited by applicant]
US 20210012486A1 · Huang et al. · 2021 [cited by applicant]
US 20210097433A1 · Olgiati et al. · 2021 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220179906A1 · Desai et al. · 2022 [cited by applicant]
US 20220198304A1 · Szczepanik et al. · 2022 [cited by applicant]
US 20220311681A1 · Palladino et al. · 2022 [cited by applicant]
US 20220318654A1 · Lin et al. · 2022 [cited by applicant]
US 20220358023A1 · Moser et al. · 2022 [cited by applicant]
US 20220366140A1 · Saito et al. · 2022 [cited by applicant]
US 20230009999A1 · Higuchi et al. · 2023 [cited by applicant]
US 20230019072A1 · Okunlola · 2023 [cited by applicant]
US 20230028339A1 · Sloane · 2023 [cited by applicant]
US 20230033317A1 · Lin et al. · 2023 [cited by applicant]
US 20230039855A1 · Greene · 2023 [cited by applicant]
US 20230076795A1 · Indani et al. · 2023 [cited by applicant]
US 20230113621A1 · Griffin et al. · 2023 [cited by applicant]
US 20230164158A1 · Fellows et al. · 2023 [cited by applicant]
US 20230171282A1 · Bollinger · 2023 [cited by applicant]
US 20230177441A1 · Durvasula et al. · 2023 [cited by applicant]
US 20230177613A1 · Crabtree et al. · 2023 [cited by applicant]
US 20230252393A1 · Orzechowski et al. · 2023 [cited by applicant]
US 20230269272A1 · Dambrot et al. · 2023 [cited by applicant]
US 20230359789A1 · Andre et al. · 2023 [cited by applicant]
US 20240020538A1 · Socher et al. · 2024 [cited by applicant]
US 20240095077A1 · Singh et al. · 2024 [cited by applicant]
US 20240129345A1 · Kassam et al. · 2024 [cited by applicant]
US 20240144082A1 · Tarapov et al. · 2024 [cited by applicant]
US 20240202442A1 · Saito et al. · 2024 [cited by applicant]
US 20240346283A1 · Ayachitula et al. · 2024 [cited by applicant]
US 20240370476A1 · Madisetti et al. · 2024 [cited by applicant]
CN 106502890A · 2017 [cited by applicant]
WO 2022125803A1 · 2022 [cited by applicant]
WO 2024020416A1 · 2024 [cited by applicant]
Aka et al., Measuring Model Biases in the Absence of Ground Truth, AIES '21, May 19-21, 2021; pp. 327-335 (Year: 2021). [cited by examiner]
AI Risk Management Framework NIST, retrieved on Jun. 17, 2024, https://www.nist.gov/itl/ai-risk-management-framework. [cited by applicant]
Empower Your Team with a Compliance Co-Pilot, Sedric, retrieved on Sep. 25, 2024. https://www.sedric.ai/. [cited by applicant]
Independent analysis of AI language models and API providers. Artificial Analysis, retrieved on Jun. 13, 2024, https://artificialanalysis.ai/, 11 pages. [cited by applicant]
What is AI Verify?, AI Verify Foundation, Jun. 11, 2024, 3 pages, https://aiverifyfoundation.sg/. [cited by applicant]
Brown, D., et al., “The Great AI Challenge: We Test Five Top Bots on Useful, Everyday Skills,” The Wall Street Journal, published May 25, 2024. [cited by applicant]
Cranium, Adopt & Accelerate AI Safely, retrieved on Nov. 7, 2024, from https://cranium.ai/. [cited by applicant]
Dong, Y., et al., “Building Guardrails for Large Language Models,” https://ar5iv.labs.arxiv.org/html/2402.01822v1, published May 29, 2024, 20 pages. [cited by applicant]
Futurism, “Sam Altman Admits That OpenAI Doesn't Actually Understand How Its AI Works”, Jun. 11, 2024, 4 pages, https://futurism.com/sam-altman-admits-openai-understand-ai. [cited by applicant]
Generative machine learning models; IPCCOM000272835D, Aug. 17, 2023. (Year: 2023). [cited by applicant]
Guldimann, P., et al. “COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act,” arXiv:2410.07959v1 [cs.CL] Oct. 10, 2024, 38 pages. [cited by applicant]
Hu, Q., J., et al., “ROUTERBENCH: A Benchmark for Multi-LLM Routing System,” arXiv:2403.12031v2 [cs.LG] Mar. 28, 2024, 16 pages. [cited by applicant]
International Search Report and Written Opinion received in Application No. PCT/US24/47571, dated Dec. 9, 2024, 10 pages. [cited by applicant]
Kojima, Takeshi, et al. “Large Language Models are Zero-Shot Reasoners,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2205.11916 [cs.CL], Jan. 29, 2023, 42 pages. [cited by applicant]
Mathews, A. W., “What AI Can Do in Healthcare—and What It Should Never Do,” The Wall Street Journal, published on Aug. 21, 2024, retrieved on Sep. 5, 2024. https://www.wsj.com. [cited by applicant]
Mavrepis, P., et al., “XAI for All: Can Large Language Models Simplify Explainable AI?,” https://arxiv.org/abs/2401.13110, Jan. 23, 2024, 10 pages. [cited by applicant]
Mollick, E., “Latent Expertise: Everyone is in R&D,” One Useful Thing, published on Jun. 20, 2024, https://www.oneusefulthing.org/p/latent-expertise-everyone-is-in-r. [cited by applicant]
Nauta, M., et al., “From Anecdotal Evidence to Quantative Evaluation Methods: A Systematic Review of Evaluating Explainable AI” ACM Computing Surveys, vol. 55 No. 13s Article 295, 2023 [retrieved Jul. 3, 2024]. [cited by applicant]
Peers, M., “What California AI Bill Could Mean,” The Briefing, published and retrieved Aug. 30, 2024, 8 pages, https://www.theinformation.com/articles/what-california-ai-bill-could-mean. [cited by applicant]
Wei, Jason, et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2201.11903 [cs.CL], Jan. 10, 2023, 43 pages. [cited by applicant]
Zhao, H., et al., “Explainability for Large Language Models: A Survey,” https://arxiv.org/abs/2309.01029, Nov. 28, 2024, 38 pages. [cited by applicant]
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T., & Yu, P. S. (2024). Trustworthiness in Retrieval-Augmented Generation Systems: A Survey. ArXiv./abs/2409.10102. [cited by applicant]
Aggarwal, Nitin, “Why measuring your new Al is essential to its succes”, KPIs for gen AI: Why measuring your new AI is essential to its succes, 7 pages. [cited by applicant]
ANTHROP/C, “Mapping the Mind of a Large Language Model”, Mapping the Mind of a Large Language Model, May 21, 2024. [cited by applicant]
Claburn, Thomas, “OpenAI's GPT-4 can exploit real vulnerabilities by reading security advisories”, OpenAI's GPT-4 can exploit real vulnerabilities by reading security advisories, Apr. 17, 2024, 3 pages. [cited by applicant]
Marshall, Andrew, “Threat Modeling AI/ML Systems and Dependencies”, Threat Modeling AI/ML Systems and Dependencies, Nov. 2, 2022, 27 pages. [cited by applicant]
Roose, Kevin, “A.I. Has a Measurement Problem”, A.I. Has a Measurement Problem, Apr. 15, 2024, 5 pages. [cited by applicant]
Roose, Kevin, “A.I.'s Black Boxes Just Got a Little Less Mysterious”, A.I.'s Black Boxes Just Got a Little Less Mysterious, May 21, 2024, 5 pages. [cited by applicant]
Shah, Harshay, “Decomposing and Editing Predictions by Modeling Model Computation”, Decomposing and Editing Predictions by Modeling Model Computation, 5 pages. [cited by applicant]
Shankar, Ram, “Failure Modes in Machine Learning”, Nov. 2019, 14 pages. [cited by applicant]
Teo, Josephine, “Singapore launches Project Moonshot”, Singapore launches Project Moonshot—a generative Artificial Intelligence testing toolkit to address LLM safety and security challenges, May 31, 2024, 8 pages. [cited by applicant]
Aka et al., Measuring Model Biases in the Absence of Ground Truth, AIES '21, May 19-21, 2021, Virtual Event, USA.; pp. 327-335 (Year: 2021). [cited by applicant]
Lai et al., Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies, arXiv:2112.11471v1 [cs.AI] Dec. 21, 2021; Total pp. 36 (Year: 2021). [cited by applicant]
Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools, 37th Conference on Neural Information Processing Systems (NeurIPS 2023); Total pp. 13 (Year: 2023). [cited by applicant]
Yuan et al., R-Judge: Benchmarking Safety Risk Awareness for LLM Agents, arXiv:2401.10019v1 [cs.CL] Jan. 18, 2024; Total pp. 23 (Year: 2024). [cited by applicant]