IP Library › Granted Patent US 12,524,508
Granted Patent B2
US 12,524,508 · App. 19/188,116 · Granted Jan 13, 2026

Prompt refinement to improve accuracy of outputs from machine learning models

Inventors: Nigil Satish Jeyashekar (Irving, TX); Jason Engelbrecht (London, GB); Zheyu Wang (Shanghai, CN); Haolin Jin (Shanghai, CN); Payal Jain (London, GB); Tariq Husayn Maonah (London, GB); Mariusz Saternus (Cracow, PL); Daniel Lewandowski (Cracow, PL); Biraj Krushna Rath (London, GB); Stuart Murray (London, GB); Philip Davies (London, GB); Sourabh Deb (Tampa, FL)
G06F21/31G06F21/6218G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,508
App. No.
19/188,116
Granted
Jan 13, 2026
Kind
B2
Abstract

Systems and methods for restructuring prompts in order to improve accuracy of outputs from models are disclosed herein. The system receives a user prompt indicating a request for data. The system generates a first and second output using a model, the first output generated based on the user prompt and the second output generated based on pseudocode. The system compares the first and second outputs to determine a match accuracy between the two outputs. If the two outputs sufficiently match, the system approves the user prompt. If the two outputs do not sufficiently match, the system initiates a prompt restructuring process, whereby the user prompt is restructured using pseudocode to improve the accuracy of the first output. The process is repeated iteratively until the restructured user prompt generates a first output that sufficiently matches the second output generated based on the pseudocode.

Claims (79)

1 . One or more non-transitory, computer-readable storage media comprising instructions recorded thereon, wherein the instructions, when executed by at least one data processor of a system, cause the system to:

receive a user prompt indicating a request for a report summarizing data over a time period;

generate a first output by inputting, into a generative model, the user prompt to cause the generative model to generate the first output based on the user prompt, the first output comprising a first plurality of text-based analytics, a first plurality of visual analytics, and a first plurality of queries;

generate a second output by causing the system to:

retrieve, from a database associated with the data requested in the user prompt, a rule-based pseudocode;

generate a pseudocode prompt based on the rule-based pseudocode and the data requested in the user prompt; and

input, into the generative model, the pseudocode prompt to cause the generative model to generate the second output based on the pseudocode prompt, the second output comprising a second plurality of text-based analytics, a second plurality of visual analytics, and a second plurality of queries;

perform a comparison between (i) the first plurality of text-based analytics and the second plurality of text-based analytics, (ii) the first plurality of visual analytics and the second plurality of visual analytics, and (iii) the first plurality of queries and the second plurality of queries, to determine a match accuracy between the first output and the second output;

determine whether the match accuracy satisfies an accuracy threshold;

based on the match accuracy failing to satisfy the accuracy threshold, generate a restructured prompt based on the user prompt and the rule-based pseudocode; and

approve the restructured prompt for use in conjunction with the generative model.

2 . The one or more non-transitory, computer-readable storage media of claim 1 , wherein, to approve the restructured prompt, the instructions further cause the system to:

generate an updated first output by inputting, into the generative model, the restructured prompt to cause the generative model to generate the updated first output based on the restructured prompt, the updated first output comprising an updated first plurality of text-based analytics, an updated first plurality of visual analytics, and an updated first plurality of queries;

perform an updated comparison of the updated first output and the second output to determine an updated match accuracy;

determine whether the updated match accuracy satisfies the accuracy threshold; and

based on the updated match accuracy satisfying the accuracy threshold, approve the restructured prompt.

3 . The one or more non-transitory, computer-readable storage media of claim 2 , wherein performing the updated comparison comprises causing the system to compare (i) the updated first plurality of text-based analytics with the second plurality of text-based analytics, (ii) the updated first plurality of visual analytics with the second plurality of visual analytics, and (iii) the updated first plurality of queries with the second plurality of queries.

4 . The one or more non-transitory, computer-readable storage media of claim 1 , wherein performing the comparison comprises causing the system to determine a measure of similarity between (i) the first plurality of text-based analytics and the second plurality of text-based analytics, (ii) the first plurality of visual analytics and the second plurality of visual analytics, and (iii) the first plurality of queries and the second plurality of queries.

5 . The one or more non-transitory, computer-readable storage media of claim 1 , wherein the instructions further cause the system to:

determine a ranking of categories of the first output and the second output according to a measure of importance of each category, the categories comprising text-based analytics, visual analytics, and queries,

wherein the comparison is performed according to the ranking.

6 . The one or more non-transitory, computer-readable storage media of claim 5 , wherein, to perform the comparison according to the ranking, the instructions further cause the system to:

compare a higher-ranked category between the first output and the second output, wherein the categories of the first output and the second output comprise at least a higher-ranked category and a lower-ranked category;

based on determining that the higher-ranked category fails to match between the first output and the second output, generate the restructured prompt; and

based on determining that the higher-ranked category matches between the first output and the second output, approve the user prompt.

7 . A method comprising:

receiving a user prompt indicating a request for data over a time period;

generating a first output by inputting, into a generative model, the user prompt to cause the generative model to generate the first output based on the user prompt, the first output comprising a first plurality of text-based analytics, a first plurality of visual analytics, and a first plurality of queries;

generating a second output by:

retrieving, from a database associated with the data requested in the user prompt, a structured pseudocode;

generating a pseudocode prompt based on the structured pseudocode and the data requested in the user prompt; and

inputting, into the generative model, the pseudocode prompt to cause the generative model to generate the second output based on the pseudocode prompt, the second output comprising a second plurality of text-based analytics, a second plurality of visual analytics, and a second plurality of queries;

performing a comparison of the first output and the second output; and

based on the comparison indicating that the first output does not match the second output, generating a restructured prompt based on the user prompt and the structured pseudocode.

8 . The method of claim 7 , wherein performing the comparison further comprises:

determining, based on the comparison, a match accuracy between the first output and the second output; and

determining whether the match accuracy satisfies an accuracy threshold,

wherein the restructured prompt is generated based on the comparison indicating that the match accuracy does not satisfy the accuracy threshold.

9 . The method of claim 8 , further comprising:

generating an updated first output by inputting, into the generative model, the restructured prompt to cause the generative model to generate the updated first output based on the restructured prompt, the updated first output comprising an updated first plurality of text-based analytics, an updated first plurality of visual analytics, and an updated first plurality of queries;

performing an updated comparison of the updated first output and the second output to determine an updated match accuracy;

determining whether the updated match accuracy satisfies the accuracy threshold; and

based on the updated match accuracy satisfying the accuracy threshold, approving the restructured prompt.

10 . The method of claim 9 , wherein performing the updated comparison comprises comparing (i) the updated first plurality of text-based analytics with the second plurality of text-based analytics, (ii) the updated first plurality of visual analytics with the second plurality of visual analytics, and (iii) the updated first plurality of queries with the second plurality of queries.

11 . The method of claim 7 , wherein performing the comparison comprises comparing (i) the first plurality of text-based analytics with the second plurality of text-based analytics, (ii) the first plurality of visual analytics with the second plurality of visual analytics, and (iii) the first plurality of queries with the second plurality of queries.

12 . The method of claim 7 , further comprising:

determining a ranking of categories of the first output and the second output according to a measure of importance of each category, the categories comprising text-based analytics, visual analytics, and queries,

wherein the comparison is performed according to the ranking.

13 . The method of claim 12 , wherein performing the comparison according to the ranking comprises:

comparing a higher-ranked category between the first output and the second output, wherein the categories of the first output and the second output comprise at least a higher-ranked category and a lower-ranked category;

based on determining that the higher-ranked category fails to match between the first output and the second output, generating the restructured prompt; and

based on determining that the higher-ranked category matches between the first output and the second output, approving the user prompt.

14 . A system comprising:

a storage device; and

one or more processors communicatively coupled to the storage device storing instructions thereon that cause the one or more processors to:

receive a prompt indicating a request for data over a time period;

generate a first output by inputting, into a machine learning model, the prompt to cause the machine learning model to generate the first output based on the prompt;

generate a second output by causing the one or more processors to:

generate a pseudocode prompt based on a structured pseudocode and the data requested in the prompt; and

input, into the machine learning model, the pseudocode prompt to cause the machine learning model to generate the second output based on the pseudocode prompt;

perform a comparison of the first output and the second output; and

based on the comparison, generate a restructured prompt based on the prompt and the structured pseudocode.

15 . The system of claim 14 , wherein, to perform the comparison, the instructions further cause the one or more processors to:

determine, based on the comparison, a match accuracy between the first output and the second output; and

determine whether the match accuracy satisfies an accuracy threshold, wherein the restructured prompt is generated based on the comparison indicating that the match accuracy does not satisfy the accuracy threshold.

16 . The system of claim 15 , wherein the instructions further cause the one or more processors to:

generate an updated first output by inputting, into the machine learning model, the restructured prompt to cause the machine learning model to generate the updated first output based on the restructured prompt, the updated first output comprising an updated first plurality of text-based analytics, an updated first plurality of visual analytics, and an updated first plurality of queries;

perform an updated comparison of the updated first output and the second output to determine an updated match accuracy;

determine whether the updated match accuracy satisfies the accuracy threshold; and

based on the updated match accuracy satisfying the accuracy threshold, approve the restructured prompt.

17 . The system of claim 16 , wherein the second output comprises a second plurality of text-based analytics, a second plurality of visual analytics, and a second plurality of queries, and wherein performing the updated comparison comprises causing the one or more processors to compare (i) the updated first plurality of text-based analytics with the second plurality of text-based analytics, (ii) the updated first plurality of visual analytics with the second plurality of visual analytics, and (iii) the updated first plurality of queries with the second plurality of queries.

18 . The system of claim 14 , wherein the first output comprises a first plurality of text-based analytics, a first plurality of visual analytics, and a first plurality of queries, wherein the second output comprises a second plurality of text-based analytics, a second plurality of visual analytics, and a second plurality of queries, and wherein performing the comparison comprises causing the one or more processors to compare (i) the first plurality of text-based analytics with the second plurality of text-based analytics, (ii) the first plurality of visual analytics with the second plurality of visual analytics, and (iii) the first plurality of queries with the second plurality of queries.

19 . The system of claim 14 , wherein the instructions further cause the one or more processors to:

determine a ranking of categories of the first output and the second output according to a measure of importance of each category, the categories comprising text-based analytics, visual analytics, and queries,

wherein the comparison is performed according to the ranking.

20 . The system of claim 19 , wherein, to perform the comparison according to the ranking, the instructions further cause the one or more processors to:

compare a higher-ranked category between the first output and the second output, wherein the categories of the first output and the second output comprise at least a higher-ranked category and a lower-ranked category;

based on determining that the higher-ranked category fails to match between the first output and the second output, generate the restructured prompt; and

based on determining that the higher-ranked category matches between the first output and the second output, approve the prompt.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2025
From: LEWANDOWSKI, DANIEL
To: CITIBANK, N.A.
Reel/Frame 072509/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2025
From: JEYASHEKAR, NIGIL SATISH; ENGELBRECHT, JASON; WANG, ZHEYU; JIN, HAOLIN; JAIN, PAYAL; MAONAH, TARIQ HUSAYN; SATERNUS, MARIUSZ; RATH, BIRAJ KRUSHNA; MURRAY, STUART; DAVIES, PHILIP; DEB, SOURABH
To: CITIBANK, N.A.
Reel/Frame 072204/0209 →
Continuity (10)
Continuation In Part 18951120 · Nov 18, 2024
Continuation 18633293 · Apr 11, 2024
Continuation In Part 18907414 · Oct 4, 2024
Continuation 18661532 · May 10, 2024
Continuation In Part 18661519 · May 10, 2024
Continuation In Part 18633293 · Apr 11, 2024
Continuation In Part 18954389 · Nov 20, 2024
Continuation 18812913 · Aug 22, 2024
Continuation In Part 18661532 · May 10, 2024
Related Publication 20250322047A1 · Oct 16, 2025
References Cited (147)
US 8380817B2 · Okada · 2013 [cited by applicant]
US 8387020B1 · Maclachlan et al. · 2013 [cited by applicant]
US 9842045B2 · Heorhiadi · 2017 [cited by examiner]
US 10620988B2 · Lauderdale et al. · 2020 [cited by applicant]
US 10764150B1 · Hermoni et al. · 2020 [cited by applicant]
US 10949337B1 · Yalla et al. · 2021 [cited by applicant]
US 10951485B1 · Hermoni et al. · 2021 [cited by applicant]
US 11133942B1 · Griffin · 2021 [cited by applicant]
US 11153177B1 · Hermoni et al. · 2021 [cited by applicant]
US 11271822B1 · Hermoni et al. · 2022 [cited by applicant]
US 11410136B2 · Cook et al. · 2022 [cited by applicant]
US 11573848B2 · Linck et al. · 2023 [cited by applicant]
US 11652839B1 · Aloisio et al. · 2023 [cited by applicant]
US 11656852B2 · Mazurskiy · 2023 [cited by applicant]
US 11681811B1 · Dixit · 2023 [cited by applicant]
US 11683333B1 · Dominessy et al. · 2023 [cited by applicant]
US 11706241B1 · Cross et al. · 2023 [cited by applicant]
US 11720686B1 · Cross et al. · 2023 [cited by applicant]
US 11734418B1 · Epstein · 2023 [cited by applicant]
US 11750717B2 · Walsh et al. · 2023 [cited by applicant]
US 11875123B1 · Ben David · 2024 [cited by examiner]
US 11875130B1 · Bosnjakovic · 2024 [cited by examiner]
US 11924027B1 · Mysore · 2024 [cited by examiner]
US 11947435B2 · Boulineau et al. · 2024 [cited by applicant]
US 11960515B1 · Pallakonda · 2024 [cited by examiner]
US 11983806B1 · Ramesh et al. · 2024 [cited by applicant]
US 11990139B1 · Sandrew · 2024 [cited by applicant]
US 11995412B1 · Mishra · 2024 [cited by examiner]
US 12001463B1 · Pallakonda · 2024 [cited by examiner]
US 12026599B1 · Lewis, II · 2024 [cited by examiner]
US 20030007178A1 · Jeyachandran et al. · 2003 [cited by applicant]
US 20040098454A1 · Trapp et al. · 2004 [cited by applicant]
US 20050204348A1 · Horning et al. · 2005 [cited by applicant]
US 20060095918A1 · Hirose · 2006 [cited by applicant]
US 20070067848A1 · Gustave et al. · 2007 [cited by applicant]
US 20100275263A1 · Bennett et al. · 2010 [cited by applicant]
US 20100313189A1 · Beretta et al. · 2010 [cited by applicant]
US 20140137257A1 · Martinez et al. · 2014 [cited by applicant]
US 20140258998A1 · Adl-tabatabai et al. · 2014 [cited by applicant]
US 20170061132A1 · Hovor et al. · 2017 [cited by applicant]
US 20170262164A1 · Jain · 2017 [cited by examiner]
US 20170295197A1 · Parimi et al. · 2017 [cited by applicant]
US 20180239903A1 · Bodin et al. · 2018 [cited by applicant]
US 20180343114A1 · Ben-ari · 2018 [cited by applicant]
US 20190188706A1 · Mccurtis · 2019 [cited by applicant]
US 20190236661A1 · Hogg et al. · 2019 [cited by applicant]
US 20190286816A1 · Fu · 2019 [cited by applicant]
US 20200012493A1 · Sagy · 2020 [cited by applicant]
US 20200043164A1 · Fuchs et al. · 2020 [cited by applicant]
US 20200074470A1 · Deshpande et al. · 2020 [cited by applicant]
US 20200153855A1 · Kirti et al. · 2020 [cited by applicant]
US 20200259852A1 · Wolff et al. · 2020 [cited by applicant]
US 20200309767A1 · Loo et al. · 2020 [cited by applicant]
US 20200314191A1 · Madhavan et al. · 2020 [cited by applicant]
US 20200334326A1 · Zhang et al. · 2020 [cited by applicant]
US 20200349054A1 · Dai et al. · 2020 [cited by applicant]
US 20210012486A1 · Huang et al. · 2021 [cited by applicant]
US 20210049288A1 · Li · 2021 [cited by applicant]
US 20210133182A1 · Anderson et al. · 2021 [cited by applicant]
US 20210185094A1 · Waplington et al. · 2021 [cited by applicant]
US 20210211431A1 · Albero et al. · 2021 [cited by applicant]
US 20210264547A1 · Li · 2021 [cited by applicant]
US 20210273957A1 · Boyer et al. · 2021 [cited by applicant]
US 20210390465A1 · Werder et al. · 2021 [cited by applicant]
US 20220114251A1 · Guim Bernat et al. · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220147636A1 · Mahuli et al. · 2022 [cited by applicant]
US 20220179906A1 · Desai et al. · 2022 [cited by applicant]
US 20220198304A1 · Szczepanik et al. · 2022 [cited by applicant]
US 20220286438A1 · Burke et al. · 2022 [cited by applicant]
US 20220311681A1 · Palladino · 2022 [cited by examiner]
US 20220318654A1 · Lin et al. · 2022 [cited by applicant]
US 20220334818A1 · Mcfarland · 2022 [cited by applicant]
US 20220358023A1 · Moser et al. · 2022 [cited by applicant]
US 20220366140A1 · Saito et al. · 2022 [cited by applicant]
US 20220398149A1 · Mcfarland et al. · 2022 [cited by applicant]
US 20220414536A1 · M L et al. · 2022 [cited by applicant]
US 20220417274A1 · Madanahalli et al. · 2022 [cited by applicant]
US 20230019072A1 · Okunlola · 2023 [cited by applicant]
US 20230032686A1 · Williams et al. · 2023 [cited by applicant]
US 20230033317A1 · Lin et al. · 2023 [cited by applicant]
US 20230035321A1 · Vijayaraghavan · 2023 [cited by applicant]
US 20230039855A1 · Greene · 2023 [cited by applicant]
US 20230052608A1 · Wattiau et al. · 2023 [cited by applicant]
US 20230067128A1 · Engelberg et al. · 2023 [cited by applicant]
US 20230071264A1 · Hakala et al. · 2023 [cited by applicant]
US 20230076372A1 · Engelberg et al. · 2023 [cited by applicant]
US 20230077527A1 · Sarkar · 2023 [cited by applicant]
US 20230113621A1 · Griffin et al. · 2023 [cited by applicant]
US 20230114719A1 · Thomas et al. · 2023 [cited by applicant]
US 20230117962A1 · Kaimal et al. · 2023 [cited by applicant]
US 20230118388A1 · Crabtree et al. · 2023 [cited by applicant]
US 20230123314A1 · Crabtree et al. · 2023 [cited by applicant]
US 20230132703A1 · Marsenic et al. · 2023 [cited by applicant]
US 20230135660A1 · Chapman et al. · 2023 [cited by applicant]
US 20230164158A1 · Fellows et al. · 2023 [cited by applicant]
US 20230171282A1 · Bollinger · 2023 [cited by applicant]
US 20230177441A1 · Durvasula et al. · 2023 [cited by applicant]
US 20230177613A1 · Crabtree et al. · 2023 [cited by applicant]
US 20230205888A1 · Tyagi et al. · 2023 [cited by applicant]
US 20230205891A1 · Yellapragada et al. · 2023 [cited by applicant]
US 20230208869A1 · Bisht et al. · 2023 [cited by applicant]
US 20230208870A1 · Yellapragada et al. · 2023 [cited by applicant]
US 20230208871A1 · Yellapragada et al. · 2023 [cited by applicant]
US 20230229542A1 · Watkins et al. · 2023 [cited by applicant]
US 20230252393A1 · Orzechowski et al. · 2023 [cited by applicant]
US 20230259860A1 · Sarkar · 2023 [cited by applicant]
US 20230269272A1 · Dambrot et al. · 2023 [cited by applicant]
US 20240012734A1 · Lee et al. · 2024 [cited by applicant]
US 20240020538A1 · Socher · 2024 [cited by examiner]
US 20240095077A1 · Singh · 2024 [cited by examiner]
US 20240129345A1 · Kassam et al. · 2024 [cited by applicant]
US 20240202442A1 · Saito et al. · 2024 [cited by applicant]
CN 106502890A · 2017 [cited by applicant]
WO 2022125803A1 · 2022 [cited by applicant]
WO 2024020416A1 · 2024 [cited by applicant]
AI Risk Management Framework NIST, retrieved on Jun. 17, 2024, https://www.nist.gov/itl/ai-risk-management-framework. [cited by applicant]
Empower Your Team with a Compliance Co-Pilot, Sedric, retrieved on Sep. 25, 2024. https://www.sedric.ai/. [cited by applicant]
Independent analysis of AI language models and API providers. Artificial Analysis, retrieved on Jun. 13, 2024, https://artificialanalysis.ai/, 11 pages. [cited by applicant]
What is AI Verify?, AI Verify Foundation, Jun. 11, 2024, 3 pages, https://aiverifyfoundation.sg/. [cited by applicant]
Brown, D., et al., “The Great AI Challenge: We Test Five Top Bots on Useful, Everyday Skills,” The Wall Street Journal, published May 25, 2024. [cited by applicant]
Cranium, Adopt & Accelerate AI Safely, retrieved on Nov. 7, 2024, from https://cranium.ai/. [cited by applicant]
Dong, Y., et al., “Building Guardrails for Large Language Models,” https://ar5iv.labs.arxiv.org/html/2402.01822v1, published May 29, 2024, 20 pages. [cited by applicant]
Futurism, “Sam Altman Admits That OpenAI Doesn't Actually Understand How Its AI Works”, Jun. 11, 2024, 4 pages, https://futurism.com/sam-altman-admits-openai-understand-ai. [cited by applicant]
Generative machine learning models; IPCCOM000272835D, Aug. 17, 2023. (Year: 2023). [cited by applicant]
Guldimann, P., et al. “COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act,” arXiv:2410.07959v1 [cs.CL] Oct. 10, 2024, 38 pages. [cited by applicant]
Hu, Q., J., et al., “Routerbench: A Benchmark for Multi-LLM Routing System,” arXiv:2403.12031v2 [cs.LG] Mar. 28, 2024, 16 pages. [cited by applicant]
International Search Report and Written Opinion received in Application No. PCT/US24/47571, dated Dec. 9, 2024, 10 pages. [cited by applicant]
International Search Report and Written Opinion Received received in Application No. PCT/US23/85942, dated Feb. 15, 2024, 6 pages. [cited by applicant]
Kojima, Takeshi, et al. “Large Language Models are Zero-Shot Reasoners,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2205.11916 [cs.CL], Jan. 29, 2023, 42 pages. [cited by applicant]
Mathews, A. W., “What AI Can Do in Healthcare—and What It Should Never Do,” The Wall Street Journal, published on Aug. 21, 2024, retrieved on Sep. 5, 2024 https://www.wsj.com. [cited by applicant]
Mavrepis, P., et al., “XAI for All: Can Large Language Models Simplify Explainable AI?,” https://arxiv.org/abs/2401.13110, Jan. 23, 2024, 10 pages. [cited by applicant]
Mollick, E., “Latent Expertise: Everyone is in R&D,” One Useful Thing, published on Jun. 20, 2024, https://www.oneusefulthing.org/p/latent-expertise-everyone-is-in-r. [cited by applicant]
Nauta, M., et al., “From Anecdotal Evidence to Quantative Evaluation Methods: A Systematic Review of Evaluating Explainable AI” ACM Computing Surveys, vol. 55 No. 13s Article 295, 2023 [retrieved Jul. 3, 2024]. [cited by applicant]
Peers, M., “What California AI Bill Could Mean,” The Briefing, published and retrieved Aug. 30, 2024, 8 pages, https://www.theinformation.com/articles/what-california-ai-bill-could-mean. [cited by applicant]
Wei, Jason, et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), arXiv:2201.11903 [cs.CL], Jan. 10, 2023, 43 pages. [cited by applicant]
Zhao, H., et al., “Explainability for Large Language Models: A Survey,” https://arxiv.org/abs/2309.01029, Nov. 28, 2024, 38 pages. [cited by applicant]
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T., & Yu, P. S. (2024). Trustworthiness in Retrieval-Augmented Generation Systems: A Survey. ArXiv./abs/2409.10102. [cited by applicant]
“Singapore launches Project Moonshot”, a generative Artificial Intelligence testing toolkit to address LLM safety and security challenges, https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/press-r… [cited by applicant]
Aggarwal, Nitin , KPIs for gen AI: Why measuring your new AI is essential to its success, https://cloud.google.com/transform/kpis-for-gen-ai-why-measuring-your-new-ai-is-essential-to-its-success. [cited by applicant]
Anthrop/C , Mapping the Mind of a Large Language Model, https://www.anthropic.com/research/mapping-mind-language-model, May 21, 2024. [cited by applicant]
Claburn, Thomas , OpenAI's GPT-4 can exploit real vulnerabilities by reading security advisories, The Register, https://www.theregister.com/2024/04/17/gpt4_can_exploit_real_vulnerabilities/?utm_source=tldrai, Apr. 17, 2… [cited by applicant]
Marshall, Andrew , Threat Modeling AI/ML Systems and Dependencies, Nov. 2, 2022, 27 pages. [cited by applicant]
Roose, Kevin , “A.I. Has a Measurement Problem”, The New York Times, Apr. 15, 2024, 5 pages. [cited by applicant]
Roose, Kevin , “A.I.'s Black Boxes Just Got a Little Less Mysterious”, The New York Times, May 21, 2024, 5 pages. [cited by applicant]
Shah, Harshay , Decomposing and Editing Predictions by Modeling Model Computation, arXiv:2404.11534v1 [cs.LG] Apr. 17, 2024. [cited by applicant]
Shankar, Ram , “Failure Modes in Machine Learning”, , Nov. 2019, 14 pages. [cited by applicant]