IP Library › Granted Patent US 12,572,856
Granted Patent B1
US 12,572,856 · App. 19/202,544 · Granted Mar 10, 2026

System and method for evaluating generative artificial intelligence outcomes

Inventors: Anshuman Behera (Kearny, NJ); Raka Rajanigandha (Jersey City, NJ); Vidhan Janak Dhagai (Mumbai, IN); Mazeed Ahamed Mohammed (Anantapur, IN); Sujit Eapen (Plainsboro, NJ)
Assignee: Morgan Stanley Services Group Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,856
App. No.
19/202,544
Granted
Mar 10, 2026
Kind
B1
Abstract

A computer-implemented method of assessing a large language model (LLM) includes receiving user inputs concerning the LLM including selected hyperparameters, a use case, at least one prompt, and examples. The user inputs are mapped with a glossary of metrics to determine recommended metrics and recommended prompts for the LLM. A minimum recommended sample size is determined based on the user inputs, recommended at least one metric and expected confidence and accuracy. An LLM-generated dataset related to the use case is augmented when it is determined that the LLM-generated dataset has fewer entries than the minimum recommended sample size. An evaluation report is then generated for assessing the recommended at least one metric for determining the accuracy of: i) the LLM based on the user inputs, recommended at least one metric and the LLM-generated dataset, ii) the at least one user prompt, and iii) at least one recommended prompt.

Claims (28)

1 . A computer-implemented method of assessing a large language model (LLM) comprising:

receiving user inputs concerning the LLM including selected hyperparameters, a use case, at least one prompt, and user provided examples;

mapping the user inputs with a glossary of LLM metrics to determine at least one recommended metric and at least one recommended prompt for the LLM;

determining a minimum recommended sample size based on the user inputs, recommended at least one metric and expected confidence and accuracy;

augmenting an LLM-generated dataset related to the use case when it is determined that the user inputs and LLM-generated dataset has fewer entries than the minimum recommended sample size; and

generating an evaluation report for assessing the recommended at least one metric for determining the accuracy of: i) the LLM based on the user inputs, recommended at least one metric and the LLM-generated dataset, ii) the at least one user prompt, and iii) at least one recommended prompt.

2 . The computer-implemented method of claim 1 , further comprising:

receiving the augmented LLM-generated dataset, recommended at least one metric, recommended prompts and user selected metrics;

generating a final objective metric ranking, final subjective metric ranking, and final safety metric ranking using an LLM based on the received augmented LLM-generated dataset, recommended at least one metric, recommended prompts and user selected metrics; and

generating an aggregated prompt ranking that includes the final objective metric ranking, final subjective metric ranking, and final safety metric ranking.

3 . The computer-implemented method of claim 2 , wherein the final objective metric ranking, final subjective metric ranking, and final safety metric ranking are each determined by extracting metrics related to the use case using an additional, second LLM and executing a structured prompt using the first LLM which includes instructions for evaluating content according to a respective criterion.

4 . The computer-implemented method of claim 2 , further comprising:

receiving the final objective metric ranking, final subjective metric ranking, and final safety metric ranking; and

generating a prompt by dynamically aligning with the preferences of the final objective metric ranking, final subjective metric ranking, and final safety metric rating using an LLM to balance evaluations across categories.

5 . The computer-implemented method of claim 1 , wherein the at least one recommended metric and at least one recommended prompt for the LLM determined by the mapping includes a selected number of objective metrics, a selected number of subjective metrics and a selected number of safety metrics.

6 . The computer-implemented method of claim 1 , wherein the mapping includes the steps of:

comparing an embedded model of the user inputs concerning the LLM and an embedded topic model of the glossary of LLM metrics using cosine textual similarity to obtain a cosine similarity ranking list; and

determining a topic similarity between the embedded model of the user input concerning the LLM and the embedded topic model of the glossary of LLM metrics to obtain a topic similarity ranking list; and

inputting to user inputs, the cosine similarity ranking list and topic similarity ranking list to an LLM to determine an LLM-based similarity ranking list.

7 . The computer-implemented method of claim 6 , further comprising aggregating the cosine similarity ranking list, the topic similarity ranking list and the LLM-based similarity ranking list with selected weights for each respective list to determine ranked lists of recommended objective metrics, subjective metrics and safety metrics.

8 . The computer-implemented method of claim 6 , further comprising synthesizing dynamic prompts based on the at least one prompt provided in the user input, the dynamic prompts including a plurality of variations that align with different approaches.

9 . The computer-implemented method of claim 6 , wherein the dynamic prompts include a directive prompt that provides explicit instructions to the LLM; a Scenario-Based prompt that situates the LLM within a hypothetical but plausible scenario, and an Expertise-Affirming prompt that underscores a role of the LLM as an expert with a specific domain.

10 . The computer-implemented method of claim 6 , wherein the steps of augmenting an LLM-generated dataset includes:

comparing a vectorized version of the user provided examples with vectorized examples from a global dataset;

selecting a group of closest samples from the user provided examples and vectorized examples based on the comparison; and

performing cluster analysis, feature engineering and dimensionality reduction on the group of closest samples, yielding a modified example set; and

generating additional samples based on the modified example set using an LLM that employs and adversarial text generation process; and

generating additional samples based on the modified example set using an LLM that employs and cooperative text generation process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2025
From: BEHERA, ANSHUMAN; RAJANIGANDHA, RAKA; DHAGAI, VIDHAN JANAK; MOHAMMED, MAZEED AHAMED; EAPEN, SUJIT
To: MORGAN STANLEY SERVICES GROUP INC.
Reel/Frame 071065/0210 →
References Cited (12)
US 12147513B1 · Jain et al. · 2024 [cited by applicant]
US 20240289395A1 · Zhou et al. · 2024 [cited by applicant]
US 20240289558A1 · Muraoka et al. · 2024 [cited by applicant]
US 20240296314A1 · Singh et al. · 2024 [cited by applicant]
US 20240296315A1 · Singh et al. · 2024 [cited by applicant]
US 20240311618A1 · Sukhavasi et al. · 2024 [cited by applicant]
US 20240330655A1 · Hearty et al. · 2024 [cited by applicant]
US 20240362417A1 · Imani et al. · 2024 [cited by applicant]
US 20240378396A1 · Bhupati et al. · 2024 [cited by applicant]
Chang, Yupeng, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen et al. “A survey on evaluation of large language models.” ACM transactions on intelligent systems and technology 15, No. 3 (2024): 1-45. (Y… [cited by examiner]
Banerjee, Debarag, Pooja Singh, Arjun Avadhanam, and Saksham Srivastava. “Benchmarking LLM powered chatbots: methods and metrics.” arXiv preprint arXiv:2308.04624 (2023). (Year: 2023). [cited by examiner]
Liu, Zhiwei, Weiran Yao, Jianguo Zhang, Le Xue, Shelby Heinecke, Rithesh Murthy, Yihao Feng et al. “Bolaa: Benchmarking and orchestrating IIm-augmented autonomous agents.” arXiv preprint arXiv:2308.05960 (2023). (Year: … [cited by examiner]