IP Library Granted Patent US 12,645,883
Granted Patent B2
US 12,645,883 · App. 18/774,571 · Granted Jun 2, 2026

Application specific auto-evaluation for large language models (LLMs)

Inventors: Liyu Gong (Lexington, KY); Michael Avendi (Irvine, CA); Yuying Wang (Redmond, WA); Tao Sheng (Bellevue, WA); Jun Qian (Bellevue, WA); Vinod Mamtani (Bellevue, WA)
Assignee: Oracle International Corporation
G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,883
App. No.
18/774,571
Granted
Jun 2, 2026
Kind
B2
Abstract

In one embodiment, a non-transitory computer-readable media stores instructions executable by processors for generating a prompt configured for eliciting outputs from large language models (LLMs) based on information associated with a task, inputting the prompt to a first LLM configured to output a response based on processing the prompt, determining metrics for evaluating the first LLM based on the task, wherein each of the metrics is associated with a scoring guideline, generating metric prompts based on the respective metrics and the scoring guidelines associated with the respective metrics, inputting the response and the metric prompts to second LLMs configured to output scores corresponding to the respective metrics based on processing the response and the metric prompts, and generating an analysis report based on the metrics and their corresponding scores.

Claims (64)

1 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause performance of:

generating, based on information associated with a first task, a prompt configured for eliciting outputs from large language models (LLMs);

inputting the prompt to a first LLM, wherein the first LLM is configured to output a response based on processing the prompt;

determining, based on the first task, one or more metrics for evaluating the first LLM, wherein each of the metrics is associated with a scoring guideline;

generating one or more metric prompts based on the one or more respective metrics and the one or more scoring guidelines associated with the one or more respective metrics;

inputting the response and the one or more metric prompts to one or more second LLMs, wherein the one or more second LLMs are configured to output one or more scores corresponding to the one or more respective metrics based on processing the response and the metric prompts; and

generating an analysis report based on the one or more metrics and their corresponding scores.

2 . The media of claim 1 , wherein the instructions when executed by the processors, cause further performance of:

generating one or more examples for each of the one or more metrics; and

injecting the examples to each of the metric prompts associated with the respective metric.

3 . The media of claim 2 , wherein each of the examples comprises one or more of:

a demo;

a reference score for the demo; or

an explanation for the reference score.

4 . The media of claim 1 , wherein the instructions when executed by the processors, cause further performance of:

determining one or more thresholds for the one or more metrics, respectively;

comparing the one or more scores with the one or more thresholds, respectively; and

determining whether the first LLM passes or fails the evaluation based on the comparison.

5 . The media of claim 4 , wherein the instructions when executed by the processors, cause further performance of:

determining all the metrics exceed their respective thresholds;

based at least in part on said determining all the metrics exceed their respective thresholds, determining the first LLM passes the evaluation; and

wherein the analysis report comprises an indication that the first LLM passes the evaluation.

6 . The media of claim 4 , wherein the instructions when executed by the processors, cause further performance of:

determining one or more of the metrics are lower than their corresponding thresholds;

based at least in part on said determining one or more of the metrics are lower than their corresponding thresholds, determining the first LLM fails the evaluation; and

wherein the analysis report comprises an indication that the first LLM fails the evaluation.

7 . The media of claim 4 , wherein the one or more thresholds are determined based on one or more requirements associated with the first task.

8 . The media of claim 4 , wherein the instructions when executed by the processors, cause further performance of:

inputting the prompt to the first LLM for a plurality of respective times, wherein the first LLM is configured to output a plurality of respective responses based on processing the prompt for the plurality of respective times;

determining, by the one or more second LLMs, a plurality of sets of one or more scores corresponding to the one or more respective metrics based on processing the plurality of response and the metric prompts; and

calculating a pass-rate for the first LLM based on determining whether the first LLM passes or fails the evaluation for each of the plurality of times;

wherein the analysis report comprises the pass-rate.

9 . The media of claim 1 , wherein the instructions when executed by the processors, cause further performance of:

generating, based on one or more of the metric prompts, one or more metric functions configured to be re-used for one or more second tasks.

10 . The media of claim 1 , wherein generating the prompt is based on a prompt template and one or more input variables associated with the first task.

11 . The media of claim 10 , wherein each of the input variables comprises a pair of variable name and value.

12 . The media of claim 1 , wherein the metrics comprise one or more of tone, natural language quality, variety, repetitiveness, content classification, content relevance, JSON format, or length.

13 . The media of claim 1 , wherein one or more of the metrics are shareable between the first task and one or more second tasks.

14 . The media of claim 1 , wherein the one or more second LLMs are based on a same model.

15 . The media of claim 1 , wherein one or more of the second LLMs are based on different models.

16 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions, when executed using the processors, cause the system to execute:

generating, based on information associated with a first task, a prompt configured for eliciting outputs from large language models (LLMs);

inputting the prompt to a first LLM, wherein the first LLM is configured to output a response based on processing the prompt;

determining, based on the first task, one or more metrics for evaluating the first LLM, wherein each of the metrics is associated with a scoring guideline;

generating one or more metric prompts based on the one or more respective metrics and the one or more scoring guidelines associated with the one or more respective metrics;

inputting the response and the one or more metric prompts to one or more second LLMs, wherein the one or more second LLMs are configured to output one or more scores corresponding to the one or more respective metrics based on processing the response and the metric prompts; and

generating an analysis report based on the one or more metrics and their corresponding scores.

17 . The system of claim 16 , wherein the instructions when executed using the processors, cause the processors to further execute:

generating one or more examples for each of the one or more metrics; and

injecting the examples to each of the metric prompts associated with the respective metric.

18 . The system of claim 16 , wherein the instructions when executed using the processors, cause the processors to further execute:

determining one or more thresholds for the one or more metrics, respectively;

comparing the one or more scores with the one or more thresholds, respectively; and

determining whether the first LLM passes or fails the evaluation based on the comparison.

19 . A method comprising, by one or more computing systems:

generating, based on information associated with a first task, a prompt configured for eliciting outputs from large language models (LLMs);

inputting the prompt to a first LLM, wherein the first LLM is configured to output a response based on processing the prompt;

determining, based on the first task, one or more metrics for evaluating the first LLM, wherein each of the metrics is associated with a scoring guideline;

generating one or more metric prompts based on the one or more respective metrics and the one or more scoring guidelines associated with the one or more respective metrics;

inputting the response and the one or more metric prompts to one or more second LLMs, wherein the one or more second LLMs are configured to output one or more scores corresponding to the one or more respective metrics based on processing the response and the metric prompts; and

generating an analysis report based on the one or more metrics and their corresponding scores.

20 . The method of claim 19 , further comprising:

generating one or more examples for each of the one or more metrics; and

injecting the examples to each of the metric prompts associated with the respective metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2024
From: GONG, LIYU; AVENDI, MICHAEL; WANG, YUYING; SHENG, TAO; QIAN, JUN; MAMTANI, VINOD
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 068002/0256 →
Continuity (1)
Related Publication 20260023929A1 · Jan 22, 2026
References Cited (28)
US 11875130B1 · Bosnjakovic · 2024 [cited by examiner]
US 12515075B2 · Kuusela · 2026 [cited by examiner]
US 20220201021A1 · Joshi · 2022 [cited by examiner]
US 20230316090A1 · Chakraborty · 2023 [cited by examiner]
US 20240362422A1 · Callegari · 2024 [cited by examiner]
US 20240428015A1 · Yoon · 2024 [cited by examiner]
US 20250077844A1 · Lin · 2025 [cited by examiner]
US 20250086211A1 · Bolcer · 2025 [cited by examiner]
US 20250111147A1 · Pryzant · 2025 [cited by examiner]
US 20250111167A1 · Mcintyre · 2025 [cited by examiner]
US 20250260707A1 · Belgi · 2025 [cited by examiner]
US 20250291934A1 · Chan · 2025 [cited by examiner]
US 20250335858A1 · Stavarache · 2025 [cited by examiner]
CN 117171536A · 2023 [cited by applicant]
CN 117194258A · 2023 [cited by applicant]
CN 117573846A · 2024 [cited by applicant]
CN 117667635A · 2024 [cited by applicant]
WO WO2024177227A1 · 2024 [cited by examiner]
WO WO2025017427A1 · 2025 [cited by examiner]
Yongqiang Ma; et al.; From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications; Apr. 11, 2024. [cited by applicant]
Arthur Team; LLM-Guided Evaluation: Using LLMs to Evaluate LLMs; https://www.arthur.ai/blog/llm-guided-evaluation-using-llms-to-evaluate-llms; Sep. 29, 2023. [cited by applicant]
Yang Liu, et al.; G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment; May 23, 2023. [cited by applicant]
Qingqing Zhu et al.; How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses; Jun. 18, 2024. [cited by applicant]
Jeffrey Zhou et al.; Instruction-Following Evaluation for Large Language Models; Nov. 14, 2023. [cited by applicant]
Yen-Ting Lin, et al.; National Taiwan University, Taipei, Taiwan; “LLM-EVAL: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models;” May 23, 2023. [cited by applicant]
Zingyao Wang, et al.; “MINT: Evaluating LLMs in Multi-Turn Interaction with Tools and Language Feedback;” Published as a conference paper at ICLR 2024. [cited by applicant]
Patentability Search Report; Ref. No: IDF-138335; Apr. 12, 2024. [cited by applicant]
Arthur Team; “The Most Robust Way to Evaluate LLMs;” https://www.arthur.ai/product/bench; Arthur Bench; Printed Feb. 3, 2025. [cited by applicant]