IP Library › Granted Patent US 12,737,537
Granted Patent B2
US 12,737,537 · App. 18/585,145 · Granted Sep 15, 2026

Governance and confidence assessment of LLM

Inventors: Warren Nicholas Mante (Sachse, TX); Warren Thomas Bloomer Lucas (Middletown, CT); Sourav Mazumder (Contra Costa, CA); Andrew R. Freed (Cary, NC); Stefan A. G. Van Der Stockt (Austin, TX)
Assignee: International Business Machines Corporation
G06F40/20G06Q30/018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,537
App. No.
18/585,145
Granted
Sep 15, 2026
Kind
B2
Abstract

An approach for governing responses generated by a language learning model (LLM) model. The approach defines a reference set of inputs and output pairs for the LLM wherein the reference set of inputs and output pairs are actual inputs and reference outputs. The approach defines a set of metadata associated with the reference set of input and output pairs and assigns the metadata to each pair of the reference set of inputs and output pairs. The approach also defines a set of evaluation criteria, assigns the evaluation criteria to organizational risk framework and associates the set of metadata to the evaluation criteria.

Claims (83)

1 . A computer-implemented method for governing responses generated by a language large learning model (LLM), the computer-implemented method comprising:

setting up an initial risk assessment framework for a model of an LLM in a question repository, wherein setting up further comprises:

defining a reference set of inputs and output pairs for the LLM wherein the reference set of inputs and output pairs are actual inputs and reference outputs;

defining a set of metadata associated with the reference set of inputs and output pairs, wherein the set of metadata including at least prioritization, sensitive-content designation, and topic area;

assigning the set of metadata to each pair of the reference set of inputs and output pairs;

defining a set of evaluation criteria including

(i) an exact match criterion computed using a normalized Damerau-Levenshtein score, (ii) a semantic match criterion computed using a cosine similarity score, and

(iii) a prohibited word or phrase criterion determined by detecting presence of a specified word and/or phrase; assigning the set of evaluation criteria to an organizational risk framework;

associating the set of metadata to the set of evaluation criteria;

evaluating the initial risk assessment framework and the model of the LLM after a predetermined interval, further comprising:

passing all inputs from the reference set of inputs and output pairs to the LLM, and collecting all outputs from the reference set of inputs and output pairs at a predetermined interval;

evaluating and producing an evaluation metric of all the inputs and all of the outputs based on one or more policies, further comprising:

generating new outputs and comparing the new outputs against the reference set of inputs and output pairs; and

generating scores for the new outputs and adding the scores to a historical data for future analysis;

summarizing results of the evaluation metric across all of the reference set of inputs and output pairs; and

notifying one or more users when the results of the evaluation metric exceed a predetermined metric threshold;

responsive to determining that the evaluation metric exceeds the predetermined metric threshold, automatically initiating, in a governance risk and compliance (GRC) system, a workflow that includes the actual input, the reference output, actual output and confidence score for review; and

saving data associated with the result of the evaluation metric and all of the reference set of inputs and output pairs to a data repository for trend analysis.

2 . The computer-implemented method of claim 1 , wherein the predetermined interval comprises every hour or every week.

3 . The computer-implemented method of claim 1 , wherein the set of evaluation criteria further comprises:

the exact match criterion in which the normalized Damerau-Levenshtein score is required to equal 1.0 to pass;

ii) the semantic match criterion in which the cosine similarity score is required to exceed a threshold X to pass; and

(iii) the prohibited word or phrase criterion in which the presence of the specified word and/or phrase triggers a fail.

4 . The computer-implemented method of claim 1 , wherein assigning the set of metadata to each pair of the reference set of inputs and output pairs further comprises assessing the set of metadata using a rules engine or classifier including at least one of: keyword evaluation to mark an input/output pair as sensitive content; and a topic classifier to assign the topic area.

5 . The computer-implemented method of claim 1 , wherein the set of metadata comprises, prioritization, sensitive content and topic areas.

6 . The computer-implemented method of claim 1 , wherein automatically initiating the workflow further comprises routing the workflow to at least one user selected from an administrator, a data scientist, risk management personnel, and an internal auditor, and wherein the workflow is reviewed to determine whether the actual output is within specification under the evaluation criteria or whether the model of the LLM is to be fine-tuned.

7 . The computer-implemented method of claim 1 , wherein associating the set of metadata to the set of evaluation criteria can include associating high priority to the exact match criterion; associating low priority and low sensitivity content to the semantic match criterion; and associating sensitive content to the prohibited word or phrase criterion.

8 . A computer program product for governing responses generated by a language large learning model (LLM), the computer program product comprising:

one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising the steps of:

setting up an initial risk assessment framework for a model of an LLM in a question repository, wherein setting up further comprises:

defining a reference set of inputs and output pairs for the LLM wherein the reference set of inputs and output pairs are actual inputs and reference outputs;

defining a set of metadata associated with the reference set of inputs and output pairs, wherein the set of metadata including at least prioritization, sensitive-content designation, and topic area;

assigning the set of metadata to each pair of the reference set of inputs and output pairs;

defining a set of evaluation criteria including (i) an exact match criterion computed using a normalized Damerau-Levenshtein score, (ii) a semantic match criterion computed using a cosine similarity score, and

(iii) a prohibited word or phrase criterion determined by detecting presence of a specified word and/or phrase;

assigning the set of evaluation criteria to an organizational risk framework;

and

associating the set of metadata to the set of evaluation criteria; evaluating the initial risk assessment framework and the model of the LLM after a predetermined interval, further comprising:

passing all inputs from the reference set of inputs and output pairs to the LLM, and collecting all outputs from the reference set of inputs and output pairs at a predetermined interval;

evaluating and producing an evaluation metric of all the inputs and all of the outputs based on one or more policies, further comprising:

generating new outputs and comparing the new outputs against the reference set of inputs and output pairs; and

generating scores for the new outputs and adding the scores to a historical data for future analysis;

summarizing results of the evaluation metric across all of the reference set of inputs and output pairs; and

notifying one or more users when the results of the evaluation metric exceed a predetermined metric threshold;

responsive to determining that the evaluation metric exceeds the predetermined metric threshold, automatically initiating, in a governance risk and compliance (GRC) system, a workflow that includes the actual input, the reference output, actual output and confidence score for review; and

saving data associated with the result of the evaluation metric and all of the reference set of inputs and output pairs to a data repository for trend analysis.

9 . The computer program product of claim 8 , wherein the predetermined interval comprises every hour or every week.

10 . The computer program product of claim 8 , wherein the set of evaluation criteria further comprises (i) the exact match criterion in which the normalized Damerau-Levenshtein score is required to equal 1.0 to pass;

(ii) the semantic match criterion in which the cosine similarity score is required to exceed a threshold X to pass; and

iii) the prohibited word or phrase criterion in which the presence of the specified word and/or phrase triggers a fail.

11 . The computer program product of claim 8 , wherein assigning the set of metadata to each pair of the reference set of inputs and output pairs further comprises:

assessing the set of metadata using a rules engine or classifier including at least one of: keyword evaluation to mark an input/output pair as sensitive content; and a topic classifier to assign the topic area.

12 . The computer program product of claim 8 , wherein the set of metadata comprises, prioritization, sensitive content and topic areas.

13 . The computer program product of claim 8 , wherein automatically initiating the workflow further comprises routing the workflow to at least one user selected from an administrator, a data scientist, risk management personnel, and an internal auditor, and wherein the workflow is reviewed to determine whether the actual output is within specification under the evaluation criteria or whether the model of the LLM is to be fine-tuned.

14 . The computer program product of claim 8 , wherein associating the set of metadata to the set of evaluation criteria can include, a high priority is associated to exact match, a low priority and low sensitivity content are required a semantic match and sensitive contents must not contain a certain letter and/or phrase.

15 . A computer system for governing responses generated by a language learning model (LLM), the computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising the steps of:

setting up an initial risk assessment framework for a model of an LLM in a question repository, wherein setting up further comprises:

defining a reference set of inputs and output pairs for the LLM wherein the reference set of inputs and output pairs are actual inputs and reference outputs; defining a set of metadata associated with the reference set of inputs and output pairs, wherein the set of metadata including at least prioritization, sensitive-content designation, and topic area;

assigning the set of metadata to each pair of the reference set of inputs and output pairs;

defining a set of evaluation criteria including

(i) an exact match criterion computed using a normalized Damerau-Levenshtein score, (ii) a semantic match criterion computed using a cosine similarity score, and

(iii) a prohibited word or phrase criterion determined by detecting presence of a specified word and/or phrase; assigning the set of evaluation criteria to an organizational risk framework;

and

associating the set of metadata to the set of evaluation criteria; evaluating the initial risk assessment framework and the model of the LLM after a predetermined interval, further comprising:

passing all inputs from the reference set of inputs and output pairs to the LLM, and collecting all outputs from the reference set of inputs and output pairs at a predetermined interval;

evaluating and producing an evaluation metric of all the inputs and all of the outputs based on one or more policies, further comprising:

generating new outputs and comparing the new outputs against the reference set of inputs and output pairs; and

generating scores for the new outputs and adding the scores to a historical data for future analysis;

summarizing results of the evaluation metric across all of the reference set of inputs and output pairs; and

notifying one or more users when the results of the evaluation metric exceed a predetermined metric threshold; responsive to determining that the evaluation metric exceeds the predetermined metric threshold, automatically initiating, in a governance risk and compliance (GRC) system, a workflow that includes the actual input, the reference output, actual output and confidence score for review; and

saving data associated with the result of the evaluation metric and all of the reference set of inputs and output pairs to a data repository for trend analysis.

16 . The computer system of claim 15 , wherein the predetermined interval comprises every hour or every week.

17 . The computer system of claim 15 , wherein the set of evaluation criteria further comprises (i) the exact match criterion in which the normalized Damerau-Levenshtein score is required to equal 1.0 to pass;

(ii) the semantic match criterion in which the cosine similarity score is required to exceed a threshold X to pass; and

(iii) the prohibited word or phrase criterion in which the presence of the specified word and/or phrase triggers a fail.

18 . The computer system of claim 15 , wherein assigning the set of metadata to each pair of the reference set of inputs and output pairs further comprises:

assessing the set of metadata using a rules engine or classifier including at least one of:

keyword evaluation to mark an input/output pair as sensitive content; and a topic classifier to assign the topic area.

19 . The computer system of claim 15 , wherein the set of metadata comprises, prioritization, sensitive content and topic areas.

20 . The computer system of claim 15 , wherein automatically initiating the workflow further comprises routing the workflow to at least one user selected from an administrator, a data scientist, risk management personnel, and an internal auditor, and wherein the workflow is reviewed to determine whether the actual output is within specification under the evaluation criteria or whether the model of the LLM is to be fine-tuned.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2024
From: MANTE, WARREN NICHOLAS; LUCAS, WARREN THOMAS BLOOMER; MAZUMDER, SOURAV; FREED, ANDREW R.; VAN DER STOCKT, STEFAN A. G.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066538/0298 →
Continuity (1)
Related Publication 20250272484A1 · Aug 28, 2025
References Cited (19)
US 10592838B2 · Urban · 2020 [cited by applicant]
US 10984283B2 · Aerni · 2021 [cited by applicant]
US 11263550B2 · Vasconcelos · 2022 [cited by applicant]
US 11537875B2 · Kozhaya · 2022 [cited by applicant]
US 11568286B2 · Nourian · 2023 [cited by applicant]
US 11714842B1 · Sharma · 2023 [cited by examiner]
US 11741302B1 · Pathak · 2023 [cited by examiner]
US 11900229B1 · Swope · 2024 [cited by examiner]
US 20190279111A1 · Merrill · 2019 [cited by applicant]
US 20190340518A1 · Merrill · 2019 [cited by applicant]
US 20220108222A1 · Brannon · 2022 [cited by applicant]
US 20220114399A1 · Castiglione et al. · 2022 [cited by applicant]
US 20220400094A1 · Sampath · 2022 [cited by examiner]
US 20230281482A1 · Lantzman · 2023 [cited by examiner]
US 20230377037A1 · Kamkar · 2023 [cited by applicant]
US 20240127153A1 · Amini · 2024 [cited by examiner]
US 20250199786A1 · Silver · 2025 [cited by examiner]
CN 109947088B · 2020 [cited by applicant]
CN 116909889A · 2023 [cited by applicant]