IP Library › Granted Patent US 12,566,957
Granted Patent B1
US 12,566,957 · App. 19/172,987 · Granted Mar 3, 2026

System and method for automatic evaluation of artificial intelligence models

Inventors: Sreekanth Menon (Bangalore, IN); Megha Sinha (Gurgaon, IN); Ram Sagar (Bangalore, IN); Vernika Samadhiya (Delhi, IN)
Assignee: Genpact USA, Inc.
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,957
App. No.
19/172,987
Filed
Apr 8, 2025
Granted
Mar 3, 2026
Kind
B1
Art Unit
2129
USPC
706/15
Abstract

A method for automatic evaluation of an artificial intelligence (AI) model is presented. The method can include receiving, at a computer system comprising a processor and a memory storing instructions executable by the processor, a pretrained AI model, an input dataset, and an expected predictions dataset, wherein the expected predictions dataset comprises expected results of the pretrained AI model based on the input dataset. The method can include generating an actual predictions dataset using the pretrained AI model by providing, by the processor, the input dataset as input to the pretrained AI model and receiving, at the computer system, as output from the pretrained AI model a plurality of predictions based on the input dataset. The method can include calculating, by the processor, a plurality of algorithmic metrics based on the expected predictions dataset and the actual predictions dataset. The method can include determining, by the processor, a responsible AI (RAI) score based on the plurality of algorithmic metrics. The method can include calculating, by the processor, a RAI health score based on the RAI score and a safety fraction. The method can include comparing, by the processor, the RAI health score to a RAI health metric. The method can include determining, by the processor, an overall RAI health of the pretrained AI model based on the comparison.

Claims (170)

1 . A method for automatic evaluation of an artificial intelligence (AI) model, the method comprising:

receiving, at a computer system comprising a processor and a memory storing instructions executable by the processor, a pretrained AI model, an input dataset, and an expected predictions dataset, wherein the expected predictions dataset comprises expected results of the pretrained AI model based on the input dataset;

generating an actual predictions dataset using the pretrained AI model by providing, by the processor, the input dataset as input to the pretrained AI model and receiving, at the computer system, as output from the pretrained AI model a plurality of predictions based on the input dataset;

calculating, by the processor, a plurality of algorithmic metrics based on the expected predictions dataset and the actual predictions dataset;

calculating, by the processor, a responsible AI (RAI) score based on;

calculating, by the processor, a safety fraction based on the equation:

[

1

-

[

(

∑

c

⁢

(

∑

i

∈

Q

⁢

q

j

)

×

W

C

)

+

P

f

∑

max

⁢

(

∑

j

∈

Q

q

j

)

×

W

C

]

]

wherein q j is a category score, W c is a weightage constant, and P f is a penalty factor;

calculating, by the processor, a RAI health score based on the RAI score and the safety fraction;

comparing, by the processor, the RAI health score to a RAI health metric;

determining, by the processor, an overall RAI health of the pretrained AI model based on the comparison; and

based on the RAI health score, sending a signal to cause an enterprise network to prevent deploying or disable the pretrained AI model on the enterprise network.

2 . The method of claim 1 , wherein the pretrained AI model comprises at least one of predictive AI models or generative AI models.

3 . The method of claim 1 , wherein the input dataset comprises a training dataset.

4 . The method of claim 1 , wherein the expected predictions dataset comprises output data generated from a training AI model.

5 . The method of claim 1 , wherein receiving the pretrained AI model comprises receiving a regression-based AI model.

6 . The method of claim 5 , wherein the plurality of algorithmic metrics comprises at least one of an adjusted R-squared score, a mean absolute percentage error (MAPE), a median absolute error (MAE), a D-squared score, an explained variance score, a feature importance distribution, an interpretability score, a local interpretable model-agnostic explanations (LIME) score, a Cramér-von Mises (CvM) statistic, or a Regressor Uncertainty score.

7 . The method of claim 1 , wherein calculating the RAI health score comprises calculating the RAI health score based on the equation:

W

r

[

100

×

S

L

]

+

W

a

(

∑

(

m

k

×

W

m

)

×

100

)

wherein W r is a performance constant, W a is an algorithmic constant, m k is an individual algorithmic metric of the plurality of algorithmic metrics, and W m is a weightage corresponding to the individual algorithmic metric, and S L is the safety fraction.

8 . The method of claim 1 , wherein calculating the plurality of algorithmic metrics comprises calculating at least one corresponding evaluation metric for each one of the plurality of algorithmic metrics.

9 . The method of claim 1 , wherein calculating the plurality of algorithmic metrics comprises calculating an average metric from the plurality of algorithmic metrics.

10 . The method of claim 1 , further comprising:

based on the RAI health score, sending a signal to cause an enterprise network to prevent deploying the pretrained AI model on the enterprise network.

11 . The method of claim 1 , further comprising

based on the RAI health score, sending a signal to cause an enterprise network to disable the pretrained AI on the enterprise network.

12 . The method of claim 1 , wherein calculating the RAI score based on the plurality of algorithmic metrics comprises calculating the RAI score based on the equation: Σ m,k=1 m k W m wherein m k is an individual algorithmic metric of the plurality of algorithmic metrics, and W m is a weightage corresponding to the individual algorithmic metric.

13 . A method for automatic evaluation of an artificial intelligence (AI) model, the method comprising:

receiving, at a computer system comprising a processor and a memory storing instructions executable by the processor, a pretrained AI model, an input dataset, and an expected predictions dataset, wherein the expected predictions dataset comprises expected results of the pretrained AI model based on the input dataset;

generating an actual predictions dataset using the pretrained AI model by providing, by the processor, the input dataset as input to the pretrained AI model and receiving, at the computer system, as output from the pretrained AI model a plurality of predictions based on the input dataset;

calculating, by the processor, a plurality of algorithmic metrics based on the expected predictions dataset and the actual predictions dataset;

calculating, by the processor, a responsible AI (RAI) score based on the plurality of algorithmic metrics;

calculating, by the processor, a safety fraction based on the equation:

[

1

-

[

(

∑

c

⁢

(

∑

i

∈

Q

⁢

q

j

)

×

W

C

)

+

P

f

∑

max

⁢

(

∑

j

∈

Q

q

j

)

×

W

C

]

]

wherein q j is a category score, W c is a weightage constant, and P f is a penalty factor;

calculating, by the processor, a RAI health score based on the RAI score and the safety fraction;

comparing, by the processor, the RAI health score to a RAI health metric; and

based on the RAI health score, sending a signal to cause an enterprise network to prevent deploying or disable the pretrained AI model on the enterprise network.

14 . The method of claim 13 , further comprising determining, by the processor, an overall RAI health of the pretrained AI model based on the comparison.

15 . The method of claim 13 , wherein the pretrained AI model comprises at least one of predictive AI models or generative AI models.

16 . The method of claim 13 , wherein the input dataset comprises a training dataset.

17 . The method of claim 13 , wherein the expected predictions dataset comprises output data generated from a training AI model.

18 . The method of claim 13 , wherein receiving the pretrained AI model comprises receiving a regression-based AI model.

19 . The method of claim 18 , wherein the plurality of algorithmic metrics comprises at least one of an adjusted R-squared score, a mean absolute percentage error (MAPE), a median absolute error (MAE), a D-squared score, an explained variance score, a feature importance distribution, an interpretability score, a local interpretable model-agnostic explanations (LIME) score, a Cramér-von Mises (CvM) statistic, or a Regressor Uncertainty score.

20 . The method of claim 13 , wherein calculating the RAI health score comprises calculating the RAI health score based on the equation:

W

r

[

100

×

S

L

]

+

W

a

(

∑

(

m

k

×

W

m

)

×

100

)

wherein W r is a performance constant, W a is an algorithmic constant, m k is an individual algorithmic metric of the plurality of algorithmic metrics, W m is a weightage corresponding to the individual algorithmic metric, and S L is the safety fraction.

21 . The method of claim 13 , wherein calculating the RAI score based on the plurality of algorithmic metrics comprises calculating the RAI score based on the equation: Σ m,k=1 m k W m wherein m k is an individual algorithmic metric of the plurality of algorithmic metrics, and W m is a weightage corresponding to the individual algorithmic metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2025
From: MENON, SREEKANTH; SINHA, MEGHA; SAGAR, RAM; SAMADHIYA, VERNIKA
To: GENPACT USA, INC.
Reel/Frame 071008/0340 →
References Cited (13)
US 10769570B2 · Lu · 2020 [cited by applicant]
US 11399060B2 · Morin · 2022 [cited by applicant]
US 11915179B2 · Lee et al. · 2024 [cited by applicant]
US 11954112B2 · Siebel et al. · 2024 [cited by applicant]
US 20170270408A1 · Shi · 2017 [cited by examiner]
US 20220300822A1 · Shmelkin · 2022 [cited by examiner]
US 20230368868A1 · Griffin · 2023 [cited by examiner]
US 20240202351A1 · Mohammed · 2024 [cited by examiner]
US 20240202751A1 · Coleman · 2024 [cited by examiner]
US 20240256964A1 · Tay · 2024 [cited by examiner]
US 20250068982A1 · Huang · 2025 [cited by examiner]
US 20250078091A1 · Munguia Tapia · 2025 [cited by examiner]
⋅ Kereopa-Yorke et al. (“Quantifying AI Vulnerabilities: A Synthesis of Complexity, Dynamical Systems, and Game Theory”, NPL, arXiv:2404.10782, 2024) (Year: 2024). [cited by examiner]