IP Library › Granted Patent US 12,737,623
Granted Patent B2
US 12,737,623 · App. 18/190,268 · Granted Sep 15, 2026

Reverse data generation and data distribution analysis to validate artificial intelligence model

Inventors: Zhong Fang Yuan (Xi'An, CN); Tong Liu (Xi'An, CN); Shuang Yin Liu (Beijing, CN); Jun Wang (Xi'An, CN); Yan Fen Liu (Tianjin, CN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,623
App. No.
18/190,268
Granted
Sep 15, 2026
Kind
B2
Abstract

Validity of a trained artificial intelligence model is verified. The verifying the validity includes generating a training dataset from the trained artificial intelligence model using reverse data generation of the trained artificial intelligence model. The training dataset generated using the reverse data generation is compared with a test dataset used to evaluate the trained artificial intelligence model. The comparing is to determine a relationship between the training dataset that was generated and the test dataset. Data from the test dataset determined to have a predefined relationship with the training dataset is removed to obtain a new test dataset. The new test dataset is used to verify the validity of the trained artificial intelligence model.

Claims (50)

1 . A computer-implemented method of facilitating processing within a computing environment, the computer-implemented method comprising:

verifying validity of a trained artificial intelligence model used in artificial intelligence processing in one or more environments, the verifying validity including:

generating a training dataset from the trained artificial intelligence model using reverse data generation of the trained artificial intelligence model, wherein the training dataset is generated using a simulation of a distribution of data represented by the trained artificial intelligence model, wherein random simulation inputs are passed through the trained artificial intelligence model which attaches confidences to the simulation inputs and wherein the training dataset is generated based thereon;

comparing the training dataset generated using the reverse data generation with a test dataset used to evaluate the trained artificial intelligence model, the comparing to determine a relationship between the training dataset that was generated and the test dataset;

removing data from the test dataset determined to have a predefined relationship with the training dataset to obtain a new test dataset;

using the new test dataset to verify the validity of the trained artificial intelligence model; and

automatically updating, using program code including a learning agent, the trained artificial intelligence model based on feedback and selected data from one or more sources to increase accuracy of the trained artificial intelligence model in performing the artificial intelligence processing.

2 . The computer-implemented method of claim 1 , wherein the generating the training dataset includes constructing a simulation dataset of the trained artificial intelligence model, the simulation dataset being the simulation of the distribution of data represented by the trained artificial intelligence model, and using the simulation dataset to generate the training dataset.

3 . The computer-implemented method of claim 2 , wherein the constructing the simulation dataset includes randomly generating vector data to conform to a selected distribution model.

4 . The computer-implemented method of claim 2 , wherein the generating the training dataset further includes:

passing the simulation dataset through the trained artificial intelligence model to obtain confidence values for simulation data of the simulation dataset;

comparing the confidence values to a confidence comparator value; and

forming a retained dataset that includes the simulation data that have confidence values with a predetermined relationship with the confidence comparator value, the retained dataset to be used to generate the training dataset.

5 . The computer-implemented method of claim 4 , wherein the generating the training dataset further includes filtering the retained dataset to obtain the training dataset, the filtering including removing redundancy from the retained dataset to obtain the training dataset.

6 . The computer-implemented method of claim 5 , wherein the filtering includes performing density clustering to partition the retained dataset and remove the redundancy.

7 . The computer-implemented method of claim 1 , wherein the comparing is based on data distributions of the training dataset and the test dataset, and wherein test dataset data that overlaps training dataset data are removed from the test dataset.

8 . The computer-implemented method of claim 1 , wherein the comparing the training dataset and the test dataset includes performing anomaly detection on a mix of the training dataset and the test dataset to obtain the new test dataset.

9 . The computer-implemented method of claim 8 , wherein the performing the anomaly detection includes using an estimation network of a selected anomaly detection technique to detect a degree of integration in the training dataset and the test dataset.

10 . The computer-implemented method of claim 1 , wherein the generating the training dataset is performed absent availability of a dataset used to train the trained artificial intelligence model.

11 . A computer system for facilitating processing within a computing environment, the computer system comprising:

a memory; and

a computing device in communication with the memory, wherein the computer system is configured to perform a method, said method comprising:

verifying validity of a trained artificial intelligence model used in artificial intelligence processing in one or more environments, the verifying validity including:

generating a training dataset from the trained artificial intelligence model using reverse data generation of the trained artificial intelligence model, wherein the training dataset is generated using a simulation of a distribution of data represented by the trained artificial intelligence model, wherein random simulation inputs are passed through the trained artificial intelligence model which attaches confidences to the simulation inputs and wherein the training dataset is generated based thereon;

comparing the training dataset generated using the reverse data generation with a test dataset used to evaluate the trained artificial intelligence model, the comparing to determine a relationship between the training dataset that was generated and the test dataset;

removing data from the test dataset determined to have a predefined relationship with the training dataset to obtain a new test dataset;

using the new test dataset to verify the validity of the trained artificial intelligence model; and

automatically updating, using program code including a learning agent, the trained artificial intelligence model based on feedback and selected data from one or more sources to increase accuracy of the trained artificial intelligence model in performing the artificial intelligence processing.

12 . The computer system of claim 11 , wherein the generating the training dataset includes constructing a simulation dataset of the trained artificial intelligence model, the simulation dataset being the simulation of the distribution of data represented by the trained artificial intelligence model, and using the simulation dataset to generate the training dataset.

13 . The computer system of claim 12 , wherein the generating the training dataset further includes:

passing the simulation dataset through the trained artificial intelligence model to obtain confidence values for simulation data of the simulation dataset;

comparing the confidence values to a confidence comparator value; and

forming a retained dataset that includes the simulation data that have confidence values with a predetermined relationship with the confidence comparator value, the retained dataset to be used to generate the training dataset.

14 . The computer system of claim 13 , wherein the generating the training dataset further includes filtering the retained dataset to obtain the training dataset, the filtering including removing redundancy from the retained dataset to obtain the training dataset.

15 . The computer system of claim 11 , wherein the comparing is based on data distributions of the training dataset and the test dataset, and wherein test dataset data that overlaps training dataset data are removed from the test dataset.

16 . A computer program product for facilitating processing within a computing environment, the computer program product comprising:

one or more computer readable storage media and program instructions collectively stored on the one or more computer readable storage media to perform a method comprising:

verifying validity of a trained artificial intelligence model used in artificial intelligence processing in one or more environments, the verifying validity including:

generating a training dataset from the trained artificial intelligence model using reverse data generation of the trained artificial intelligence model, wherein the training dataset is generated using a simulation of a distribution of data represented by the trained artificial intelligence model, wherein random simulation inputs are passed through the trained artificial intelligence model which attaches confidences to the simulation inputs and wherein the training dataset is generated based thereon;

comparing the training dataset generated using the reverse data generation with a test dataset used to evaluate the trained artificial intelligence model, the comparing to determine a relationship between the training dataset that was generated and the test dataset;

removing data from the test dataset determined to have a predefined relationship with the training dataset to obtain a new test dataset;

using the new test dataset to verify the validity of the trained artificial intelligence model; and

automatically updating, using program code including a learning agent, the trained artificial intelligence model based on feedback and selected data from one or more sources to increase accuracy of the trained artificial intelligence model in performing the artificial intelligence processing.

17 . The computer program product of claim 16 , wherein the generating the training dataset includes constructing a simulation dataset of the trained artificial intelligence model, the simulation dataset being the simulation of the distribution of data represented by the trained artificial intelligence model, and using the simulation dataset to generate the training dataset.

18 . The computer program product of claim 17 , wherein the generating the training dataset further includes:

passing the simulation dataset through the trained artificial intelligence model to obtain confidence values for simulation data of the simulation dataset;

comparing the confidence values to a confidence comparator value; and

forming a retained dataset that includes the simulation data that have confidence values with a predetermined relationship with the confidence comparator value, the retained dataset to be used to generate the training dataset.

19 . The computer program product of claim 18 , wherein the generating the training dataset further includes filtering the retained dataset to obtain the training dataset, the filtering including removing redundancy from the retained dataset to obtain the training dataset.

20 . The computer program product of claim 16 , wherein the comparing is based on data distributions of the training dataset and the test dataset, and wherein test dataset data that overlaps training dataset data are removed from the test dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: YUAN, ZHONG FANG; LIU, TONG; LIU, SHUANG YIN; WANG, JUN; LIU, YAN FEN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063127/0266 →
Continuity (1)
Related Publication 20240330675A1 · Oct 3, 2024
References Cited (11)
US 11068453B2 · Griffith · 2021 [cited by applicant]
US 11163527B2 · Oberbreckling et al. · 2021 [cited by applicant]
US 20150039591A1 · Ding · 2015 [cited by examiner]
US 20200019890A1 · Rossi et al. · 2020 [cited by applicant]
US 20200356901A1 · Zarandioon et al. · 2020 [cited by applicant]
US 20220188703A1 · Hong · 2022 [cited by examiner]
US 20220230024A1 · Kallianpur et al. · 2022 [cited by applicant]
US 20230222360A1 · Liu · 2023 [cited by examiner]
KR 20220085740A · 2022 [cited by applicant]
D. Dahmen et al., “Generation of Virtual test scenarios for training and validation of AI-based systems”, 2021 International conference on progress in informatics and computing (PIC) , 2021 (Year: 2021). [cited by examiner]
Sun, Yiyou et al., “Out-of-Distribution Detection with Deep Nearest Neighbors,” Proceedings of the 39 [cited by applicant]