IP Library Granted Patent US 12,056,586
Granted Patent B2
US 12,056,586 · App. 17/548,070 · Granted Aug 6, 2024

Data drift impact in a machine learning model

Inventors: Jason Lopatecki (Mill Valley, CA); Aparna Dhinakaran (Dublin, CA); Michael Schiff (Mill Valley, CA)
Assignee: ARIZE AI, INC.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,056,586
App. No.
17/548,070
Granted
Aug 6, 2024
Kind
B2
Abstract

Techniques for determining a drift impact score in a machine learning model are disclosed. The techniques can include: obtaining a reference distribution of a machine learning model; obtaining a current distribution of the machine learning model; determining a statistical distance based on the reference distribution and the current distribution; determining a local feature importance parameter for each feature associated with a prediction made by the machine learning model; determining a cohort feature importance parameter for a cohort of multiple features based on the local feature importance parameter of each feature in the cohort; and determining a drift impact score for the cohort based on the statistical distance and the cohort feature importance parameter.

Claims (30)

1. A computer-implemented method for detecting a bias in a machine learning model, the method comprising:

obtaining a reference distribution of a machine learning model;

obtaining a current distribution of the machine learning model;

binning the reference distribution and the current distribution into a plurality of corresponding bins;

determining a statistical distance based on the binned reference distribution and the binned current distribution;

determining a local feature importance parameter for each feature associated with a prediction made by the machine learning model, wherein the local feature importance parameter of a feature indicates a difference between the model's prediction and an expected prediction;

determining a cohort feature importance parameter for a cohort of multiple features based on averaging values of the local feature importance parameter of each feature in the cohort;

determining a drift impact score for the cohort based on a multiplication of the statistical distance and the cohort feature importance parameter; and

detecting a bias in the machine learning model based on the drift impact score.

2. The method of claim 1 , wherein the determining of the statistical distance is based on a population stability index metric.

3. The method of claim 1 , wherein the determining of the statistical distance is based on a Kullback-Leibler (KL) divergence metric.

4. The method of claim 1 , wherein the determining of the statistical distance is based on a Jensen-Shannon (JS) divergence metric.

5. The method of claim 1 , wherein the determining of the statistical distance is based on an Earth Mover's distance (EMD) metric.

6. The method of claim 1 , wherein the reference distribution is a distribution across a fixed time window or a moving time window.

7. The method of claim 1 , wherein the reference distribution is from a training environment or a production environment.

8. A system for detecting a bias in a machine learning model, the system comprising a processor and an associated memory, the processor being configured for:

obtaining a reference distribution of a machine learning model;

obtaining a current distribution of the machine learning model;

binning the reference distribution and the current distribution into a plurality of corresponding bins;

determining a statistical distance based on the binned reference distribution and the binned current distribution;

determining a local feature importance parameter for each feature associated with a prediction made by the machine learning model, wherein the local feature importance parameter of a feature indicates a difference between the model's prediction and an expected prediction;

determining a cohort feature importance parameter for a cohort of multiple features based on averaging values of the local feature importance parameter of each feature in the cohort;

determining a drift impact score for the cohort based on a multiplication of the statistical distance and the cohort feature importance parameter; and

detecting a bias in the machine learning model based on the drift impact score.

9. The system of claim 8 , wherein the determining of the statistical distance is based on a population stability index metric.

10. The system of claim 8 , wherein the determining of the statistical distance is based on a Kullback-Leibler (KL) divergence metric.

11. The system of claim 8 , wherein the determining of the statistical distance is based on a Jensen-Shannon (JS) divergence metric.

12. The system of claim 8 , wherein the determining of the statistical distance is based on an Earth Mover's distance (EMD) metric.

13. The system of claim 8 , wherein the reference distribution is a distribution across a fixed time window or a moving time window.

14. The system of claim 8 , wherein the reference distribution is from a training environment or a production environment.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE'S NAME PREVIOUSLY RECORDED AT REEL: 059226 FRAME: 0730. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 12, 2022
From: LOPATECKI, JASON; DHINAKARAN, APARNA; SCHIFF, MICHAEL
To: ARIZE AI, INC.
Reel/Frame 062115/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: LOPATECKI, JASON; DHINAKARAN, APARNA; SCHIFF, MICHAEL
To: ARIZA AI, INC
Reel/Frame 059226/0730 →
Continuity (1)
Related Publication 20230186144A1 · Jun 15, 2023
Cited By (1)
US 12,335,116