IP Library Granted Patent US 12,056,588
Granted Patent B2
US 12,056,588 · App. 18/334,183 · Granted Aug 6, 2024

System and method for incremental training of machine learning models in artificial intelligence systems, including incremental training using analysis of network identity graphs

Inventors: Mohamed M. Badawy (Round Rock, TX); Rajat Kabra (Austin, TX); Jostine Fei Ho (Austin, TX)
Assignee: SAILPOINT TECHNOLOGIES, INC.
G06N20/00G06F16/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,056,588
App. No.
18/334,183
Granted
Aug 6, 2024
Kind
B2
Abstract

Systems and methods for embodiments of incremental training of machine learning model in artificial intelligence systems are disclosed. Specifically, embodiments of incremental training of machine learning models using drift detection models are disclosed, including embodiments that utilize drift detection models to determine drift based on identity graphs in artificial intelligence identity management systems.

Claims (59)

1. A method, comprising:

generating an identify graph by:

obtaining identity management data from one or more identity management systems in a distributed enterprise computing environment, the identity management data comprising data associated with a set of entitlements and a set of identities utilized with identity management in the distributed enterprise computing environment, wherein each entitlement of the set of entitlements relates to an access right within the distributed enterprise computing environment;

generating the identity graph from the identity management data by creating a node in the identity graph for each identity and for each entitlement; and

for each identity that is associated with an entitlement of the set of entitlements, creating an edge in the identity graph representing a relationship between the nodes representing the respective identity and respective entitlement;

deriving a first dataset from the identity graph at a first time;

training a first machine learning model used by a machine learning system based on the first dataset;

responsive to determining that an incremental training time interval has elapsed, deriving a second dataset from the identity graph;

applying a drift detection model to the second dataset to determine a drift measure between the second dataset and the first dataset; and

based on the determined drift measure, performing one of:

continuing to use the first machine learning model in the machine learning system,

incrementally training the first machine learning model using a third dataset comprised of data including data from the second dataset, or

training a second machine learning model for use in the machine learning system and replacing the first machine learning model with the second machine learning model.

2. The method of claim 1 , wherein the drift prediction model is trained based on the first dataset.

3. The method of claim 1 , wherein the first dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at the first time and the second dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at a second time.

4. The method of claim 3 , wherein the graph embedding is performed using a graph embedding model.

5. The method of claim 1 , further comprising comparing the drift measure to a warning threshold.

6. The method of claim 5 , further comprising when the drift measure is greater than the warning threshold, generating an alert to a user indicating that data drift is occurring.

7. The method of claim 1 , wherein the second dataset comprises data derived from the identity graph subsequent to a time at which the first dataset is derived.

8. A non-transitory computer readable medium, comprising instructions for:

generating an identify graph by:

obtaining identity management data from one or more identity management systems in a distributed enterprise computing environment, the identity management data comprising data associated with a set of entitlements and a set of identities utilized with identity management in the distributed enterprise computing environment, wherein each entitlement of the set of entitlements relates to an access right within the distributed enterprise computing environment;

generating the identity graph from the identity management data by creating a node in the identity graph for each identity and for each entitlement; and

for each identity that is associated with an entitlement of the set of entitlements, creating an edge in the identity graph representing a relationship between the nodes representing the respective identity and respective entitlement;

deriving a first dataset from the identity graph at a first time;

training a first machine learning model used by a machine learning system based on the first dataset;

responsive to determining that an incremental training time interval has elapsed, deriving a second dataset from the identity graph;

applying a drift detection model to the second dataset to determine a drift measure between the second dataset and the first dataset; and

based on the determined drift measure, performing one of:

continuing to use the first machine learning model in the machine learning system,

incrementally training the first machine learning model using a third dataset comprised of data including data from the second dataset, or

training a second machine learning model for use in the machine learning system and replacing the first machine learning model with the second machine learning model.

9. The non-transitory computer readable medium of claim 8 , wherein the drift prediction model is trained based on the first dataset.

10. The non-transitory computer readable medium of claim 8 , wherein the first dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at the first time and the second dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at a second time.

11. The non-transitory computer readable medium of claim 10 , wherein the graph embedding is performed using a graph embedding model.

12. The non-transitory computer readable medium of claim 8 , further comprising comparing the drift measure to a warning threshold.

13. The non-transitory computer readable medium of claim 12 , further comprising when the drift measure is greater than the warning threshold, generating an alert to a user indicating that data drift is occurring.

14. An identity management system, comprising:

a data store;

a processor;

a non-transitory, computer-readable storage medium, including computer instructions for:

generating an identify graph by:

obtaining identity management data from one or more identity management systems in a distributed enterprise computing environment, the identity management data comprising data associated with a set of entitlements and a set of identities utilized with identity management in the distributed enterprise computing environment, wherein each entitlement of the set of entitlements relates to an access right within the distributed enterprise computing environment;

generating the identity graph from the identity management data by creating a node in the identity graph for each identity and for each entitlement; and

for each identity that is associated with an entitlement of the set of entitlements, creating an edge in the identity graph representing a relationship between the nodes representing the respective identity and respective entitlement;

deriving a first dataset from the identity graph at a first time;

training a first machine learning model used by a machine learning system based on the first dataset;

responsive to determining that an incremental training time interval has elapsed, deriving a second dataset from the identity graph;

applying a drift detection model to the second dataset to determine a drift measure between the second dataset and the first dataset; and

based on the determined drift measure, performing one of:

continuing to use the first machine learning model in the machine learning system,

incrementally training the first machine learning model using a third dataset comprised of data including data from the second dataset, or

training a second machine learning model for use in the machine learning system and replacing the first machine learning model with the second machine learning model.

15. The system of claim 14 , wherein the drift prediction model is trained based on the first dataset.

16. The system of claim 14 , wherein the first dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at the first time and the second dataset comprises data generated from performing graph embedding on at least a portion of the identity graph at a second time.

17. The system of claim 16 , wherein the graph embedding is performed using a graph embedding model.

18. The system of claim 14 , further comprising comparing the drift measure to a warning threshold.

19. The system of claim 18 , further comprising when the drift measure is greater than the warning threshold, generating an alert to a user indicating that data drift is occurring.

20. The system of claim 14 , wherein the second dataset comprises data derived from the identity graph subsequent to a time at which the first dataset is derived.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded Jun 25, 2025
From: SAILPOINT TECHNOLOGIES, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 071724/0511 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: BADAWY, MOHAMED M.; KABRA, RAJAT; HO, JOSTINE FEI
To: SAILPOINT TECHNOLOGIES, INC.
Reel/Frame 064031/0488 →
Continuity (3)
Continuation 17669554 · Feb 11, 2022
Continuation 17180357 · Feb 19, 2021
Related Publication 20230325723A1 · Oct 12, 2023
Cited By (3)
US 12,254,422 US 12,294,584 US 12,413,594