IP Library › Patent Application 19250955
Patent Application
App. No. 19/250,955

SYSTEM AND METHOD FOR OBSERVABILITY AND DATA AUDIT USING IMPLICIT DATA DEPENDENCY CAPTURE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/250,955
Abstract

The disclosure is directed to systems, methods, and computer-readable media for observability and data audit using implicit data dependency capture. Data dependency information can be intercepted, for example, as a user trains or otherwise interacts with a machine learning (ML) model. Data dependency information can include information regarding files, data sources, inputs, outputs, storage buckets, storage directories, and/or other pertinent information. A log of the data dependency information can be reviewed to determine ML model provenance.

Claims (33)

1 . A method comprising:

observing, when training a machine learning model and via a software tool, actions associated with data used to train the machine learning model, wherein the actions are performed by a training command run to train the machine learning model;

generating a database comprising the data; and

determining, based on the data, machine learning model provenance without altering source code associated with executing components of the machine learning model.

2 . The method of claim 1 , wherein the data comprises one or more of file read information, source bucket information, directory information or output file information.

3 . The method of claim 2 , wherein the file read information comprises one or more of a file name, a file identifier, a file version, a date associated with a creation of a file, and a date associated with a most recent change to a file.

4 . The method of claim 2 , wherein the source bucket information comprises one or more of a path, a git repo identifier and commit version, a source computing device, input information, output information, and a database identification.

5 . The method of claim 2 , wherein directory information comprises one or more of a path, a git repo identifier and commit version, a source computing device, input information, output information, a directory identification and a database identification.

6 . The method of claim 1 , wherein the software tool comprises an implicit dependency tracking tool that observes the actions performed by a training command used to train the machine learning model.

7 . The method of claim 6 , wherein the actions comprise one or more of reading a file, performing a git repo done on a project code, obtaining a version of the project code, accessing a cloud-stored object having a date and version, and generating an output file.

8 . The method of claim 7 , wherein the data comprises information about all the actions taken by the implicit dependency tracking tool.

9 . The method of claim 1 , wherein determining, based on the data, the machine learning model provenance further comprises:

determining, via the data in the database, that all actions performed by the training command were from one or more allowed data source; and

identifying, if any, data leakage while training the machine learning model.

10 . The method of claim 6 , wherein the implicit dependency tracking tool operates at an operating system level for local accesses.

11 . The method of claim 6 , wherein the implicit dependency tracking tool acts as an HTTP proxy for remote data accesses to intercept network communications performed by the training command.

12 . A system comprising:

one or more processor; and

a computer-readable storage device storing instructions which, when executed by the one or more processor, cause the one or more processor to be configured to:

observe, when training a machine learning model and via a software tool, actions associated with data used to train the machine learning model, wherein the actions are performed by a training command run to train the machine learning model;

generate a database comprising the data; and

determine, based on the data, machine learning model provenance without altering source code associated with executing components of the machine learning model.

13 . The system of claim 12 , wherein the data comprises one or more of file read information, source bucket information, directory information or output file information.

14 . The system of claim 13 , wherein the file read information comprises one or more of a file name, a file identifier, a file version, a date associated with a creation of a file, and a date associated with a most recent change to a file.

15 . The system of claim 13 , wherein the source bucket information comprises one or more of a path, a git repo identifier and commit version, a source computing device, input information, output information, and a database identification.

16 . The system of claim 13 , wherein directory information comprises one or more of a path, a git repo identifier and commit version, a source computing device, input information, output information, a directory identification and a database identification.

17 . The system of claim 12 , wherein the software tool comprises an implicit dependency tracking tool that observes the actions performed by a training command used to train the machine learning model.

18 . The system of claim 17 , wherein the actions comprise one or more of reading a file, performing a git repo done on a project code, obtaining a version of the project code, accessing a cloud-stored object having a date and version, and generating an output file.

19 . The system of claim 18 , wherein the data comprises information about all the actions taken by the implicit dependency tracking tool.

20 . A software tool for use in connection with implementing operations associated with a companion command, wherein the software tool, when implemented, causes one or more processor to be configured to:

observe, when training a machine learning model, actions associated with data used to train the machine learning model, wherein the actions are performed by a training command run to train the machine learning model;

generate a database comprising the data; and

determine, based on the data, machine learning model provenance without altering source code associated with executing components of the machine learning model.