Autonomous supply chain data hub and platform
A system and method of autonomous data hub processing that uses semantic metadata, machine learning models, and a permissioned blockchain to autonomously standardize, identify and correct errors in supply chain data is disclosed. Embodiments input supply chain data stored in a supply chain database, train with the machine learning model trainer, one or more machine learning models to identify one or more data errors in the supply chain data, clean the one or more identified data errors from the supply chain data, and store cleaned supply chain data. Embodiments also update one or more machine learning models to identify one or more data errors in cleaned supply chain data, and join and aggregate one or more sets of cleaned supply chain data.
1 . A system for a machine learning data pipeline to provide a valid and quality aggregate data set for training one or more machine learning models, comprising:
a computer comprising a processor and memory, the computer configured to:
transfer one or more initial data sets to input data;
clean the one or more initial data sets by removing one or more first errors in the one or more initial data sets and store the one or more data sets as valid and quality base data;
join the valid and quality base data and store the joined valid and quality base data as a joined data set;
clean the joined data set by removing one or more second errors in the joined data set and store the cleaned joined data set as a valid and quality joined data set;
aggregate the valid and quality joined data set and store the aggregated valid and quality joined data set as an aggregated data set;
clean the aggregated data set by removing one or more third errors in the aggregated data set and store the cleaned aggregated data set as the valid and quality aggregate data set;
identify, using permissioned blockchain data, a specific system in a supply chain network and a date and a time at which the one or more first errors, the one or more second errors and the one or more third errors were introduced and a manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred, wherein the manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred is due to latency in the permissioned blockchain data, and wherein the permissioned blockchain data comprises blockchain records and blockchain transaction data; and
retrieve the valid and quality aggregate data set for training the one or more machine learning models.
2 . The system of claim 1 , wherein the input data comprises one or more of: item data, location data, item-location data and inventory transaction data.
3 . The system of claim 1 , wherein the one or more initial data sets further comprise the permissioned blockchain data.
4 . The system of claim 1 , wherein the one or more initial data sets further comprise semantic metadata.
5 . The system of claim 1 , wherein the computer is further configured to:
store metadata of one or more machine learning models as semantic metadata in the permissioned blockchain data.
6 . A computer implemented method for a machine learning data pipeline to provide a valid and quality aggregate data set for training one or more machine learning models, comprising:
transferring, by a computer comprising a processor and memory, one or more initial data sets to input data;
cleaning, by the computer, the one or more initial data sets by removing one or more first errors in the one or more initial data sets and store the one or more data sets as valid and quality base data;
joining, by the computer, the valid and quality base data and storing, by the computer, the joined valid and quality base data as a joined data set;
cleaning, by the computer, the joined data set by removing one or more second errors in the joined data set and storing, by the computer, the cleaned joined data set as a valid and quality joined data set;
aggregating, by the computer, the valid and quality joined data set and storing, by the computer, the aggregated valid and quality joined data set as an aggregated data set;
cleaning, by the computer, the aggregated data set by removing one or more third errors in the aggregated data set and storing, by the computer, the cleaned aggregated data set as the valid and quality aggregate data set;
identifying, by the computer using permissioned blockchain data, a specific system in a supply chain network and a date and a time at which the one or more first errors, the one or more second errors and the one or more third errors were introduced and a manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred, wherein the manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred is due to latency in the permissioned blockchain data, and wherein the permissioned blockchain data comprises blockchain records and blockchain transaction data; and
retrieve, by the computer, the valid and quality aggregate data set for training the one or more machine learning models.
7 . The computer implemented method of claim 6 , wherein the input data comprises one or more of: item data, location data, item-location data and inventory transaction data.
8 . The computer implemented method of claim 6 , wherein the one or more initial data sets further comprise the permissioned blockchain data.
9 . The computer implemented method of claim 7 , wherein the one or more initial data sets further comprise semantic metadata.
10 . The computer implemented method of claim 6 , further comprising:
storing, by the computer, metadata of one or more machine learning models as semantic metadata in the permissioned blockchain data.
11 . A non-transitory computer-readable medium embodied with computer program instructions for a machine learning data pipeline to provide a valid and quality aggregate data set for training one or more machine learning models, the computer program instructions when executed by one or more processors:
transfers one or more initial data sets to input data;
cleans the one or more initial data sets by removing one or more first errors in the one or more initial data sets and stores the one or more data sets as valid and quality base data;
joins the valid and quality base data and stores the joined valid and quality base data as a joined data set;
cleans the joined data set by removing one or more second errors in the joined data set and stores the cleaned joined data set as a valid and quality joined data set;
aggregates the valid and quality joined data set and stores the aggregated valid and quality joined data set as an aggregated data set;
cleans the aggregated data set by removing one or more third errors in the aggregated data set and stores the cleaned aggregated data set as the valid and quality aggregate data set;
identifies, using permissioned blockchain data, a specific system in a supply chain network and a date and a time at which the one or more first errors, the one or more second errors and the one or more third errors were introduced and a manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred, wherein the manner in which the one or more first errors, the one or more second errors and the one or more third errors occurred is due to latency in the permissioned blockchain data, and wherein the permissioned blockchain data comprises blockchain records and blockchain transaction data; and
retrieves the valid and quality aggregate data set for training the one or more machine learning models.
12 . The non-transitory computer-readable medium of claim 11 , wherein the input data comprises one or more of: item data, location data, item-location data and inventory transaction data.
13 . The non-transitory computer-readable medium of claim 11 , wherein the one or more initial data sets further comprise the permissioned blockchain data.
14 . The non-transitory computer-readable medium of claim 11 , wherein the one or more initial data sets further comprise semantic metadata.
15 . The non-transitory computer-readable medium of claim 11 , the computer program instructions when executed by the one or more processors further:
stores metadata of one or more machine learning models as semantic metadata in the permissioned blockchain data.