Distributed computer system and method of operation thereof
Disclosed is a distributed computer system that includes a plurality of worker nodes that are coupled together via a data communication network to exchange data therebetween, wherein collective learning of the worker nodes is managed within the distributed computer system. The distributed computer system comprises a data processing arrangement operable to cluster the plurality of worker nodes into one or more clusters, wherein worker nodes of a given cluster train a computing model by employing a respective secondary distributed ledger. The collective learning from the plurality of worker nodes is coordinated using the distributed ledger arrangement.
1 . A distributed computer system to enable training of at least one computing model, the system comprising:
a plurality of worker nodes coupled via a data communication network, each worker node comprising at least one processor and a local database configured to process and store data;
a distributed ledger arrangement configured to coordinate operation of the system;
at least one secondary distributed ledger configured to coordinate collective learning of the plurality of worker nodes; and
one or more processors configured to execute instructions to:
cluster the plurality of worker nodes into one or more clusters by:
processing encrypted data associated with the worker nodes using a homomorphic encryption algorithm;
transmitting the encrypted data to at least one other worker node and receiving encrypted model predictions from the at least one other worker node; and
determining loss values based on the encrypted data and the encrypted model predictions for use in clustering the plurality of worker nodes;
synchronize, via a cross-ledger validation protocol, records between the distributed ledger arrangement and the at least one secondary distributed ledger;
execute, at the at least one secondary distributed ledger, at least one computer protocol for the collective learning of the plurality of worker nodes; and
cause the plurality of worker nodes within a given cluster to train the at least one computing model using information comprised in the at least one secondary distributed ledger,
wherein the clustering is performed before starting training of the at least one computing model.
2 . The distributed computer system of claim 1 , wherein the data processing arrangement is configured to cluster the plurality of worker nodes into one or more clusters based on estimated generalization performance of a shared classifier on a task associated with the worker nodes.
3 . The distributed computer system of claim 1 , wherein the system is configured to initiate the secondary distributed ledger by employing at least one of: staking of tokens in the distributed ledger, smart contract, off-chain contractual agreements.
4 . The distributed computer system of claim 1 , wherein each of the secondary distributed ledgers is configured to store cryptographically signed records pertaining to the respective computing model, worker nodes associated therewith, and a contribution of each of the worker nodes in their respective cluster in relation to training the respective computing model.
5 . The distributed computer system of claim 1 , wherein the data processing arrangement is configured to validate a given block in the distributed ledger arrangement when the given block is referenced by a block that has been notarized by a majority of the plurality of worker nodes.
6 . The distributed computer system of claim 1 , wherein the data processing arrangement is configured to employ verifiable delay functions to space a timing of block production according to a priority of each worker node within a given cluster.
7 . The distributed computer system of claim 1 , wherein the one or more processors are configured to apply the homomorphic encryption algorithm to model-update data vectors generated by the plurality of worker nodes prior to transmission of the encrypted data between worker nodes for determining the loss values used in clustering the plurality of worker node.
8 . The distributed computer system of claim 1 , wherein the one or more processors are configured to generate a matrix of loss values representing model performance across pairs of worker nodes and to cluster the plurality of worker nodes based on the generated matrix of loss values.
9 . The distributed computer system of claim 1 , wherein the one or more processors are configured to:
transmit encrypted data from a first worker node to a second worker node;
process the encrypted data at the second worker node using a machine learning model associated with the second worker node to generate encrypted model predictions; and
transmit the encrypted model predictions to the first worker node.
10 . A method for operating a distributed computer system to enable training of at least one computing model, wherein the distributed computer system comprises a plurality of worker nodes that are coupled via a data communication network to exchange data therebetween, each worker node comprising at least one processor and a local database configured to process and store data therein, and wherein operation of the distributed computer system is coordinated by a distributed ledger arrangement, the method comprising:
clustering the plurality of worker nodes into one or more clusters by:
processing encrypted data associated with the worker nodes using a homomorphic encryption algorithm;
transmitting the encrypted data to at least one other worker node and receiving encrypted model predictions from the at least one other worker node; and
determining loss values based on the encrypted data and the encrypted model predictions for use in clustering the plurality of worker nodes;
synchronizing, via a cross-ledger validation protocol, records between the distributed ledger arrangement and at least one secondary distributed ledger;
executing, at the at least one secondary distributed ledger, at least one computer protocol for collective learning of the plurality of worker nodes; and
training the at least one computing model at the plurality of worker nodes within a given cluster using information comprised in the at least one secondary distributed ledger,
wherein the clustering is performed before starting training of the at least one computing model.
11 . The method of claim 10 , wherein the method comprises clustering the plurality of worker nodes into one or more clusters based on estimated generalization performance of a shared classifier on a task associated with the worker nodes.
12 . The method of claim 10 , wherein the method comprises initiating at least one secondary distributed ledger by employing at least one of: staking of tokens in the distributed ledger arrangement, smart contract logic, off-chain contractual agreements.
13 . The method of claim 10 , wherein each of the secondary distributed ledgers stores cryptographically signed records pertaining to the respective computing model, worker nodes associated therewith, and a contribution of each of the worker nodes in the respective cluster to train the respective computing model.
14 . The method of claim 10 , wherein the method comprises validating a given block in the distributed ledger arrangement when the given block is referenced by a block that has been notarized by a majority of the plurality of worker nodes.
15 . The method of claim 10 , wherein the method comprises employing verifiable delay functions to space a timing of block production according to a priority of each worker node within a given cluster.
16 . The method of claim 10 , further comprising checking finality of the distributed ledger arrangement and the at least one secondary distributed ledger and relaying the information between the distributed ledger arrangement and the at least one secondary distributed ledger.
17 . The method of claim 10 , further comprising generating a matrix of loss values representing model performance across pairs of worker nodes and clustering the plurality of worker nodes based on the generated matrix of loss values.
18 . The method of claim 10 , further comprising:
transmitting encrypted data from a first worker node to a second worker node;
processing the encrypted data at the second worker node using a machine learning model associated with the second worker node to generate encrypted model predictions; and
transmitting the encrypted model predictions to the first worker node.