System and method for training machine-learning algorithms for processing biology-related data, a microscope and a trained machine learning algorithm
A system ( 100 ) comprises one or more processors ( 110 ) and one or more storage devices ( 120 ), wherein the system ( 100 ) is configured to generate a first high-dimensional representation of the biology-related language-based input training data ( 102 ) by a language recognition machine-learning algorithm executed by the one or more processors ( 110 ). Further, the system ( 100 ) is configured to generate biology-related language-based output training data based on the first high-dimensional representation by the language recognition machine-learning algorithm and adjust the language recognition machine-learning algorithm based on a comparison of the biology-related language-based input training data ( 102 ) and the biolo-gy-related language-based output training data. Additionally, the system ( 100 ) is configured to generate a second high-dimensional representation of the biology-related image-based input training data ( 104 ) by a visual recognition machine-learning algorithm executed by the one or more processors ( 110 ) and adjust the visual recognition machine-learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.
1 . A microscope comprising a system, the system comprising one or more processors and one or more storage devices, wherein the system is configured to:
receive biology-related language-based input training data, wherein the biology-related language-based input training data is a biological sequence;
generate a first high-dimensional representation of the biology-related language-based input training data by a language recognition machine-learning algorithm executed by the one or more processors, wherein the first high-dimensional representation comprises at least three entries each having a different value;
generate biology-related language-based output training data based on the first high-dimensional representation by the language recognition machine-learning algorithm executed by the one or more processors, wherein the biology-related language-based output training data comprises a prediction on a next element in the biological sequence;
adjust the language recognition machine-learning algorithm based on a comparison of the biology-related language-based input training data and the biology-related language-based output training data;
receive biology-related image-based input training data associated with the biology-related language-based input training data from the microscope and generate a second high-dimensional representation of the biology-related image-based input training data by a visual recognition machine-learning algorithm executed by the one or more processors, wherein the second high-dimensional representation comprises at least three entries each having a different value; and
adjust the visual recognition machine-learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.
2 . The system of claim 1 , wherein the biology-related image-based input training data is image training data of an image of at least one of a biological structure comprising a nucleotide or a nucleotide sequence, a biological structure comprising a protein or a protein sequence, a biological molecule, a biological tissue, a biological structure with a specific behavior, or a biological structure with a specific biological function or a specific biological activity.
3 . The system of claim 1 , wherein values of one or more entries of the first high-dimensional representation are proportional to a likelihood of a presence of a specific biological function or a specific biological activity.
4 . The system of claim 1 , wherein values of one or more entries of the second high-dimensional representation are proportional to a likelihood of a presence of a specific biological function or a specific biological activity.
5 . The system of claim 1 , wherein the first high-dimensional representation and the second high-dimensional representation are numerical representations.
6 . The system of claim 1 , wherein the first high-dimensional representation and the second high-dimensional representation each comprise more than 100 dimensions.
7 . The system of claim 1 , wherein the first high-dimensional representation is a first vector and the second high-dimensional representation is a second vector.
8 . The system of claim 1 , wherein more than 50% of values of the entries of the first high-dimensional representation and more than 50% of values of the entries of the second high-dimensional representation are non-zero.
9 . The system of claim 1 , wherein the values of more than 5 entries of the first high-dimensional representation are larger than 10% of a largest absolute value of the entries of the first high-dimensional representation and the values of more than 5 entries of the second high-dimensional representation are larger than 10% of a largest absolute value of the entries of the second high-dimensional representation.
10 . The system of claim 1 , wherein the comparison of the biology-related language-based input training data and the biology-related language-based output training data for the adjustment of the language recognition machine-learning algorithm is based on a cross entropy loss function.
11 . The system of claim 1 , wherein the comparison of the first high-dimensional representation and the second high-dimensional representation for the adjustment of the visual recognition machine-learning algorithm is based on a cosine similarity loss function.
12 . The system of claim 1 , wherein the biology-related language-based input training data comprises a length of more than 20 characters.
13 . The system of claim 1 , wherein the adjustment of the language recognition machine-learning algorithm comprises an adjustment of a plurality of language recognition neural network weights, wherein a final set of language recognition neural network weights is stored by the one or more storage devices.
14 . The system of claim 1 , wherein the adjustment of the visual recognition machine-learning algorithm comprises an adjustment of a plurality of visual recognition neural network weights, wherein a final set of visual neural network weights is stored by the one or more storage devices.
15 . The system of claim 1 , wherein the language recognition machine-learning algorithm comprises a language recognition neural network.
16 . The system of claim 15 , wherein the language recognition neural network comprises more than 30 layers.
17 . The system of claim 15 , wherein the language recognition neural network is a recurrent neural network.
18 . The system of claim 15 , wherein the language recognition neural network is a long short-term memory network.
19 . The system of claim 1 , wherein the visual recognition machine-learning algorithm comprises a visual recognition neural network.
20 . The system of claim 19 , wherein the visual recognition neural network comprises more than 30 layers.
21 . The system of claim 19 , wherein the visual recognition neural network is a convolutional neural network or a capsule network.
22 . The system of claim 19 , wherein the visual recognition neural network comprises a plurality of convolution layers and a plurality of pooling layers.
23 . The system of claim 19 , wherein the visual recognition neural network uses a rectified linear unit activation function.
24 . The system of claim 1 , wherein the system is configured to repeat, for each biology-related language-based input training data of a training group of biology-related language-based input training data sets, the steps of:
generating the first high-dimensional representation,
generating the biology-related language-based output training data, and
adjusting the language recognition machine-learning algorithm.
25 . The system of claim 24 , wherein a length of the first biology-related language-based input training data of the training group of biology-related language-based input training data sets differs from a length of second biology-related language-based input training data of the training group of biology-related language-based input training data sets.
26 . The system of claim 1 , wherein the system is configured to repeat, for each biology-related image-based input training data of a training group of biology-related image-based input training data sets, the steps of:
generating a second high-dimensional representation, and
adjusting the visual recognition machine-learning algorithm.
27 . The system of claim 26 , wherein the training group of biology-related language-based input training data sets comprises more entries than the training group of biology-related image-based input training data sets.
28 . A method for training machine-learning algorithms for processing biology-related data using a microscope, the method comprising:
receiving biology-related language-based input training data, wherein the biology-related language-based input training data is a biological sequence;
generating a first high-dimensional representation of the biology-related language-based input training data by a language recognition machine-learning algorithm, wherein the first high-dimensional representation comprises at least three entries each having a different value;
generating biology-related language-based output training data based on the first high-dimensional representation by the language recognition machine-learning algorithm, wherein the biology-related language-based output training data comprises a prediction on a next element in the biological sequence;
adjusting the language recognition machine-learning algorithm based on a comparison of the biology-related language-based input training data and the biology-related language-based output training data;
receiving biology-related image-based input training data associated with the biology-related language-based input training data from the microscope;
generating a second high-dimensional representation of the biology-related image-based input training data by a visual recognition machine-learning algorithm, wherein the second high-dimensional representation comprises at least three entries each having a different value; and
adjusting the visual recognition machine-learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.
29 . A non-transitory, computer-readable medium storing a trained machine learning algorithm, the trained machine learning algorithm trained by:
receiving biology-related language-based input training data, wherein the biology-related language-based input training data is a biological sequence;
generating a first high-dimensional representation of the biology-related language-based input training data by a language recognition machine-learning algorithm, wherein the first high-dimensional representation comprises at least three entries each having a different value;
generating biology-related language-based output training data based on the first high-dimensional representation by the language recognition machine-learning algorithm, wherein the biology-related language-based output training data comprises a prediction on a next element in the biological sequence;
adjusting the language recognition machine-learning algorithm based on a comparison of the biology-related language-based input training data and the biology-related language-based output training data;
receiving biology-related image-based input training data associated with the biology-related language-based input training data from a microscope;
generating a second high-dimensional representation of the biology-related image-based input training data by a visual recognition machine-learning algorithm, wherein the second high-dimensional representation comprises at least three entries each having a different value; and
adjusting the visual recognition machine-learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.