Methods, Systems, and Software for Identifying Functional Bio-Molecules
The present invention generally relates to methods of rapidly and efficiently searching biologically-related data space. More specifically, the invention includes methods of identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. The invention also provides methods of modeling sequence-activity relationships. As many of the methods are computer-implemented, the invention additionally provides digital systems and software for performing these methods.
1 . A method for identifying amino acid residues for variation in a protein variant library in order to affect a desired activity, said method comprising:
(a) receiving data characterizing a training set of a protein variant library,
wherein the data provides activity and sequence information for each protein variant in the training set;
(b) from the data, developing a sequence-activity model that predicts activity as a function of amino acid residue type and corresponding position in a protein sequence,
wherein the sequence-activity model includes one or more non-linear terms, each representing an interaction between two or more amino acid residues in the protein sequence; and
(c) using the sequence-activity model to identify one or more amino acid residues at specific positions for variation to impact the desired activity.