Methods, systems, and software for identifying bio-molecules using models of multiplicative form
Abstract
The present invention provides methods for identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. More specifically, some embodiments of the present invention provide methods for building sequence-activity models comprising multiplicative terms and using the models to guide directed evolution. In some embodiments, the sequence-activity models include one or more interaction terms, each of which including an interaction coefficient representing the contribution to activity of two or more defined residues. In some embodiments, the models describe relation between protein or nucleic acid sequences and protein activities. In some embodiments, the present invention also provides methods for preparing sequence-activity models, including but not limited to stepwise addition or subtraction techniques, Bayesian regression, ensemble regression and other methods. The present invention further provides digital systems and software for performing the methods provided herein.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, implemented using a computer system comprising one or more processors and system memory, of conducting directed evolution of proteins, the method comprising,
(a) obtaining, by the computer system, sequence data and activity data of an activity for each of a plurality of protein variants; (b) fitting, by the one or more processors, a sequence-activity model to the sequence data and the activity data, wherein
the sequence-activity model relates amino acids of a protein variant or nucleotides encoding the protein variant to an activity of the protein variant, and
the sequence-activity model comprises a product of multiple terms, wherein two or more terms of the multiple terms each comprise (i) an independent variable representing a defined amino acid of the protein variant or a defined nucleotide encoding the protein variant and (ii) a coefficient representing the contribution to the activity of the protein variant by only the defined amino acid or only the defined nucleotide;
(c) identifying, using the sequence-activity model, one or more polypeptides or one or more polynucleotides; and (d) modifying the one or more polypeptides or one or more polynucleotides in a round of directed evolution.
2 . The method of claim 1 , wherein each of the multiplicative terms is provided in the form of (coefficient×independent variable).
3 . The method of claim 1 , wherein each of the multiplicative terms is provided in the form of (1+coefficient×independent variable).
4 . The method of claim 1 , wherein (c) comprises identifying the one or more polypeptides or the one or more polynucleotides encoding the one or more polypeptides, wherein the one or more polypeptides are predicted by the sequence-activity model to have a desired level of activity.
5 . The method of claim 4 , wherein (d) comprises performing mutagenesis on the one or more polypeptides or the one or more polynucleotides.
6 . The method of claim 1 , wherein the each of the one or more identified biomolecules comprises one or more amino acids or nucleotides associated with one or more coefficients of the sequence-activity model, and wherein the one or more coefficients meet one or more criteria.
7 . The method of claim 1 , wherein (c) comprises:
evaluating coefficients of the sequence-activity model to identify one or more residues at defined sequence positions that contribute to the activity; selecting one or more mutations of the one or more residues; and identifying the one or more polypeptide comprising the one or more mutations or the one or more polynucleotides encoding the one or more mutations.
8 . The method of claim 7 , wherein the evaluating the coefficients comprises identifying one or more coefficients that are determined to be larger than other coefficients.
9 . The method of claim 7 , wherein (d) comprises shuffling the one or more polynucleotides encoding the one or more mutations.
10 . The method of claim 1 , wherein the method further comprises synthesizing the one or more polypeptides or the one or more polynucleotides.
11 . The method of claim 1 , wherein the method further comprises assaying the one or more polypeptides.
12 . The method of claim 1 , further comprising fragmenting and recombining the one or more polynucleotides.
13 . The method of claim 1 , wherein the coefficients of the multiple terms are provided in a look-up table.
14 . The method of claim 1 , wherein at least one of the multiple terms of the sequence-activity model comprises an interaction coefficient representing the contribution to activity of a defined combination of (i) a first amino acid or nucleotide at a first position in the sequence, and (ii) a second amino acid or nucleotide at a second position in the sequence, and
wherein the interaction coefficient represents the contribution to activity of said defined combination.
15 . The method of any of claim 1 , wherein fitting the sequence-activity model comprises performing a stepwise addition or subtraction of terms of the sequence-activity model.
16 . The method of claim 1 , wherein fitting the sequence-activity model comprises using a genetic algorithm to select one or more terms of the sequence-activity model.
17 . The method of claim 1 , further comprising generating two or more sequence-activity models each having the form recited in (b).
18 . The method of claim 17 , further comprising generating an ensemble model including terms from the two or more sequence-activity models, wherein the terms of the ensemble model are weighted by the ability of the two or more sequence-activity models to predict activity.
19 . A computer system comprising one or more processors and system memory, the one or more processors being configured to:
(a) obtain sequence data and activity data of an activity for each of a plurality of protein variants; (b) fit a sequence-activity model to the sequence data and the activity data, wherein
the sequence-activity model relates amino acids of a protein variant or nucleotides encoding the protein variant to an activity of the protein variant, and
the sequence-activity model comprises a product of multiple terms, wherein two or more terms of the multiple terms each comprise (i) an independent variable representing a defined amino acid of the protein variant or a defined nucleotide encoding the protein variant and (ii) a coefficient representing the contribution to the activity of the protein variant by only the defined amino acid or only the defined nucleotide; and
(c) identify, using the sequence-activity model, one or more polypeptides or one or more polynucleotides to modify in a round of directed evolution.
20 . A computer program product comprising a non-transitory machine readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for of conducting directed evolution of proteins, said program code comprising:
(a) code for obtaining sequence data and activity data of an activity for each of a plurality of protein variants; (b) code for fitting a sequence-activity model to the sequence data and the activity data, wherein
the sequence-activity model relates amino acids of a protein variant or nucleotides encoding the protein variant to an activity of the protein variant, and
the sequence-activity model comprises a product of multiple terms, wherein two or more terms of the multiple terms each comprise (i) an independent variable representing a defined amino acid of the protein variant or a defined nucleotide encoding the protein variant and (ii) a coefficient representing the contribution to the activity of the protein variant by only the defined amino acid or only the defined nucleotide; and
(c) code for identifying, using the sequence-activity model, one or more polypeptides or one or more polynucleotides to modify in a round of directed evolution.Join the waitlist — get patent alerts
Track US2017204405A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.