Methods, systems, and software for identifying bio-molecules with interacting components
Abstract
The present invention provides methods for rapidly and efficiently searching biologically-related data space. More specifically, the present invention provides methods for identifying bio-molecules with desired properties, or which are most suitable for acquiring such properties, from complex bio-molecule libraries or sets of such libraries. The present invention also provides methods for modeling sequence-activity relationships, including but not limited to stepwise addition or subtraction techniques, Bayesian regression, ensemble regression and other methods. The present invention further provides digital systems and software for performing the methods provided herein.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, implemented using a computer system comprising one or more processors and system memory, of conducting directed evolution of bio-molecules, the method comprising:
(a) receiving, by the computer system, sequence data and activity data for a plurality of bio-molecules; (b) fitting, by the one or more processors, a first or second base model to the sequence data and activity data, wherein
the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule,
the first base model includes one or more linear terms but no interaction term,
the second base model includes one or more linear terms and a defined pool of interaction terms, and
each interaction term represents an interaction between two or more interacting sub-units;
(c) obtaining, by the one or more processors, a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms; (d) determining, by the one or more processors, an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units; (e) identifying, by the one or more processors, at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms; (f) identifying, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.
2 . The method of claim 1 , wherein obtaining the plurality of new models in (c) comprises using prior information of parameters of the plurality of new models to determine posterior probability distributions of the parameters of the plurality of new models.
3 . The method of claim 2 , wherein the fitting a first or second base model and/or obtaining the plurality of new models comprises using Gibbs sampling to fit a model to the sequence and activity data.
4 . The method of claim 1 , wherein the at least one best model comprises two or more best models, each of which includes different interaction terms.
5 . The method of claim 1 , further comprising preparing an ensemble model based on the two or more best models, wherein
the ensemble model includes interaction terms from the two or more best models, and the interaction terms are weighted by the ability of the two or more best models to predict activity as determined in (d).
6 . The method of claim 1 , further comprising:
(i) setting the at least one best model as an updated model and repeating (c) using the updated model in place of the first or second base model; and (ii) repeating (d) and (e).
7 . The method of claim 6 , further comprising repeating (i) and (ii) one or more times.
8 . The method of claim 1 , wherein the ability of the plurality of new models to predict the activity in (d) is measured by Akaike Information Criterion or Bayesian Information Criterion.
9 . The method of claim 1 , wherein the sequence is a whole genome, whole chromosome, chromosome segment, a collection of gene sequences for interacting genes, gene, or protein.
10 . The method of any of claim 1 , wherein the sub-units are chromosomes, chromosome segments, haplotypes, genes, nucleotides, codons, mutations, amino acids, or residues.
11 . The method of claim 1 , wherein the plurality of bio-molecules comprises a protein variant library.
12 . The method of claim 1 , wherein (f) comprises identifying the one or more identified bio-molecules that are predicted by the at least one best model to have a desired level of activity.
13 . The method of claim 12 , further comprising performing saturation mutagenesis on the one or more identified bio-molecules that are predicted by the at least one best model to have the desired level of activity.
14 . The method of claim 1 , wherein the each of the one or more identified bio-molecules comprises sub-units associated with one or more coefficients of the at least one best model, and wherein the one or more coefficients meet one or more criteria.
15 . The method of claim 1 , wherein (f) comprises:
evaluating coefficients of the one or more linear terms and/or coefficients of the one or more interaction terms of the at least one best model to identify one or more defined amino acids at defined sequence positions that contribute to the activity; selecting one or more mutations of the defined amino acids; and identifying one or more oligonucleotides encoding the one or more mutations.
16 . The method of claim 15 , wherein the evaluating coefficients comprises identifying one or more coefficients that are determined to be larger than other coefficients.
17 . The method of claim 1 , wherein the one or more identified bio-molecules comprise one or more nucleic acid molecules, and wherein the method further comprises synthesizing the one or more nucleic acid molecules using a nucleic acid synthesizer.
18 . The method of claim 1 , further comprising fragmenting and recombining the one or more identified bio-molecules.
19 . A computer system comprising one or more processors and system memory, the one or more processors being configured to:
(a) receive sequence data and activity data for a plurality of bio-molecules; (b) fit a first or second base model to the sequence data and activity data, wherein
the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule,
the first base model includes one or more linear terms but no interaction term,
the second base model includes one or more linear terms and a defined pool of interaction terms, and
each interaction term represents an interaction between two or more interacting sub-units;
(c) obtain a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms; (d) determine an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units; (e) identify at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms; and (f) identify, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.
20 . A computer program product comprising a non-transitory machine readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for conducting directed evolution of bio-molecules, said program code comprising:
(a) code for receiving sequence data and activity data for a plurality of bio-molecules; (b) code for fitting a first or second base model to the sequence data and activity data, wherein
the first or second base model relates sub-units of a sequence of a bio-molecule to an activity of the bio-molecule,
the first base model includes one or more linear terms but no interaction term,
the second base model includes one or more linear terms and a defined pool of interaction terms, and
each interaction term represents an interaction between two or more interacting sub-units;
(c) code for obtaining a plurality of new models, wherein each new model is obtained by adding to the first base model one different interaction term in the defined pool of interaction terms or subtracting from the second base model one different interaction term in the defined pool of interaction terms; (d) code for determining an ability of each model of the plurality of new models to predict the activity as a function of the presence or absence of the sub-units; (e) code for identifying at least one best model among the plurality of new models based on the ability of the plurality of new models to predict activity as determined in (d) and with a bias against including additional interaction terms; and (f) code for identify, using the at least one best model, one or more identified bio-molecules to modify in a round of directed evolution.Join the waitlist — get patent alerts
Track US2017211206A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.