US2015134315A1PendingUtilityA1
Structure based predictive modeling
Est. expirySep 27, 2033(~7.2 yrs left)· nominal 20-yr term from priority
G16C 10/00G16C 20/50G16B 35/20C12N 15/1058G06F 19/12G06F 19/24G06F 19/16G16B 10/00G16C 20/30G16B 15/00G16B 15/30G16B 5/00G16B 40/00G16B 15/20
56
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Disclosed are methods for building a sequence activity model with reference to structural data, which model can be used to guide directed evolution of proteins having beneficial properties. Some embodiments use genetic algorithms and structural data to filter out uninformative data. Some embodiments use a support vector machine to train the sequence activity model. The filtering and training methods can generate a sequence activity model having higher predictive power than conventional modeling methods. Systems and computer program products implementing the methods are also provided.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of conducting directed evolution, the method comprising:
(a) receiving a data set having information from physical measurements of molecules, wherein the data set comprises the following information for each of a plurality of variant biomolecules: (i) activity of the variant biomolecule on a ligand in a binding site of the variant biomolecule, (ii) a sequence of the variant biomolecule, and (iii) one or more geometric parameters characterizing the geometry of the ligand in the binding site; (b) filtering the data set to produce a filtered data subset by removing information for one or more of the variant biomolecules, wherein the filtering comprises testing the predictive power of sequence activity models trained with a plurality of selected data subsets, each selected data subset having information for a particular set of variant biomolecules removed from the data set of (a); and (c) training an improved sequence activity model using the filtered data subset.
2 . The method of claim 1 , wherein filtering the data set comprises removing at least one of the geometric parameters from the data set.
3 . The method of claim 1 , wherein the filtering the data set is performed with a genetic algorithm.
4 . The method of claim 3 , wherein the genetic algorithm varies thresholds for removing information associated with the geometric parameters for one or more of the variant biomolecules.
5 . The method of claim 1 , further comprising applying the improved sequence activity model to identify one or more new biomolecule variants predicted by the improved sequence activity model to have activity meeting one or more criteria, wherein each of the one or more new biomolecule variants has a sequence that differs from the sequences of the biomolecule variants providing information for the data set of (a).
6 . The method of claim 5 , wherein applying the improved sequence activity model to identify one or more new biomolecule variants comprises performing a genetic algorithm in which potential new biomolecule variants are evaluated using the improved sequence activity model as a fitness function.
7 . The method of claim 5 , further comprising assaying the new biomolecule variants for activity.
8 . The method of claim 5 , further comprising
producing a structural model for each of the new biomolecule variants; and using the structural models to generate geometric parameters for binding sites of the new biomolecule variants, wherein the geometric parameters characterize the geometry of the ligand in the binding sites of the new biomolecule variants.
9 . The method of claim 1 , further comprising measuring the activity of the variant biomolecules by an in vitro assay.
10 . The method of claim 1 , further comprising receiving structural models of biomolecule variants and determining the one or more geometric parameters using the structural models.
11 . The method of claim 10 , wherein the structural models are homology models.
12 . The method of claim 10 , wherein the homology models are prepared using physical structural measurement details of biomolecules.
13 . The method of claim 12 , wherein the physical structural measurement details of biomolecules comprise three-dimensional positions of atoms obtained by NMR or x-ray crystallography.
14 . The method of claim 10 , further comprising using a docker to determine the one or more geometric parameters.
15 . The method of claim 1 , wherein the information for each of a plurality of variant biomolecules further comprises (iv) an interaction energy characterizing the interaction of the ligand in the binding site.
16 . The method of claim 15 , further comprising using a docker to determine the interaction energy
17 . The method of claim 1 , wherein the improved sequence activity model is obtained by a support vector machine, a multiple linear regression, a principal component regression, a partial least square regression, or a neural network.
18 . The method of claim 17 , wherein the sequence activity model is obtained by a support vector machine.
19 . The method of claim 1 , wherein the plurality of variant biomolecules comprises a plurality of enzymes.
20 . The method of claim 19 , wherein the activity of the variant biomolecule on a ligand is the activity of an enzyme on a substrate.
21 . The method of claim 20 , wherein the activity of an enzyme on a substrate comprises one or more features of a catalytic conversion of the substrate by the enzyme.
22 . The method of claim 1 , further comprising using the improved sequence activity model to identifying one or more biomolecules having desired activity.
23 . The method of claim 22 , further comprising synthesizing the biomolecules having desired activity.
24 . A computer program product comprising one or more computer-readable non-transitory storage media having stored thereon computer-executable instructions that, when executed by one or more processors of a computer system, cause the computer system to implement a method for conducting directed evolution, the method comprising:
(a) receiving, by the computer system, a data set having information from physical measurements of molecules, wherein the data set comprises the following information for each of a plurality of variant biomolecules: (i) activity of the variant biomolecule on a ligand in a binding site of the variant biomolecule, (ii) a sequence of the variant biomolecule, and (iii) one or more geometric parameters characterizing the geometry of the ligand in the binding site; (b) filtering, by the computer system, the data set to produce a filtered data subset by removing information for one or more of the variant biomolecules, wherein the filtering comprises testing the predictive power of sequence activity models trained with a plurality of selected data subsets, each selected data subset having information for a particular set of variant biomolecules removed from the data set of (a); and (c) training, by the computer system, an improved sequence activity model using the filtered data subset.
25 . A computer system, comprising:
one or more processors; system memory; and one or more computer-readable storage media having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computer system to implement a method for conducting directed evolution, the method comprising:
(a) receiving a data set having information from physical measurements of molecules, wherein the data set comprises the following information for each of a plurality of variant biomolecules: (i) activity of the variant biomolecule on a ligand in a binding site of the variant biomolecule, (ii) a sequence of the variant biomolecule, and (iii) one or more geometric parameters characterizing the geometry of the ligand in the binding site;
(b) filtering the data set to produce a filtered data subset by removing information for one or more of the variant biomolecules, wherein the filtering comprises testing the predictive power of sequence activity models trained with a plurality of selected data subsets, each selected data subset having information for a particular set of variant biomolecules removed from the data set of (a); and
(c) training an improved sequence activity model using the filtered data subset.Join the waitlist — get patent alerts
Track US2015134315A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.