US2023118842A1PendingUtilityA1

Mining all atom simulations for diagnosing and treating disease

Assignee: UNIV GEORGE MASONPriority: Dec 20, 2017Filed: Dec 16, 2022Published: Apr 20, 2023
Est. expiryDec 20, 2037(~11.4 yrs left)· nominal 20-yr term from priority
G16B 20/20G16B 40/20C12Q 1/6827C12Q 1/6883G16B 40/30G16B 35/00G16B 20/00G16B 10/00G16B 5/20G16B 50/50G16B 50/10
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure describes methods for determining the functional consequences of mutations. The methods include the use of machine learning to identify and quantify features of all atom molecular dynamics simulations to obtain the disruptive severity of genetic variants on molecular function.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 identifying at least one structure of a wildtype macromolecule;   seeding and executing a plurality of all atom molecular dynamics (MD) simulations using the at least one structure to generate a conformational landscape of the wildtype macromolecule and a plurality of variants of the wildtype macromolecule, wherein generating the conformational landscape comprises generating, via the plurality of all atom MD simulations, a plurality of variant structures corresponding to each of the plurality of variants of the wildtype macromolecule;   generating a first dataset by determining, from the conformational landscape, structural features and energetic features of the at least one structure and the plurality of variant structures;   processing the first dataset via a first clustering algorithm to generate a wildtype cluster of the at least one structure and a plurality of variant clusters, wherein each of the plurality of variant clusters defines a respective conformational population comprising a different subset of the plurality of variant structures;   generating a second dataset by quantifying structural and energetic features of the wildtype cluster and the respective conformational population defined by each of the plurality of variant clusters; and   performing principal component analysis on the second dataset to generate a principal component feature space of the wildtype macromolecule and the plurality of variants of the wildtype macromolecule.   
     
     
         2 . The method of  claim 1 , wherein the structural features and energetic features comprise global structural features and dynamics of structure subdomains. 
     
     
         3 . The method of  claim 1 , wherein the structural features and energetic features comprise energetic interactions and overall statistical characteristics of structures. 
     
     
         4 . The method of  claim 1 , wherein the plurality of variants of the wildtype macromolecule are variant proteins. 
     
     
         5 . The method of  claim 4 , wherein the variant proteins cause a disease. 
     
     
         6 . The method of  claim 5 , wherein the disease is a genetic disease. 
     
     
         7 . The method of  claim 6 , wherein the genetic disease is Long QT Syndrome or Polymorphic Ventricular Tachycardia. 
     
     
         8 . A system comprising:
 memory comprising structural data of at least one wildtype macromolecule; and   at least one computing device in communication with the memory and configured to:
 identify at least one structure of the at least one wildtype macromolecule based on the structural data; 
 seed and execute a plurality of all atom molecular dynamics (MD) simulations using the at least one structure to generate a conformational landscape of the at least one wildtype macromolecule and a plurality of variants of the at least one wildtype macromolecule, wherein the step of generating the conformational landscape comprises the at least one computing device being configured to generate, via the plurality of all atom MD simulations, a plurality of variant structures corresponding to each of the plurality of variants of the at least one wildtype macromolecule; 
 generate a first dataset by determining, from the conformational landscape, structural features and energetic features of the at least one structure and the plurality of variant structures; 
 process the first dataset via a first clustering algorithm to generate a wildtype cluster of the at least one structure and a plurality of variant clusters, wherein each of the plurality of variant clusters defines a respective conformational population comprising a different subset of the plurality of variant structures; 
 generate a second dataset by quantifying structural and energetic features of the wildtype cluster and the respective conformational population defined by each of the plurality of variant clusters; and 
 perform principal component analysis on the second dataset to generate a principal component feature space of the at least one wildtype macromolecule and the plurality of variants of the at least one wildtype macromolecule. 
   
     
     
         9 . The system of  claim 8 , wherein the structural features and energetic features comprise global structural features and dynamics of structure subdomains. 
     
     
         10 . The system of  claim 8 , wherein the structural features and energetic features comprise energetic interactions and overall statistical characteristics of structures. 
     
     
         11 . The system of  claim 8 , wherein the plurality of variants of the wildtype macromolecule are variant proteins. 
     
     
         12 . The system of  claim 11 , wherein the variant proteins cause a disease. 
     
     
         13 . The system of  claim 12 , wherein the disease is a genetic disease. 
     
     
         14 . The system of  claim 13 , wherein the genetic disease is Long QT Syndrome or Polymorphic Ventricular Tachycardia. 
     
     
         15 . A non-transitory, computer-readable medium embodying a program that, when executed by at least one computing device, causes the at least one computing device to:
 identify at least one structure of a wildtype macromolecule;   seed and execute a plurality of all atom molecular dynamics (MD) simulations using the at least one structure to generate a conformational landscape of the wildtype macromolecule and a plurality of variants of the wildtype macromolecule, wherein the step of generating the conformational landscape comprises the program, when executed by the at least one computing device, causing the at least one computing device to generate, via the plurality of all atom MD simulations, a plurality of variant structures corresponding to each of the plurality of variants of the wildtype macromolecule;   generate a first dataset by determining, from the conformational landscape, structural features and energetic features of the at least one structure and the plurality of variant structures;   process the first dataset via a first clustering algorithm to generate a wildtype cluster of the at least one structure and a plurality of variant clusters, wherein each of the plurality of variant clusters defines a respective conformational population comprising a different subset of the plurality of variant structures;   generate a second dataset by quantifying structural and energetic features of the wildtype cluster and the respective conformational population defined by each of the plurality of variant clusters; and   perform principal component analysis on the second dataset to generate a principal component feature space of the wildtype macromolecule and the plurality of variants of the wildtype macromolecule.   
     
     
         16 . The non-transitory, computer-readable medium of  claim 15 , wherein the structural features and energetic features comprise global structural features and dynamics of structure subdomains. 
     
     
         17 . The non-transitory, computer-readable medium of  claim 15 , wherein the structural features and energetic features comprise energetic interactions and overall statistical characteristics of structures. 
     
     
         18 . The non-transitory, computer-readable medium of  claim 15 , wherein the plurality of variants of the wildtype macromolecule are variant proteins. 
     
     
         19 . The non-transitory, computer-readable medium of  claim 18 , wherein the variant proteins cause a disease. 
     
     
         20 . The non-transitory, computer-readable medium of  claim 19 , wherein the disease is a genetic disease.

Join the waitlist — get patent alerts

Track US2023118842A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.