US2023178186A1PendingUtilityA1

Generation of protein sequences using machine learning techniques

Assignee: JUST EVOTEC BIOLOGICS INCPriority: May 19, 2019Filed: Jan 13, 2023Published: Jun 8, 2023
Est. expiryMay 19, 2039(~12.8 yrs left)· nominal 20-yr term from priority
G06N 20/00G16B 35/10G16B 40/20G16B 30/00G06N 3/123G16B 20/30G06N 3/0475G06N 3/0455G06N 3/0464G06N 3/08
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Amino acid sequences of antibodies can be generated using a generative adversarial network that includes a first generating component that generates amino acid sequences of antibody light chains and a second generating component that generates amino acid sequences of antibody heavy chains. Amino acid sequences of antibodies call be produced by combining the respective amino acid sequences produced by the first generating component and the second generating component. The training of the first generating component and the second generating component can proceed at different rates. Additionally, the antibody amino acids produced by combining amino acid sequences front the first generating component and the second generating component may be evaluated according to complentarity-determining regions of the antibody amino acid sequences. Training datasets may be produced using amino acid sequences that correspond to antibodies have particular binding affinities with respect to molecules, such as binding affinity with major histocompatibility complex (MHC) molecules.

Claims

exact text as granted — not AI-modified
1 - 16 . (canceled) 
     
     
         17 . A method comprising:
 obtaining, by a computing system including one or more computing devices having one or more processors and memory, a training dataset including amino acid sequences of proteins;   generating, by the computing system, structured amino acid sequences based on the training dataset;   generating, by the computing system, a model to produce additional amino acid sequences that correspond to the amino acid sequences included in the training dataset using the structured amino acid sequences and a generative adversarial network;   generating, by the computing system, the additional amino acid sequences using the model and an input vector; and   evaluating, by the computing system, the additional amino acid sequences according to one or more criteria to determine metrics for the additional amino acid sequences.   
     
     
         18 . The method of  claim 17 , comprising determining, by the computing system, a number of variations of an amino acid sequence included in the additional amino acid sequences with respect to a protein derived from a gene of a germline. 
     
     
         19 . The method of  claim 17 , comprising:
 determining, by the computing system an amount of similarity between individual additional amino acid sequences and an amino acid sequence of an antibody produced in relation to expression of a germline gene;   determining, by the computing system, a length of respective complementarity-determining region (CDR) H3 regions of the individual additional amino acid sequences; and   evaluating, by the computing system, the additional amino acid sequences based on respective amounts of similarity and respective lengths of the CDR H3 regions of the additional amino acid sequences.   
     
     
         20 . The method of  claim 19 , comprising evaluating, by the computing system, the additional amino acid sequences based on a measure of immunogenicity of the additional amino acid sequences. 
     
     
         21 . The method of  claim 20 , wherein the measure of immunogenicity corresponds to a measure of major histocompatibility complex (MHC) Class II binding. 
     
     
         22 . The method of  claim 17 , wherein the structured amino acid sequences are represented in a matrix that includes a first number of rows and a second number of columns, individual rows of the first number of rows corresponding to a position of a sequence, and individual columns of the second number of columns corresponding to individual amino acids. 
     
     
         23 . The method of  claim 17 , wherein one or more characteristics of proteins corresponding to the additional amino acid sequences have at least a threshold similarity to one or more characteristics of the proteins included in the training dataset. 
     
     
         24 . The method of  claim 23 , wherein the one or more characteristics include at least one of structural position features, tertiary structure features, or biophysical properties. 
     
     
         25 . The method of  claim 17 , wherein the proteins include antibodies, affibodies, affilins, affimers, affitins, alphabodies, anticalins, avimers, monobodies, designed ankyrin repeat proteins (DARPins), nanoCLAMP (clostridal antibody mimetic proteins), antibody fragments, or combinations thereof. 
     
     
         26 . The method of  claim 17 , wherein the training dataset includes first data indicating a first amino acid sequence of a first antigen and second data indicating binding interactions between individual antibodies of a first plurality of antibodies and one or more antigens. 
     
     
         27 . The method of  claim 26 , comprising:
 determining, by the computing system, a second amino acid sequence of a second antigen of the one or more antigens that has at least a threshold amount of identity with respect to at least a portion of the first amino acid sequence of the first antigen; and   determining, by the computing system, a plurality of antibodies included in the second data that have at least a first threshold amount of binding interaction with the second antigen.   
     
     
         28 . The method of  claim 27 , wherein the additional amino acid sequences correspond to additional antibodies having at least a threshold probability of having at least a second threshold amount of binding interaction with the first antigen. 
     
     
         29 . The method of  claim 28 , wherein at least one of the first threshold amount of binding interaction or the second threshold amount of binding interaction correspond to binding interactions comprising at least one of a binding affinity or a binding avidity. 
     
     
         30 . The method of  claim 29 , wherein the binding interactions indicate an amino acid sequence of a binding region of an antibody and an additional amino acid sequence of an epitope region of an antigen, and the binding region binds to the epitope region. 
     
     
         31 . The method of  claim 30 , wherein the binding interactions indicate couplings between amino acids included in the binding region and additional amino acids included in the epitope region. 
     
     
         32 . The method of  claim 29 , wherein the binding interactions correspond to equilibrium constants between individual antibodies and one or more antigens. 
     
     
         33 . The method of  claim 17 , comprising wherein the metrics of the additional amino acid sequences include at least one of a number of hydrophobic amino acids included in individual amino acid sequences of the additional amino acid sequences, a number of positively charged amino acids included in individual amino acid sequences of the additional amino acid sequences, a number of negatively charged amino acids included in individual amino acid sequences of the additional amino acid sequences, a number of uncharged amino acids included in individual amino acid sequences of the additional amino acid sequences, a level of expression of individual proteins that correspond to the additional amino acid sequences, a melting temperature of individual proteins that correspond to the additional amino acid sequences, or a level of self-aggregation of individual proteins that correspond to the additional amino acid sequences. 
     
     
         34 . A system comprising:
 one or more hardware processors; and   one or more non-transitory computer readable media storing computer-executable instructions that, when executed by the one or more hardware processors, cause the one or more processor to perform operations comprising:   obtaining a training dataset including amino acid sequences of proteins;   generating structured amino acid sequences based on the training dataset;   generating a model to produce additional amino acid sequences that correspond to the amino acid sequences included in the training dataset using the structured amino acid sequences and a generative adversarial network;   generating the additional amino acid sequences using the model and an input vector; and   evaluating the additional amino acid sequences according to one or more criteria to determine metrics for the additional amino acid sequences.   
     
     
         35 . The system of  claim 34 , wherein the generative adversarial network includes a Wasserstein generative adversarial network. 
     
     
         36 . The system of  claim 34 , wherein the one or more non-transitory computer readable media store additional computer-executable instructions that, when executed by the one or more hardware processors, cause the one or more processor to perform additional operations comprising:
 analyzing the additional amino acid sequences with respect to the amino acid sequences included in the training dataset to evaluate a loss function of the generative adversarial network; and   modifying one or more components of the generative adversarial network to minimize the loss function.

Join the waitlist — get patent alerts

Track US2023178186A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.