US2025210146A1PendingUtilityA1

Sequence optimization

Assignee: UAB BIOMATTER DESIGNSPriority: Mar 17, 2022Filed: Mar 16, 2023Published: Jun 26, 2025
Est. expiryMar 17, 2042(~15.6 yrs left)· nominal 20-yr term from priority
G16B 20/50G06N 20/00G16B 40/00G16B 40/30
42
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Provided herein is an apparatus for generating an optimized protein or nucleic acid sequence from a target protein or nucleic acid sequence, wherein the optimized protein or nucleic acid sequence has an improved function over the target sequence. The apparatus comprises at least one processor and at least one memory including a computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause the apparatus to at least: operate a machine learning model configured to receive the target protein sequence or a nucleic acid sequence, and to generate therefrom one or more corresponding optimized sequences, wherein the machine learning model has been trained on a set of training data comprising native or engineered protein or nucleic acid sequences, and additionally, at least a subset of the sequences comprising one or more masked portions and/or at least a subset of the sequences comprising one or more mutations introduced therein.

Claims

exact text as granted — not AI-modified
1 . An apparatus for generating an optimized protein or nucleic acid sequence from a target protein or nucleic acid sequence, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause the apparatus to at least:
 operate a machine learning model configured to receive an input target protein sequence or a nucleic acid sequence, and to generate therefrom one or more corresponding optimized sequences having an improved function over the target sequence, each optimized sequence having one or more mutations with respect to the target sequence, wherein the machine learning model has been trained on a set of training data comprising native or engineered protein or nucleic acid sequences, and additionally, at least a subset of the sequences comprising one or more masked portions and/or at least a subset of the sequences comprising one or more mutations introduced therein; and   generate output data relating to the one or more optimized sequences.   
     
     
         2 . The apparatus of  claim 1 , wherein the apparatus is configured to perform the following steps after the target sequence has been received:
 i) determine likelihoods of substitutions with a pre-defined set of proteinogenic amino acids at one or more positions within the target sequence when the target sequence is a protein sequence, or with a pre-defined set of bases at one or more positions within the target sequence when the target sequence is a nucleic acid sequence;   ii) calculate scores for the likelihoods of the substitutions at the one or more positions based on a scoring function;   iii) select one or more substitutions based on the calculated scores to generate one or more new mutated target sequences each comprising one or more substitutions; and   iv) repeat steps i) to iii) until no further substitutions with an improved score are yielded at each corresponding position; or repeat steps i) to iii) until the apparatus has performed for a selected amount of time.   
     
     
         3 . The apparatus of  claim 2 , wherein the apparatus is configured to determine the likelihoods of substitutions with the pre-defined set of proteinogenic amino acids at two or more respective positions within the target sequence when the target sequence is a protein sequence, or with the pre-defined set of bases at two or more respective positions within the target sequence when the target sequence is a nucleic acid sequence, wherein the likelihoods are based on the combined substitutions at the two or more positions. 
     
     
         4 . The apparatus of  claim 2 , wherein the scoring function is equal to the likelihood. 
     
     
         5 . The apparatus of  claim 2 , wherein the scoring function further comprises an input that is independent of likelihood. 
     
     
         6 . The apparatus of  claim 2 , wherein the apparatus is configured to introduce one or more mutations into the received target sequence and/or into the one or more new target sequences obtained in step iii). 
     
     
         7 . The apparatus of  claim 6 , wherein the introduction of one or more mutations is based on probability derived from a position-specific scoring matrix. 
     
     
         8 . The apparatus of  claim 1 , wherein the input target sequence comprises a protein sequence. 
     
     
         9 . The apparatus of  claim 8 , wherein the protein is an enzyme. 
     
     
         10 . The apparatus of  claim 8 , wherein the one or more optimized sequences exhibit one or more improved functions relative to the target sequence selected from: K cat /K m , K cat , K m , thermostability, pH stability, specificity, ionic strength stability, solvent stability, resistance to one or more inhibitors, resistance to a chaotropic agent, resistance to an ionic detergent, shelf-life, expressibility in recombinant systems, and adsorption to a plastic. 
     
     
         11 . (canceled) 
     
     
         12 . The apparatus of  claim 1 , wherein the machine learning model is trained on 20% to 100% cluster representatives of the native sequences. 
     
     
         13 . The apparatus of  claim 12 , wherein the machine learning model is trained on about 50% cluster representatives of the native sequences. 
     
     
         14 . The apparatus of  claim 1 , wherein the machine learning model is trained on distances between amino acid residues within structural representations of the native or engineered protein sequences. 
     
     
         15 . The apparatus of  claim 1 , wherein the machine learning model is constructed from multi-attention blocks combined with convolution layers. 
     
     
         16 . The apparatus of  claim 1 , wherein the machine learning model has been further trained on a set of homologous sequences that have greater than about 35% sequence identity to the target sequence. 
     
     
         17 . The apparatus of  claim 1 , wherein the apparatus further comprises a user interface which is configured to receive the input target sequence and to supply the input target sequence to the machine learning model, and wherein the user interface is configured to present the one or more optimized sequences to a user. 
     
     
         18 . (canceled) 
     
     
         19 . A computer-implemented method of generating one or more optimized protein sequences or nucleic acid sequences from a target protein sequence or nucleic acid sequence, using a trained machine learning model;
 wherein the machine learning model has been trained based on:   a) a set of training data comprising native or engineered protein or nucleic acid sequences; and   b) at least a subset of the native or engineered protein or nucleic acid sequences comprising one or more masked portions and/or at least a subset of the native or engineered protein or nucleic acid sequences comprising one or mutations introduced therein; and   wherein the machine learning model is configured use the training data in order to become a trained machine learning model; and   wherein the method comprises:   i) receiving as input the target sequence;   ii) causing the trained machine learning model to evaluate the inputted target sequence; and   iii) based on the evaluation, generating one or more optimized sequences corresponding to the target sequence, each optimized sequence having an improved function over the target sequence, and one or more mutations with respect to the target sequence; and   iv) outputting data relating to the one or more generated optimized sequences.   
     
     
         20 . The method of  claim 19 , wherein in step (ii), evaluation of the target sequence comprises:
 v) determining likelihoods of substitutions with a pre-defined set of proteinogenic amino acids at one or more positions within the target sequence when the target sequence is a protein sequence, or with a pre-defined set of bases at one or more positions within the target sequence when the target sequence is a nucleic acid sequence;   vi) calculating scores for the likelihoods based on a scoring function;   vii) selecting one or more substitutions based on the calculated scores to generate one or more new mutated target sequences each comprising one or more substitutions; and   viii) repeating steps v) to vii) until no further mutations with an improved score are yielded at each corresponding position; or repeating steps v) to vii) until the apparatus has performed for a selected period of time.   
     
     
         21 - 25 . (canceled) 
     
     
         26 . The method of  claim 19 , wherein in step b), the introduced mutations are random. 
     
     
         27 - 35 . (canceled) 
     
     
         36 . The method of  claim 19 , wherein the machine learning model has been trained based on a loss function related to predicting the original amino acid or original base at any given position of the protein sequence or nucleic acid sequence, predicting a mutated amino acid or a mutated base at any given position of the protein sequence or nucleic acid sequence, and if the training data comprises native protein sequences, the distance between Ca atoms of the protein sequences. 
     
     
         37 - 38 . (canceled)

Join the waitlist — get patent alerts

Track US2025210146A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.