Neoantigen identification using hotspots
Abstract
A method for identifying neoantigens that are likely to be presented on a surface of tumor cells of a subject. Peptide sequences of tumor neoantigens are obtained by sequencing the tumor cells of the subject. The peptide sequence of each of the neoantigens is associated with one or more k-mer blocks of a plurality of k-mer blocks of the nucleotide sequencing data of the subject; The peptide sequences and the associated k-mer blocks are input into a machine-learned presentation model to generate presentation likelihoods for the tumor neoantigens, each presentation likelihood representing the likelihood that a neoantigen is presented by an MHC allele on the surfaces of the tumor cells of the subject. A subset of the neoantigens is selected based on the presentation likelihoods.
Claims
exact text as granted — not AI-modified1 . A method for identifying one or more antigens from one or more cells of a subject that are likely to be presented on a surface of the cells, the method comprising the steps of:
(a) obtaining data representing peptide sequences of each of a set of antigens, wherein said data further comprises a value indicating a likelihood of a presentation hotspot for one or more k-mer blocks associated with the peptide sequences; (b) determining, using a neural network model, a set of presentation likelihoods for the set of antigens, each presentation likelihood in the set representing the likelihood that a corresponding antigen is presented by one or more MHC alleles on the surface of the cells of the subject, the neural network model comprising:
(i) two or more layers comprising a first layer and a second layer, each layer comprising one or more nodes, wherein said nodes comprise a memory location for one or more input values;
(ii) a plurality of connections between nodes of said first layer and one or more nodes of said second layer;
(iii) optimized parameters stored in memory locations, wherein the optimized parameters transform input values of nodes of the first layer into input values for nodes of the second layer connected to the nodes of the first layer,
(iv) wherein the optimized parameters are generated using a training data set comprising:
(A) training peptide sequences or data derived from training peptide sequences;
(B) at least one MHC allele associated with the training peptide sequences;
(C) a value indicating a likelihood of a presentation hotspot for one or more k-mer blocks of a plurality of k-mer blocks associated with the training peptide sequences, and
(D) for each of one or more of the training peptide sequences, a label indicating whether the training peptide was presented by the at least one MHC allele; and
(c) wherein said determining comprises:
forward feeding the data representing peptide sequences of each of a set of antigens, using a computer processor, through nodes of the first layer and the second layer of the neural network model, said forward feeding comprising transforming the data as they are fed from nodes of the first layer to nodes of the second layer using the optimized parameters;
generating, using a computer processor, the set of presentation likelihoods for the set of antigens from the transformed data;
(d) selecting a subset of the set of antigens based on the set of presentation likelihoods to generate a set of selected antigens; and (e) returning the set of selected antigens.
2 . The method of claim 1 , wherein forward feeding the data representing peptide sequences further comprises inputting into the neural network model the value indicating the likelihood of a presentation hotspot for one or more k-mer blocks associated with the peptide sequences.
3 . The method of claim 1 , wherein a larger value of a parameter of the optimized parameters indicates a greater likelihood that a corresponding k-mer block gives rise to a presented peptide.
4 . The method of claim 1 , wherein a smaller value of a parameter of the optimized parameters indicates a smaller likelihood that a corresponding k-mer block gives rise to a presented peptide.
5 . The method of claim 1 , wherein at least one of the k-mer blocks corresponds to a proteomic location.
6 . The method of claim 5 , wherein the proteomic location comprises a block of n adjacent peptides, wherein n represents a hyperparameter of the neural network model.
7 . The method of claim 1 , wherein the peptide sequences or training peptide sequences comprise sequences having lengths between 8-15 amino acids.
8 . The method of claim 1 , wherein generating the set of presentation likelihoods for the set of antigens comprises:
generating a dependency score for each of the one or more class I MHC alleles, the dependency scores indicating whether the class I MHC alleles will present the antigen based on the particular amino acids at the particular positions of the peptide sequence.
9 . The method of claim 8 , wherein generating the set of presentation likelihoods for the set of antigens further comprises:
transforming the dependency scores to generate a corresponding per-allele likelihood for each class I MHC allele indicating a likelihood that the corresponding class I MHC allele will present the corresponding antigen; and combining the per-allele likelihoods to generate the presentation likelihood of the antigen.
10 . The method of claim 9 , wherein the transforming the dependency scores models the presentation of the antigen as mutually exclusive across the one or more class I MHC alleles.
11 . The method of claim 10 , generating the set of presentation likelihoods for the set of antigens further comprises:
transforming a combination of the dependency scores to generate the presentation likelihood, wherein transforming the combination of the dependency scores models the presentation of the antigen as interfering between the one or more class I MHC alleles.
12 . The method of claim 8 , wherein the set of presentation likelihoods are further identified by at least one or more allele noninteracting features, and further comprising:
applying the neural network model to the allele noninteracting features to generate a dependency score for the allele noninteracting features indicating whether the peptide sequence of the corresponding antigen will be presented based on the allele noninteracting features.
13 . The method of claim 12 , further comprising:
combining the dependency score for each class I MHC allele in the one or more class I MHC alleles with the dependency score for the allele noninteracting feature; and transforming the combined dependency scores for each class I MHC allele to generate a per-allele likelihood for each class I MHC allele indicating a likelihood that the corresponding class I MHC allele will present the corresponding antigen; and combining the per-allele likelihoods to generate the presentation likelihood.
14 . The method of claim 13 , further comprising:
transforming a combination of the dependency scores for each of the class I MHC alleles and the dependency score for the allele noninteracting features to generate the presentation likelihood.
15 . The method of claim 1 , wherein the one or more class I MHC alleles include two or more class I MHC alleles.
16 . The method of claim 1 , wherein the plurality of samples comprise at least one of:
(a) one or more cell lines engineered to express a single MHC class I allele; (b) one or more cell lines engineered to express a plurality of MHC class I alleles; (c) one or more human cell lines obtained or derived from a plurality of patients; (d) fresh or frozen tumor samples obtained from a plurality of patients; and (e) fresh or frozen tissue samples obtained from a plurality of patients.
17 . The method of claim 1 , wherein the set of presentation likelihoods are further identified by at least expression levels of the one or more class I MHC alleles in the subject, as measured by RNA-seq or mass spectrometry.
18 . The method of claim 1 , wherein the set of numerical likelihoods are further identified by features comprising at least one of:
(a) the C-terminal sequences flanking the antigen encoded peptide sequence within its source protein sequence; and (b) the N-terminal sequences flanking the antigen encoded peptide sequence within its source protein sequence.
19 . The method of claim 1 , further comprising generating an output for constructing a personalized cancer vaccine from the set of selected antigens.
20 . A computer system comprising:
a computer processor; a memory storing computer program instructions that when executed by the computer processor cause the computer processor to:
(a) obtain data representing peptide sequences of each of a set of antigens wherein said data further comprises a value indicating a likelihood of a presentation hotspot for one or more k-mer blocks associated with the peptide sequences;
(b) determine, using a neural network model, a set of presentation likelihoods for the set of neoantigens, each presentation likelihood in the set representing the likelihood that a corresponding neoantigen is presented by one or more MHC alleles on the surface of the cells of the subject, the neural network model comprising:
(i) two or more layers comprising a first layer and a second layer, each layer comprising one or more nodes, wherein said nodes comprise a memory location for one or more input values;
(ii) a plurality of connections between nodes of said first layer and one or more nodes of said second layer,
(iii) optimized parameters stored in memory locations, wherein the optimized parameters transform input values of nodes of the first layer into input values for nodes of the second layer connected to the nodes of the first layer,
(iv) wherein the optimized parameters are generated using a training data set comprising:
(A) training peptide sequences or data derived from training peptide sequences;
(B) at least one MHC allele associated with the training peptide sequences;
(C) a value indicating likelihood of a presentation hotspot for one or more k-mer blocks of a plurality of k-mer blocks associated with the peptide sequences; and
(D) for each of one or more of the training peptide sequences, a label indicating whether the training peptide was presented by the at least one MHC allele,
(c) wherein said determination of the set of presentation likelihoods comprises:
forward feeding the data representing peptide sequences of each of a set of antigens, using a computer processor, through nodes of the first layer and the second layer of the neural network model, said forward feeding comprising transforming the data as they are fed from nodes of the first layer to nodes of the second layer using the optimized parameters;
generating, using a computer processor, the set of presentation likelihoods for the set of antigens from the transformed data;
(d) select a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; and
(e) return the set of selected neoantigens.Join the waitlist — get patent alerts
Track US2022148681A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.