US2022228208A1PendingUtilityA1
Systems and Methods for Sequencing T Cell Receptors and Uses Thereof
Est. expiryDec 9, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G16B 50/30G16B 40/00G16B 30/00G16B 25/20G16B 20/30C12Q 1/6869A61K 40/42A61K 40/32A61K 40/11A61K 2239/50G16B 20/20C12Q 2600/106C07K 16/2818A61K 35/17G16B 20/00C12Q 1/6886C12Q 2600/158C12Q 1/6883A61K 2039/505C12Q 2600/118C07K 16/2878
72
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Disclosed herein are methods and systems that can reconstruct, extract, and/or analyze TCR sequences using short reads. The methods and systems can be applied to both single cell and bulk sequencing data.
Claims
exact text as granted — not AI-modified1 . A method for identifying a T-cell receptor (TCR), comprising:
a) sequencing, using a high-throughput sequencing device, reads of RNA obtained from a T-cell and storing, in a system memory of a computing device, a sequence data structure comprising the reads and a reference data structure comprising a reference sequence that does not contain a TCR gene sequence; b) aligning, by the computing device, the reads in the sequence data structure with the reference sequence in the reference data structure, thereby generating, in the sequence data structure, mapped reads and unmapped reads; c) discarding, by the computing device, the mapped reads from the sequence data structure; and d) identifying, by the computing device, a first mapped read in the sequence data structure that aligns to a TCR V gene reference sequence as a candidate TCR V gene sequence and a second mapped read in the sequence data structure that aligns to a TCR J gene reference sequence as a candidate TCR J gene sequence, wherein the candidate TCR V gene sequence combined with the candidate TCR J gene sequence comprise a TCR sequence.
2 . The method of claim 1 , wherein the reads comprise short reads of less than about 100 base pairs.
3 . The method of claim 1 , further comprising:
assembling, by the computing device, the unmapped short reads into one or more long reads by aligning the unmapped short reads in the sequence data structure to one or more reference TCR sequences from a reference database of TCR sequences; and translating, by the computing device, the one or more long reads into corresponding amino acid sequences.
4 . The method of claim 3 , wherein identifying, by the computing device, a first mapped read in the sequence data structure that aligns to a TCR V gene reference sequence as a candidate TCR V gene sequence and a second mapped read in the sequence data structure that aligns to a TCR J gene reference sequence as a candidate TCR J gene sequence comprises:
fractioning, by the computing device, TCR V region and TCR J region amino acid reference sequences, from the reference database of TCR sequences, into k-strings of about six amino acids; aligning, by the computing device, the k-strings with the corresponding amino acid sequences; detecting, by the computing device, one or more conserved TCR CDR3 residues in the k-strings that map to the corresponding amino acid sequences; scoring, by the computing device, based on the one or more conserved TCR CDR3 residues that map to the corresponding amino acid sequences, a level of conservation for each of the corresponding amino acid sequences; selecting, by the computing device, one or more of the corresponding amino acid sequences, wherein the level of conservation for the one or more corresponding amino acid sequences is above a threshold conservation score; and detecting, by the computing device, a candidate CDR3 region amino acid sequence in the selected corresponding amino acid sequences.
5 . The method of claim 4 , further comprising:
identifying, by the computing device, a nucleic acid sequence of the candidate CDR3 region amino acid sequence in the one or more long reads as a candidate CDR3 region nucleic acid sequence.
6 . The method of claim 5 , further comprising:
aligning, by the computing device, a nucleic acid sequence of the one or more long reads, that is upstream of the candidate CDR3 region nucleic acid sequence with one or more TCR V gene reference sequences from the reference database of TCR sequences; scoring, by the computing device, a degree of the alignment of the nucleic acid sequence of the one or more long reads that is upstream of the candidate CDR3 region nucleic acid sequence with the one or more TCR V gene reference sequences from the reference database of TCR sequences; and identifying, by the computing device, at least one portion of the one or more long reads as comprising a candidate TCR V gene sequence, wherein the scored degree of alignment for the at least one portion of the one or more long reads that is upstream of the candidate CDR3 region nucleic acid sequence is above a threshold alignment score.
7 . The method of claim 6 , further comprising:
aligning, by the computing device, a nucleic acid sequence of the one or more long reads that is downstream of the candidate CDR3 region nucleic acid sequence with one or more TCR J gene reference sequences from the reference database of TCR sequences; scoring a degree of the alignment of the nucleic acid sequence of the one or more long reads that is downstream of the candidate CDR3 region nucleic acid sequence with the one or more TCR J gene reference sequences from the reference database of TCR sequences; and identifying at least one portion of the one or more long reads as comprising a candidate TCR J gene sequence, wherein the scored degree of alignment for the at least one portion of the one or more long reads, in the sequence data structure in the system memory, that is downstream of the candidate CDR3 region nucleic acid sequence is above the threshold alignment score.
8 . The method of claim 1 , wherein discarding the mapped reads from the sequence data structure further comprises discarding unmapped reads that are less than about 35 base pairs.
9 . The method of claim 1 further comprising, prior to sequencing the reads of RNA obtained from the T cell, administering an immunotherapy to a subject from which the T cell is obtained.
10 . The method of claim 9 , wherein the immunotherapy comprises a monotherapy or a combination therapy.
11 . The method of claim 10 , wherein the combination therapy comprises a costimulatory agonist and a coinhibitory antagonist.
12 . The method of claim 1 , further comprising:
performing steps a-d for a first plurality of T cells of a subject, wherein the T cells are collected prior to administration of a treatment; determining a number of occurrences of unique TCR sequences present in the first plurality of T cells; administering the treatment to the subject; performing steps a-d for a second plurality of T cells of the subject, wherein the T cells are collected after the administration of the treatment; determining a number of occurrences of unique TCR sequences present in the second plurality of T cells; and determining, based on the number of occurrences of unique TCR sequences present in the first plurality of T cells being less than the number of occurrences of unique TCR sequences present in the second plurality of T cells, one or more unique TCR sequences that experienced clonal expansion.
13 . The method of claim 12 , further comprising determining a T cell clonal expansion signature based on the one or more unique TCR sequences that experienced clonal expansion.
14 . The method of claim 13 , further comprising:
querying a database of T cell clonal expansion signatures and corresponding treatment responses using the T cell clonal expansion signature; and determining, based on the query, the subject's likelihood of responding to the treatment.
15 . The method of claim 13 , further comprising:
determining the subject's response to the treatment; storing the T cell clonal expansion signature in a database; and associating the subject's response to the treatment with the T cell clonal expansion signature in the database.
16 . The method of claim 1 , further comprising:
determining that the TCR sequence is present in a T cell clone that expands in response to a treatment; producing one or more T cells containing the TCR sequence; administering the one or more T cells to a subject; and administering the treatment to the subject.
17 . A method for identifying a B-cell receptor (BCR), comprising:
a) sequencing, using a high-throughput sequencing device, reads of RNA obtained from a B-cell and storing, in a system memory of a computing device, a sequence data structure comprising the reads and a reference data structure comprising a reference sequence that does not contain a BCR gene sequence; b) aligning, by the computing device, the reads in the sequence data structure with the reference sequence in the reference data structure, thereby generating, in the sequence data structure, mapped reads and unmapped reads; c) discarding, by the computing device, the mapped reads from the sequence data structure; and d) identifying, by the computing device, a first mapped read in the sequence data structure that aligns to a BCR V gene reference sequence as a candidate BCR V gene sequence and a second mapped read in the sequence data structure that aligns to a BCR J gene reference sequence as a candidate BCR J gene sequence, wherein the candidate BCR V gene sequence combined with the candidate BCR J gene sequence comprise a BCR sequence.
18 . The method of claim 17 further comprising, prior to sequencing the reads of RNA obtained from the B cell, administering an immunotherapy to a subject from which the B cell is obtained.
19 . The method of claim 18 , wherein the immunotherapy comprises a monotherapy or a combination therapy.
20 . The method of claim 19 , wherein the combination therapy comprises a costimulatory agonist and a coinhibitory antagonist.Join the waitlist — get patent alerts
Track US2022228208A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.