Estimation of circulating tumor fraction using off-target reads of targeted-panel sequencing
Abstract
Methods, systems, and software are provided for estimating a circulating tumor fraction for a test subject. Sequence reads are obtained from a panel-enriched sequencing reaction, including sequences for a first plurality of cfDNA fragments corresponding to probe sequences and a second plurality of cfDNA fragments not corresponding to probe sequences. Bin-level coverage ratios are determined from the sequences. Segments are formed by grouping adjacent bins based on similar coverage ratios and segment-level coverage ratios are determined based on bin-level coverage ratios for bins in the segment. For each simulated circulating tumor fraction in a plurality of circulating tumor fractions, segments are fitted to an integer copy state by identifying the integer copy state that best matches the segment-level coverage ratio. The circulating tumor fraction for the test subject is determined using error optimization between segment-level coverage ratios and integer copy states across the simulated circulated tumor fractions.
Claims
exact text as granted — not AI-modified1 - 30 . (canceled)
31 . A method of estimating a circulating tumor fraction for a test subject comprising:
at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors: A) obtaining, from a panel-enriched sequencing reaction, a first plurality of at least 10,000 nucleic acid sequences, wherein the first plurality of at least 10,000 nucleic acid sequences comprises:
(i) a corresponding sequence for each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from the test subject, wherein each respective cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences used to enrich cell-free DNA fragments in the liquid biopsy sample in the panel-enriched sequencing reaction; and
(ii) a corresponding sequence for each cell-free DNA fragment in a second plurality of cell-free DNA fragments obtained from the liquid biopsy sample, wherein each respective cell-free DNA fragment in the second plurality of cell-free DNA fragments does not correspond to any probe sequence in the plurality of probe sequences;
B) determining a plurality of at least 500 bin-level coverage values using the plurality of at least 10,000 nucleic acid sequences, each respective bin-level coverage value in the plurality of bin-level coverage values corresponding to a respective bin in a plurality of at least 500 bins, wherein:
each respective bin in the plurality of bins represents a corresponding region of the genome for the species of the test subject,
the plurality of bins collectively covers at least 25 Mb of the genome for the species of the test subject, and
each respective bin-level coverage value in the plurality of bin-level coverage values is determined from a comparison of (i) a number of nucleic acid sequences in the first plurality of nucleic acid sequences that map to the corresponding bin and (ii) a number of nucleic acid sequences from one or more reference samples that map to the corresponding bin;
C) determining a plurality of segment-level coverage values by:
forming, using the at least 500 bin-level coverage values, a plurality of segments by grouping respective subsets of adjacent bins in the plurality of bins based on a similarity between the respective bin-level coverage values of the subset of adjacent bins, and
determining, for each respective segment in the plurality of segments, a segment-level coverage value based on the corresponding bin-level coverage values for each bin in the respective segment; and
D) estimating the circulating tumor fraction for the test subject to be a first simulated circulating tumor fraction based on a measure of fit between corresponding values in (i) the plurality of segment-level coverage values and (ii) a first set of integer copy states that includes a respective integer copy state for each respective segment in the plurality of segments that is determined by fitting the respective segment, given the first simulated circulating tumor fraction, to a respective integer copy state, in a plurality of integer copy states, that best matches the segment-level coverage value.
32 . The method of claim 31 , wherein the plurality of probe sequences collectively map to at least 25 different genes in a human reference genome.
33 . The method of claim 31 , wherein the plurality of integer copy states comprises a 1-copy state, a 2-copy state, a 3-copy state, and a 4-copy state.
34 . The method of claim 31 , wherein the estimating D) includes:
determining, for each respective integer copy state in the plurality of integer copy states, a corresponding expected coverage value; comparing, for each respective segment in the plurality of segments, the corresponding segment-level coverage value to the expected coverage value for each respective integer copy state in the plurality of integer copy states; and assigning, for each respective segment in the plurality of segments, a corresponding integer copy state based on the comparison.
35 . The method of claim 31 , wherein the first simulated circulating tumor fraction is identified from among a plurality of simulated circulating tumor fractions that comprises at least 25 simulated circulating tumor fractions by:
determining, for each respective simulated circulating tumor fraction in the plurality of simulated circulating tumor fractions, a corresponding measure of fit; and selecting the first simulated circulating tumor fraction from among the plurality of simulated circulating tumor fractions on the basis that the first simulated circulating tumor fraction has the best measure of fit in the plurality of simulated circulating tumor fractions.
36 . The method of claim 35 , wherein the plurality of simulated circulating tumor fractions spans a range of at least from 5% to 25%.
37 . The method of claim 35 , wherein the plurality of simulated circulating tumor fractions spans a range of at least from 1% to 50%.
38 . The method of claim 37 , wherein the span between each consecutive pair of simulated tumor fractions is no more than 5%.
39 . The method of claim 35 , wherein the determining and selecting is performed using an expectation-maximization algorithm.
40 . The method of claim 31 , wherein the measure of fit is based on an aggregate of a respective difference, for each respective segment in the plurality of segments, between the respective segment-level coverage value and an expected coverage value for the corresponding copy state fit to the respective segment.
41 . The method of claim 31 , further comprising generating a report for the test subject comprising the circulating tumor fraction estimated for the test subject.
42 . The method of claim 41 , wherein the report further comprises a therapeutic recommendation for the test subject based on the circulating tumor fraction estimated for the test subject.
43 . The method of claim 31 , wherein the panel-enriched sequencing reaction is performed at a read depth of at least 1,000×.
44 . The method of claim 31 , wherein the panel-enriched sequencing reaction enriches for at least 50 genes.
45 . The method of claim 31 , wherein the panel-enriched sequencing reaction enriches for at least 10 genes selected from the group consisting of ALK, B2M, ERRFI1, IDH2, MSH6, PIK3R1, SPOP, FGFR2, BAP1, ESR1, JAK1, MTOR, PMS2, STK11, FGFR3, BRCA1, EZH2, JAK2, MYCN, PTCH1, TERT, NTRK1, BRCA2, FBXW7, JAK3, NF1, PTEN, TP53, RET, BTK, FGFR1, KDR, NF2, PTPN11, TSC1, ROS1, CCND1, FGFR4, KEAP1, NFE2L2, RAD51C, TSC2, BRAF, CCND2, FLT3, KIT, NOTCH1, RAF1, UGT1A1, AKT1, CCND3, FOXL2, KRAS, NPM1, RB1, VHL, AKT2, CDH1, GATA3, MAP2K1, NRAS, RHEB, CCNE1, APC, CDK4, GNA11, MAP2K2, PALB2, RHOA, CD274, AR, CDK6, GNAQ, MAPK1, PBRM1, RIT1, EGFR, ARAF, CDKN2A, GNAS, MLH1, PDCD1LG2, RNF43, ERBB2, ARID1A, CTNNB1, HNF1A, MPL, PDGFRA, SDHA, MET, ATM, DDR2, HRAS, MSH2, PDGFRB, SMAD4, MYC, ATR, DPYD, IDH1, MSH3, PIK3CA, SMO, and KMT2A.
46 . The method of claim 31 , wherein the panel-enriched sequencing reaction enriches for at least 10 genes selected from the group consisting of AKT1, ALK, APC, AR, ARAF, ARID1A, ATM, BRAF, BRCA1, BRCA2, CCND1, CCND2, CCNE1, CDH1, CDK4, CDK6, CDKN2A, CTNNB1, DDR2, EGFR, ERBB2, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, GATA3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KIT, KRAS, MAP2K1, MAP2K2, MAPK1, MAPK3, MET, MLH1, MPL, MTOR, MYC, NF1, NFE2L2, NOTCH1, NPM1, NRAS, NTRK1, NTRK3, PDGFRA, PIK3CA, PTEN, PTPN11, RAF1, RB1, RET, RHEB, RHOA, RIT1, ROS1, SMAD4, SMO, STK11, TERT, TP53, TSC1, and VHL.
47 . The method of claim 31 , wherein the panel-enriched sequencing reaction enriches for at least 10 genes selected from the group consisting of ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1 (FAM123B), APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, C11orf30 (EMSY), C17orf39 (GID4), CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274 (PD-L1), CD70, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, EZH2, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A, KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1 (MEK1), MAP2K2 (MEK2), MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYC, MYCL (MYCL1), MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NSD3 (WHSC1L1), NT5C2, NTRK1, NTRK2, NTRK3, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1 (PD-1), PDCD1LG2 (PD-L2), PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, ncRNA, Promoter, TGFBR2, TIPARP, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WT1, XPO1, XRCC2, ZNF217, and ZNF703.
48 . The method of claim 31 , wherein the plurality of bin-level coverage values comprises at least 1000 bin-level coverage values and the plurality of bins comprises at least 1000 bins.
49 . The method of claim 31 , wherein the plurality of nucleic acid sequences comprises at least 100,000 nucleic acid sequences and the plurality of at least 500 bin-level coverage values using the plurality of at least 100,000 nucleic acid sequences.
50 . A computer system comprising:
one or more processors; and a non-transitory computer-readable medium including computer-executable instructions that, when executed by the one or more processors, cause the processors to perform a method comprising: A) obtaining, from a panel-enriched sequencing reaction, a first plurality of at least 10,000 nucleic acid sequences, wherein the first plurality of at least 10,000 nucleic acid sequences comprises:
(i) a corresponding sequence for each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from the test subject, wherein each respective cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences used to enrich cell-free DNA fragments in the liquid biopsy sample in the panel-enriched sequencing reaction; and
(ii) a corresponding sequence for each cell-free DNA fragment in a second plurality of cell-free DNA fragments obtained from the liquid biopsy sample, wherein each respective cell-free DNA fragment in the second plurality of cell-free DNA fragments does not correspond to any probe sequence in the plurality of probe sequences;
B) determining a plurality of at least 500 bin-level coverage values using the plurality of at least 10,000 nucleic acid sequences, each respective bin-level coverage value in the plurality of bin-level coverage values corresponding to a respective bin in a plurality of at least 500 bins, wherein:
each respective bin in the plurality of bins represents a corresponding region of the genome for the species of the test subject,
the plurality of bins collectively covers at least 25 Mb of the genome for the species of the test subject, and
each respective bin-level coverage value in the plurality of bin-level coverage values is determined from a comparison of (i) a number of nucleic acid sequences in the plurality of nucleic acid sequences that map to the corresponding bin and (ii) a number of nucleic acid sequences from one or more reference samples that map to the corresponding bin;
C) determining a plurality of segment-level coverage values by:
forming, using the at least 500 bin-level coverage values, a plurality of segments by grouping respective subsets of adjacent bins in the plurality of bins based on a similarity between the respective coverage ratios of the subset of adjacent bins, and
determining, for each respective segment in the plurality of segments, a segment-level coverage value based on the corresponding bin-level coverage values for each bin in the respective segment; and
D) estimating the circulating tumor fraction for the test subject to be a first simulated circulating tumor fraction based on a measure of fit between corresponding values in (i) the plurality of segment-level coverage values and (ii) a first set of integer copy states that includes a respective integer copy state for each respective segment in the plurality of segments that is determined by fitting the respective segment, given the first simulated circulating tumor fraction, to a respective integer copy state, in a plurality of integer copy states, that best matches the segment-level coverage value.
51 . A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method comprising:
A) obtaining, from a panel-enriched sequencing reaction, a first plurality of at least 10,000 nucleic acid sequences, wherein the first plurality of at least 10,000 nucleic acid sequences comprises:
(i) a corresponding sequence for each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from the test subject, wherein each respective cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences used to enrich cell-free DNA fragments in the liquid biopsy sample in the panel-enriched sequencing reaction; and
(ii) a corresponding sequence for each cell-free DNA fragment in a second plurality of cell-free DNA fragments obtained from the liquid biopsy sample, wherein each respective cell-free DNA fragment in the second plurality of cell-free DNA fragments does not correspond to any probe sequence in the plurality of probe sequences;
B) determining a plurality of at least 500 bin-level coverage values using the plurality of at least 10,000 nucleic acid sequences, each respective bin-level coverage value in the plurality of bin-level coverage values corresponding to a respective bin in a plurality of at least 500 bins, wherein:
each respective bin in the plurality of bins represents a corresponding region of the genome for the species of the test subject,
the plurality of bins collectively covers at least 25 Mb of the genome for the species of the test subject, and
each respective bin-level coverage value in the plurality of bin-level coverage ratios is determined from a comparison of (i) a number of nucleic acid sequences in the plurality of nucleic acid sequences that map to the corresponding bin and (ii) a number of nucleic acid sequences from one or more reference samples that map to the corresponding bin;
C) determining a plurality of segment-level coverage values by:
forming, using the at least 500 bin-level coverage values, a plurality of segments by grouping respective subsets of adjacent bins in the plurality of bins based on a similarity between the respective coverage values of the subset of adjacent bins, and
determining, for each respective segment in the plurality of segments, a segment-level coverage value based on the corresponding bin-level coverage values for each bin in the respective segment; and
D) estimating the circulating tumor fraction for the test subject to be a first simulated circulating tumor fraction based on a measure of fit between corresponding values in (i) the plurality of segment-level coverage values and (ii) a first set of integer copy states that includes a respective integer copy state for each respective segment in the plurality of segments that is determined by fitting the respective segment, given the first simulated circulating tumor fraction, to a respective integer copy state, in a plurality of integer copy states, that best matches the segment-level coverage value.Join the waitlist — get patent alerts
Track US2022328133A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.