Quality score compression for improving downstream genotyping accuracy
Abstract
This disclosure provides for a highly-efficient and scalable compression tool that compresses quality scores, preferably by capitalizing on sequence redundancy. In one embodiment, compression is achieved by smoothing a large fraction of quality score values based on k-mer neighborhood of their corresponding positions in read sequences. The approach exploits the intuition that any divergent base in a k-mer likely corresponds to either a single-nucleotide polymorphism (SNP) or sequencing error; thus, a preferred approach is to only preserve quality scores for probable variant locations and compress quality scores of concordant bases, preferably by resetting them to a default value. By viewing individual read datasets through the lens of k-mer frequencies in a corpus of reads, the approach herein ensures that compression “lossiness” does not affect accuracy in a deleterious way.
Claims
exact text as granted — not AI-modified1 . A genotyping analysis pipeline, comprising: one or more computing systems comprising computer hardware, and computer software stored in memory and executing on the computer hardware, the computer software configured as a set of computer program code, the computer program code comprising: first program code configured as a read data set acquisition tool; second program code configured as a genotyping tool; and third program code configured as a quality score compression tool and positioned between the read data set acquisition tool and the genotyping tool, the quality score compression tool configured to: receive, as a dictionary, a set of data comprising commonly-occurring k-mers extracted from the read dataset by the read data set acquisition tool; compress quality scores in a given read dataset by identifying k-mers from each read within a given mismatch distance from other k-mers in the dictionary to generate compressed quality scores; and provide the compressed quality scores for further processing by the genotyping tool.
Join the waitlist — get patent alerts
Track US2024004838A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.