US2024004838A1PendingUtilityA1

Quality score compression for improving downstream genotyping accuracy

Assignee: LEIGHTON BONNIE BERGERPriority: Apr 26, 2014Filed: Sep 19, 2023Published: Jan 4, 2024
Est. expiryApr 26, 2034(~7.8 yrs left)· nominal 20-yr term from priority
G06F 16/1744G16B 30/00G16B 50/00G16C 99/00G16B 50/50G16C 10/00G16B 30/10
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure provides for a highly-efficient and scalable compression tool that compresses quality scores, preferably by capitalizing on sequence redundancy. In one embodiment, compression is achieved by smoothing a large fraction of quality score values based on k-mer neighborhood of their corresponding positions in read sequences. The approach exploits the intuition that any divergent base in a k-mer likely corresponds to either a single-nucleotide polymorphism (SNP) or sequencing error; thus, a preferred approach is to only preserve quality scores for probable variant locations and compress quality scores of concordant bases, preferably by resetting them to a default value. By viewing individual read datasets through the lens of k-mer frequencies in a corpus of reads, the approach herein ensures that compression “lossiness” does not affect accuracy in a deleterious way.

Claims

exact text as granted — not AI-modified
1 . A genotyping analysis pipeline, comprising:
 one or more computing systems comprising computer hardware, and computer software stored in memory and executing on the computer hardware, the computer software configured as a set of computer program code, the computer program code comprising:
 first program code configured as a read data set acquisition tool; 
 second program code configured as a genotyping tool; and 
 third program code configured as a quality score compression tool and positioned between the read data set acquisition tool and the genotyping tool, the quality score compression tool configured to:
 receive, as a dictionary, a set of data comprising commonly-occurring k-mers extracted from the read dataset by the read data set acquisition tool; 
 compress quality scores in a given read dataset by identifying k-mers from each read within a given mismatch distance from other k-mers in the dictionary to generate compressed quality scores; and 
 provide the compressed quality scores for further processing by the genotyping tool.

Join the waitlist — get patent alerts

Track US2024004838A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.