US2017316154A1PendingUtilityA1

Parallel-processing systems and methods for highly scalable analysis of biological sequence data

Assignee: RES INST AT NATIONWIDE CHILDREN'S HOSPITALPriority: Nov 21, 2014Filed: Nov 20, 2015Published: Nov 2, 2017
Est. expiryNov 21, 2034(~8.3 yrs left)· nominal 20-yr term from priority
G06F 19/22G06F 19/28G06F 3/0641G06F 3/0608G06F 3/067G16B 50/30G16B 30/10G16B 50/00G16B 30/20G16B 30/00
39
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An apparatus includes a memory configured to store a sequence that includes an estimation of a biological sequence. The sequence includes a set of elements. The apparatus also includes an assignment module implemented in a hardware processor. The assignment module is configured to receive the sequence from the memory, and assign each element to at least one segment from a set of segments, including, when an element maps to at least a first segment and a second segment, assigning the element set of segments specific to that hardware processor, and substantially simultaneous with the remaining hardware processors, remove at least a portion of duplicate elements in that segment to generate a deduplicated segment. Reorder the elements in the deduplicated segment to generate a realigned segment that has a reduced likelihood for alignment errors

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An apparatus, comprising:
 a memory configured to store a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements; and   a plurality of hardware processors operatively coupled to the memory, each hardware processor from the plurality of hardware processors configured to implement a segment processing module,   an assignment module implemented in a hardware processor from the plurality of hardware processors, the assignment module configured to:
 receive the sequence from the memory, and 
 assign each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to at least a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment, 
   the segment processing module for each hardware processor from the plurality of hardware processors operatively coupled to the assignment module, the segment processing module for each hardware processor from the plurality of hardware processors configured to, for each segment from a set of segments specific to that hardware processor and from the plurality of segments, and substantially simultaneous with the remaining hardware processors from the plurality of hardware processors:
 remove at least a portion of duplicate elements in that segment from that set of segments to generate a deduplicated segment; and 
 reorder the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment. 
   
     
     
         2 . The apparatus of  claim 1 , wherein the sequence is a target sequence, the assignment module further configured to receive the target sequence by:
 receiving a first sequence, the first sequence including a forward estimation of the biological sequence;   receiving a second sequence, the second sequence including a reverse estimation of the biological sequence;   generating a paired sequence based on the first sequence and the second sequence; and   aligning the paired sequence with a reference sequence to generate the target sequence.   
     
     
         3 . The apparatus of  claim 1 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format. 
     
     
         4 . The apparatus of  claim 1 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence. 
     
     
         5 . A method, comprising:
 receiving a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements;   assigning each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to both a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment; and   for each segment from the plurality of segments:
 removing at least a portion of duplicate elements in the segment to generate a deduplicated segment; 
 reordering the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment; and 
 transmitting the realigned segment to one or more of a storage module and a genotyping module. 
   
     
     
         6 . The method of  claim 5 , wherein the sequence is a target sequence, the receiving the target sequence including:
 receiving a first sequence, the first sequence including a forward estimation of the biological sequence;   receiving a second sequence, the second sequence including a reverse estimation of the biological sequence;   generating a paired sequence based on the first sequence and the second sequence; and   aligning the paired sequence with a reference sequence to generate the target sequence.   
     
     
         7 . The method of  claim 5 , further comprising:
 prior to the assigning, splitting the sequence into a plurality of subsequences, the assigning including assigning each subsequence from the plurality of subsequences to at least one segment from the plurality of segments; and   for each segment from the plurality of segments, subsequent to the assigning and prior to the removing, combining subsequences within the segment.   
     
     
         8 . The method of  claim 5 , further comprising, when an element from the plurality of elements maps to the first segment and the second segment from the plurality of segments, assigning the first segment and the second segment to an intersegmental sequence. 
     
     
         9 . The method of  claim 5 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format. 
     
     
         10 . The method of  claim 5 , wherein the sequence includes quality score information. 
     
     
         11 . The method of  claim 5 , wherein each element from the plurality of elements includes a read pair. 
     
     
         12 . The method of  claim 5 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence. 
     
     
         13 . The method of  claim 5 , wherein each segment from the plurality of segments includes a portion overlapping a portion of at least one remaining segment from the plurality of segments. 
     
     
         14 . The method of  claim 5 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
 the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size.   
     
     
         15 . The method of  claim 5 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
 the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments,   the realigned segment associated with each segment from the plurality of segments including a portion having a third size overlapping a portion of the realigned segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size, the third size being smaller than the second size.   
     
     
         16 . An apparatus, comprising:
 an assignment module, implemented in a memory or a processor, configured to:
 receive a sequence, the sequence including an estimation of a biological sequence, the sequence including a plurality of elements; 
 assign each element from the plurality of elements to at least one segment from a plurality of segments, including, when an element from the plurality of elements maps to at least a first segment and a second segment from the plurality of segments, assigning the element from the plurality of elements to both the first segment and the second segment; and 
   a segment processing module operatively coupled to the assignment module, the segment processing module configured to, for each segment from the plurality of segments:
 remove at least a portion of duplicate elements in the segment generate a deduplicated segment; and 
 reorder the elements in the deduplicated segment to generate a realigned segment, the realigned segment having a reduced likelihood for alignment errors than the deduplicated segment, 
   the segment processing module further configured to execute the removing and the reordering for at least two segments from the plurality of segments in a substantially simultaneous manner.   
     
     
         17 . The apparatus of  claim 16 , wherein the segment processing module is further configured to execute the removing and the reordering for the plurality of segments in a substantially simultaneous manner. 
     
     
         18 . The apparatus of  claim 16 , wherein the sequence is a target sequence, the assignment module further configured to receive the target sequence by:
 receiving a first sequence, the first sequence including a forward estimation of the biological sequence;   receiving a second sequence, the second sequence including a reverse estimation of the biological sequence;   generating a paired sequence based on the first sequence and the second sequence; and   aligning the paired sequence with a reference sequence to generate the target sequence.   
     
     
         19 . The apparatus of  claim 16 , further comprising a parallelization module configured to, prior to the assigning by the assignment module, split the sequence into a plurality of subsequences,
 the assignment module configured to assign by assigning each subsequence to at least one segment from the plurality of segments, the assignment module further configured to execute the assigning for at least two segments from the plurality of segments in a substantially simultaneous manner,   the parallelization module further configured to, for each segment from the plurality of segments, subsequent to the assigning by the assignment module and prior to the removing by the segment processing module, combine subsequences within that segment.   
     
     
         20 . The apparatus of  claim 16 , wherein each segment from the plurality of segments includes a portion having a first size overlapping a portion of at least one remaining segment from the plurality of segments,
 the deduplicated segment associated with each segment from the plurality of segments including a portion having a second size overlapping a portion of the deduplicated segment associated with a remaining segment from the plurality of segments,   the realigned segment associated with each segment from the plurality of segments including a portion having a third size overlapping a portion of the realigned segment associated with a remaining segment from the plurality of segments, the second size being smaller than the first size, the third size being smaller than the second size.   
     
     
         21 . The apparatus of  claim 16 , wherein the sequence is in a binary alignment/map (BAM) format or in a FASTQ format. 
     
     
         22 . The apparatus of  claim 16 , wherein the biological sequence is one of a deoxyribonucleic acid (DNA) sequence or a ribonucleic (RNA) sequence.

Join the waitlist — get patent alerts

Track US2017316154A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.