US2021193257A1PendingUtilityA1

Phase-aware determination of identity-by-descent dna segments

Assignee: 23ANDME INCPriority: Jul 19, 2019Filed: Mar 4, 2021Published: Jun 24, 2021
Est. expiryJul 19, 2039(~13 yrs left)· nominal 20-yr term from priority
G16B 50/30G16B 30/10G16B 30/00G16B 20/20C12Q 1/6827
72
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The disclosed embodiments concern methods, apparatus, systems and computer program products for estimating IBD segments. Some implementations use a templated positional Burrows-Wheeler transform (PBWT) technique and a phase switch error heuristic to correct genotyping errors and phase switch errors to make fast and accurate phase aware IBD estimates. In some implementations a templated PBWT technique and a probabilistic hidden Markov model (HMM) are used to correct genotyping errors and phase switch errors.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer implemented method of processing haplotypes to reduce genotyping errors when determining identity by descent (IBD) segments between haplotypes, the method comprising:
 providing a first digital template comprising a first arrangement of masked and unmasked sites in a window of consecutive haplotype sites;   providing a second digital template comprising a second arrangement of masked and unmasked sites in a window of consecutive haplotype sites, wherein the first and second arrangements are different;   providing two or more haplotypes strings for identification of IBD segments therebetween, each of the two or more haplotype strings representing a sequence of allele values at polymorphic sites in a haplotype of an organism; and   computationally identifying IBD segments between the two or more haplotype strings by (i) identifying first matches among alleles of the haplotype strings at unmasked sites produced by applying the first digital template to the two or more haplotype strings, (ii) identifying second matches among alleles of the haplotype at unmasked sites produced by applying the second digital template to the two or more haplotype strings, and (iii) merging the first and second matches among alleles to produce a merged set of IBD segments, wherein the merged set of IBD segments has reduced impact from genotyping errors compared to a set of IBD segments generated without applying the first and second digital templates.   
     
     
         2 . The method of  claim 1 , wherein the first and second templates each have a size of at least four consecutive haplotype sites. 
     
     
         3 . The method of  claim 1 , wherein identifying the first matches among alleles at unmasked sites comprises sequentially applying the first digital template to the two or more haplotype strings, each time moving to a next sequential section of the two or more haplotype strings. 
     
     
         4 . The method of  claim 1 , wherein computationally identifying IBD segments between the two or more haplotype strings further comprises:
 computationally identifying additional matches among alleles at unmasked sites produced by applying one or more additional digital templates to the two or more haplotype strings, wherein the one or more additional digital templates have additional arrangements of masked and unmasked sites in windows of consecutive haplotype sites, and each of the additional arrangements is different from both the first and the second arrangements, and   wherein merging the first and second matches among alleles to produce a merged set of IBD segments further comprises computationally merging the additional matches with the first and second matches to produce the merged set of IBD segments.   
     
     
         5 . The method of  claim 4 , wherein computationally identifying additional matches among alleles at unmasked sites employs a third digital template, a fourth digital template, a fifth digital template, and a sixth digital template. 
     
     
         6 . The method of  claim 5 , wherein the first through sixth digital templates each comprise two masked sites and two unmasked sites. 
     
     
         7 . The method of  claim 1 , wherein the first digital template and the second digital template each have a ratio of masked sites to unmasked sites of between about 2:1 to about 1:2. 
     
     
         8 . The method of  claim 1 , wherein the two or more haplotype strings comprise at least one thousand haplotype strings. 
     
     
         9 . The method of  claim 1 , wherein the two or more haplotype strings comprise at least one million haplotype strings. 
     
     
         10 . The method of  claim 1 , wherein computationally identifying IBD segments between the two or more haplotype strings comprises performing a positional Burrows-Wheeler transform (PBWT) on the unmasked sites produced by applying the first and second templates to the two or more haplotype strings. 
     
     
         11 . The method of  claim 10 , wherein computationally merging the first and second matches among alleles is performed while considering individual polymorphic sites of the two or more haplotype strings using the PBWT. 
     
     
         12 . The method of  claim 1 , wherein the total number of digital templates is between 2 and k, where k is the number of haplotype sites in the window. 
     
     
         13 . The method of  claim 1 , wherein the total number of digital templates is k!/(m!*(k-m)!), where k is the number of haplotype sites in the window and m is the number of masked sites in the window. 
     
     
         14 . The method of  claim 1 , wherein applying the first digital template comprises a deterministic process employing the first arrangement of masked and unmasked sites. 
     
     
         15 . A system for processing haplotypes to reduce genotyping errors when determining identity by descent (IBD) segments between haplotypes, the system comprising:
 (a) one or more processors and associated memory;   (b) computer readable instructions for:   providing a first digital template comprising a first arrangement of masked and unmasked sites in a window of consecutive haplotype sites;   providing a second digital template comprising a second arrangement of masked and unmasked sites in a window of consecutive haplotype sites, wherein the first and second arrangements are different;   providing two or more haplotype strings for identification of IBD segments therebetween, each of the two or more haplotype strings representing a sequence of allele values at polymorphic sites in a haplotype of an organism; and   identifying IBD segments between the two or more haplotype strings by (i) identifying first matches among alleles of the haplotype strings at unmasked sites produced by applying the first digital template to the two or more haplotype strings, (ii) identifying second matches among alleles of the haplotype at unmasked sites produced by applying the second digital template to the two or more haplotype strings, and (iii) merging the first and second matches among alleles to produce a merged set of IBD segments, wherein the merged set of IBD segments has reduced impact from genotyping errors compared to a set of IBD segments generated without applying the first and second digital templates.   
     
     
         16 . The system of  claim 15 , wherein the first and second templates each have a size of at least four consecutive haplotype sites. 
     
     
         17 . The system of  claim 15 , wherein the instructions for identifying the first matches among alleles at unmasked sites comprises instructions for sequentially applying the first digital template to the two or more haplotype strings, each time moving to a next sequential section of the two or more haplotype strings. 
     
     
         18 . The system of  claim 15 , wherein the instructions for identifying IBD segments between the two or more haplotype strings further comprise instructions for:
 computationally identifying additional matches among alleles at unmasked sites produced by applying one or more additional digital templates to the two or more haplotype strings, wherein the one or more additional digital templates have additional arrangements of masked and unmasked sites in windows of consecutive haplotype sites, and each of the additional arrangements is different from both the first and the second arrangements, and   wherein merging the first and second matches among alleles to produce a merged set of IBD segments further comprises computationally merging the additional matches with the first and second matches to produce the merged set of IBD segments.   
     
     
         19 . The system of  claim 18 , wherein the instructions for identifying additional matches among alleles at unmasked sites employ a third digital template, a fourth digital template, a fifth digital template, and a sixth digital template. 
     
     
         20 . The system of  claim 15 , wherein the two or more haplotype strings comprise at least one thousand haplotype strings. 
     
     
         21 . The system of  claim 15 , wherein the instructions for identifying IBD segments between the two or more haplotype strings comprise instructions performing a positional Burrows-Wheeler transform (PBWT) on the unmasked sites produced by applying the first and second templates to the two or more haplotype strings. 
     
     
         22 . The system of  claim 21 , wherein the instructions for merging the first and second matches among alleles comprise instructions for performing the merging while considering individual polymorphic sites of the two or more haplotype strings using the PBWT. 
     
     
         23 . A method of identifying IBD segments between two or more haplotype strings, each of the two or more haplotype strings representing a sequence of allele values at polymorphic sites in a haplotype of an organism, the method comprising:
 (a) computationally identifying IBD segments between the two or more haplotype strings by (i) identifying first matches among alleles of two or more haplotype strings at unmasked sites produced by applying a first digital template to the two or more haplotype strings, (ii) identifying second matches among alleles of the haplotype at unmasked sites produced by applying a second digital template to the two or more haplotype strings, and (iii) merging the first and second matches among alleles to produce a merged set of IBD segments, wherein the first digital template comprises a first arrangement of masked and unmasked sites in a window of consecutive haplotype sites, wherein the second digital template comprises a second arrangement of masked and unmasked sites in the window of consecutive haplotype sites, and wherein the first and second arrangements are different; and   (b) identifying a potential phase switch error in at least one of the two or more haplotype strings; and   (c) correcting the phase switch error.   
     
     
         24 . The method of  claim 23 , wherein identifying the potential phase switch error comprises identifying proximal IBD segments in at least one pair of the two or more haplotype strings. 
     
     
         25 . The method of  claim 23 , wherein identifying the potential phase switch error further comprises computationally iterating through the two or more haplotype strings, by:
 (i) identifying a first potential IBD segment between the two or more haplotype strings;   (ii) comparing the first site of the first potential IBD segment to the last site of a previously identified second potential IBD segment; and   (iii) determining that the last site of the second potential IBD segment and the first site of the first potential IBD segment are within a threshold number of sites of each other; and   wherein correcting the phase switch error comprises merging the first potential IBD segment and the second potential IBD segment to form a combined potential IBD segment based on determining that the last site of the second potential IBD segment and the first site of the first potential IBD segment are within the threshold number of sites of each other.   
     
     
         26 . The method of  claim 25 , wherein the first potential IBD segment and the second potential IBD segment each have a length of at least the threshold number of sites. 
     
     
         27 . The method of  claim 25 , wherein the threshold number of sites is between about 0 and 500 SNPs.

Join the waitlist — get patent alerts

Track US2021193257A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.