US2022084629A1PendingUtilityA1
Systems and methods for barcode design and decoding
Est. expirySep 16, 2040(~14.1 yrs left)· nominal 20-yr term from priority
Inventors:Preyas Shah
C12Q 1/68C12N 15/625C12N 15/1062G16B 30/00G16B 5/20G06K 7/1473G16B 25/20G16B 15/00G16B 35/00G16B 30/20G16B 30/10
72
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods and systems for designing large sets of barcodes that ensure robust and efficient error correction capabilities are described. Also described are methods for assigning barcodes to target analytes that minimize optical crowding in in situ detection applications. Furthermore, methods for performing barcode error correction and for performing barcode-assisted image registration and alignment are also described.
Claims
exact text as granted — not AI-modified1 . An array comprising a plurality of unique nucleic acid barcode sequences, wherein a unique nucleic acid barcode sequence, or segment thereof, of the plurality of unique nucleic acid barcode sequences has:
a specified minimum pairwise edit distance of 3 relative to other unique nucleic acid barcode sequences, or segments thereof, of the array; and at least one additional characteristic selected from a list consisting of: a total length of at least 10 nucleotides, a minimum of two segments, a segment length of at least 2 nucleotides, a guanine-cytosine (GC) content of less than 50%, a maximum length for homopolymer subsequences of 7 nucleotides, and a dilution factor of at least 10% for at least one segment.
2 . The array of claim 1 , wherein the array is a spatial array and different unique nucleic acid barcode sequences are attached to different features of the spatial array.
3 . The array of claim 1 , wherein the array is a bead array, and different unique nucleic acid barcode sequences are attached to different beads of the bead array.
4 . The array of claim 1 , wherein a unique nucleic acid barcode sequence comprises a sequence of individual nucleotides.
5 . The array of claim 1 , wherein a unique nucleic acid barcode sequence comprises a plurality of segments, and each segment comprises a plurality of nucleotides.
6 . The array of claim 5 , wherein a unique nucleic acid barcode sequence comprises at most 20 segments.
7 . The array of claim 5 , wherein each segment comprises at most 20 nucleotides.
8 . (canceled)
9 . The array of claim 1 , wherein the specified minimum pairwise edit distance comprises a specified minimum pairwise Hamming distance of at least two times an error correction capability, and wherein the error correction capability has a value of at least one.
10 . The array of claim 1 , wherein the at least one additional characteristic comprises a guanine-cytosine (GC) content of less than about 10%.
11 . The array of claim 1 , wherein the at least one additional characteristic comprises a maximum length for homopolymer subsequences of 3 nucleotides.
12 . The array of claim 1 , wherein at least one segment of at least one barcode encodes for an “OFF” state that is not visualized during a decoding process used to detect and decode the nucleic acid barcode sequences.
13 . The array of claim 1 , wherein the at least one additional characteristic comprises compatibility with a specified decoding dilution factor of at least 50%.
14 . (canceled)
15 . The array of claim 1 , wherein the array comprises at least 1,000 unique nucleic acid barcode sequences.
16 .- 18 . (canceled)
19 . A composition comprising a plurality of target-specific probe molecules, wherein a target-specific probe molecule of the plurality comprises a unique nucleic acid barcode sequence selected from a plurality of unique nucleic acid barcode sequences.
20 . The composition of claim 19 , wherein the plurality of unique nucleic acid barcode sequences comprises at least 1,000 unique nucleic acid barcode sequences, and wherein a unique nucleic acid barcode sequence, or segment thereof, of the at least 1,000 unique nucleic acid barcode sequences has:
a specified minimum pairwise edit distance of 3 relative to other unique nucleic acid barcode sequences, or segments thereof, of the array; and at least one additional characteristic selected from a list consisting of: a total length of at least 10 nucleotides, a minimum of two segments, a segment length of at least 2 nucleotides, a guanine-cytosine (GC) content of less than 50%, a maximum length for homopolymer subsequences of 7 nucleotides, and a dilution factor of at least 10% for at least one segment.
21 . The composition of claim 19 , wherein a target-specific probe molecule of the plurality further comprises a target recognition element, a unique molecular identifier, a primer binding site, a linker region, one or more detectable tags, or any combination thereof.
22 . The composition of claim 19 , wherein the unique nucleic acid barcode sequences of the plurality of unique nucleic acid barcode sequences are rank-ordered according to an average pairwise edit distance from all other unique nucleic acid barcode sequences of the plurality, and assigned to a corresponding target gene transcript of the same rank from a list of corresponding genes rank-ordered by relative expression level.
23 . The composition of claim 19 , wherein the unique nucleic acid barcode sequences of the plurality of unique nucleic acid barcode sequences are organized as a plurality of barcode tuples each comprising two unique nucleic acid barcode sequences and a pairwise edit distance between them, wherein the target gene transcripts are organized as a plurality of gene tuples each comprising two target gene transcripts and a mean expression level for their corresponding genes, and wherein the nucleic acid barcode sequences of a barcode tuple comprising the largest pairwise edit distance are assigned to the target gene transcripts of a gene tuple comprising the largest mean expression level.
24 . (canceled)
25 . The composition of claim 22 , wherein the rank-ordered unique nucleic acid barcode sequences are assigned to corresponding rank-ordered target gene transcripts such that optical crowding is reduced during a decoding process used to decode the unique nucleic acid barcode sequences.
26 . A method for generating barcode sequences comprising:
providing a plurality of candidate barcode sequences; receiving a set of design criteria that specify a total number of unique designed barcode sequences, a maximum length for the designed barcode sequences, and a minimum pairwise edit distance for each designed barcode, or segment thereof, relative to other designed barcode sequences, or segments thereof; and applying the set of design criteria, using one or more processors and a metric tree data structure, to select a set of designed barcode sequences from the plurality of candidate barcode sequences, wherein the set of designed barcode sequences comprises the specified total number of unique barcode sequences, and wherein a unique designed barcode sequence, or segment thereof, of the set has:
the specified maximum nucleotide length; and
the specified minimum pairwise edit distance relative to other designed barcode sequences, or segments thereof, of the set.
27 . The method of claim 26 , wherein the designed barcode sequences comprise nucleic acid barcode sequences.
28 . The method of claim 26 , wherein a unique designed barcode sequence of the set further exhibits at least one additional characteristic selected from a list consisting of: a specified minimum number of segments, a specified minimum segment length, a specified upper limit on guanine-cytosine (GC) content, a specified maximum length for homopolymer subsequences, and a specified dilution factor for at least one segment.
29 . (canceled)
30 . The method of claim 26 , wherein the specified pairwise edit distance comprises a specified minimum pairwise Hamming distance of at least two times a specified error correction capability.
31 .- 32 . (canceled)
33 . The method of claim 28 , wherein the at least one additional characteristic comprises a specified upper limit on guanine-cytosine (GC) content of 50%.
34 . The method of claim 28 , wherein the at least one additional characteristic comprises a specified maximum length for homopolymer subsequences of 7 nucleotides.
35 . The method of claim 28 , wherein the at least one additional characteristic comprises a specified dilution factor of at least 10% for at least one segment.
36 . (canceled)
37 . The method of claim 26 , wherein each designed barcode sequence is rank-ordered according to an average pairwise edit distance from all other designed barcode sequences of the set, and assigned to a corresponding target gene transcript of the same rank from a list of corresponding genes rank-ordered by relative expression level.
38 . (canceled)
39 . The method of claim 37 , wherein the rank-ordered designed barcode sequences are assigned to corresponding rank-ordered target gene transcripts such that optical crowding is reduced during a decoding process used to decode the designed barcode sequences.
40 . (canceled)
41 . The method of claim 26 , wherein the metric tree data structure comprises an M-tree data structure, a vp-tree data structure, a cover tree data structure, an MVP tree data structure, or a BK-tree data structure.
42 . (canceled)
43 . The method of claim 26 , further comprising generating a set of barcode probes configured to detect the designed barcode sequences, or segments thereof, for use in decoding the set of designed barcode sequences.
44 . The method of claim 26 , further comprising incorporating each unique designed barcode sequence of the set into a target-specific probe molecule of a set of target-specific probe molecules.
45 . (canceled)
46 . The method of claim 26 , further comprising attaching each unique designed barcode sequence to a different feature of a spatial array.
47 . The method of claim 26 , further comprising attaching each unique designed barcode sequence to a different bead of a bead array.
48 . An array manufactured by attaching a unique nucleic acid barcode sequence to each array element of a plurality of array elements, wherein the unique nucleic acid barcode sequences are selected from a set of candidate nucleic acid barcode sequences based on the criteria that:
each selected nucleic acid barcode sequence has a specified maximum nucleotide length; and each selected nucleic acid barcode sequence, or segment thereof, has a specified minimum pairwise edit distance from every other selected nucleic acid barcode sequence, or segments thereof.
49 .- 50 . (canceled)
51 . A system comprising:
one or more processors; memory operably coupled to the one or more processors and comprising a metric tree data structure; and one or more programs stored in the memory that, when executed by the one or more processors, cause the system to execute a method comprising:
providing a plurality of candidate barcode sequences;
receiving a set of design criteria that specify a total number of unique designed barcode sequences, a maximum length for the designed barcode sequences, and a minimum pairwise edit distance for each designed barcode, or segment thereof, relative to other designed barcode sequences, or segments thereof; and
applying the set of design criteria, using one or more processors and a metric tree data structure, to select a set of designed barcode sequences from the plurality of candidate barcode sequences, wherein the set of designed barcode sequences comprises the specified total number of unique barcode sequences, and wherein a unique designed barcode sequence of the set, or segment thereof, has:
the specified maximum nucleotide length; and
the specified minimum pairwise edit distance relative to other designed barcode sequences, or segments thereof, of the set.
52 . (canceled)Join the waitlist — get patent alerts
Track US2022084629A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.