Dna random access storage system via ligation
Abstract
Techniques for random access of particular DNA strands from a mixture of DNA strands are described. DNA strands that encode pieces of the same digital file are labeled with the same identification sequence. The identification sequence is used to selectively separate DNA strands that contain portions of the same digital file from other DNA strands. A DNA staple positions DNA strands with the identification sequence adjacent to sequencing adaptors. DNA ligase joins the molecules to create a longer molecule with the region encoding the digital file flanked by sequencing adaptors. DNA strands that include sequencing adaptors are sequenced and the sequence data is available for further analysis. DNA strands without the identification sequence are not joined to sequencing adaptors, and thus, are not sequenced. As a result, the sequencing data produced by the DNA sequencer comes from those DNA strands that included the identification sequence.
Claims
exact text as granted — not AI-modified1 . A system comprising:
one or more processing units; memory coupled to the one or more processing units; instructions stored in the memory and executed on the one or more processing units that cause the system to:
receive an indication of a digital file;
identify a sequence of DNA nucleotides that is an identification (ID) sequence for the digital file, the ID sequence present on at least one of a 5′-end or a 3′-end of multiple DNA strands that respectively encode one of multiple portions of the digital file;
receive an indication of a sequencing technique;
identify an end sequence of a sequencing adaptor used in the sequencing technique; and
design a staple that is complementary in part to the ID sequence and complementary in part to the end sequence of the sequencing adaptor.
2 . The system of claim 1 , wherein the instructions further cause the system to receive an indication of a length of overlap between the staple and the end sequence of the sequencing adapter.
3 . The system of claim 1 , wherein the instructions further cause the system to send instructions to an oligonucleotide synthesizer to synthesize multiple copies of the staple.
4 . The system of claim 1 , wherein the instructions further cause the system to send instructions to combine a DNA pool with the staple and the sequencing adaptor, wherein the DNA pool contains the multiple DNA strands that encode a portion of the digital file and other DNA strands encoding portions of one or more different digital files.
5 . The system of claim 4 , wherein the other DNA strands encoding portions of one or more different digital files contain no sequences that are complementary to the staple.
6 . The system of claim 1 , wherein the instructions further cause the system to send instructions to ligate the sequencing adaptor to the DNA strands that encode one of the multiple portions of the digital file.
7 . The system of claim 1 , wherein the instructions further cause the system to:
receive DNA sequence read data from a DNA sequencer; and
identify reads within the DNA sequence read data that encode one of the multiple portions of the digital file based at least in part on the presence of the ID sequence in the reads.
8 . A method comprising:
receiving an indication of a digital file; identifying a sequence of DNA nucleotides that is an identification (ID) sequence for the digital file, the ID sequence present on at least one of a 5′-end or a 3′-end of multiple DNA strands that respectively encode one of multiple portions of the digital file; receiving an indication of a sequencing technique; identifying an end sequence of a sequencing adaptor used in the sequencing technique; and designing a staple that is complementary in part to the ID sequence and complementary in part to the end sequence of the sequencing adaptor.
9 . The method of claim 8 , further comprising receiving an indication of a length of overlap between the staple and the end sequence of the sequencing adapter.
10 . The method of claim 8 , further comprising sending instructions to an oligonucleotide synthesizer to synthesize multiple copies of the staple.
11 . The method of claim 8 , further comprising sending instructions to combine a DNA pool with the staple and the sequencing adaptor, wherein the DNA pool contains the multiple DNA strands that encode a portion of the digital file and other DNA strands encoding portions of one or more different digital files.
12 . The method of claim 11 , wherein the other DNA strands encoding portions of one or more different digital files contain no sequences that are complementary to the staple.
13 . The method of claim 8 , further comprising sending instructions to ligate the sequencing adaptor to the DNA strands that encode one of the multiple portions of the digital file.
14 . The method of claim 8 , further comprising:
receiving DNA sequence read data from a DNA sequencer; and identifying reads within the DNA sequence read data that encode one of the multiple portions of the digital file based at least in part on the presence of the ID sequence in the reads.
15 . A computer-readable storage media encoding instructions which when executed by a processing unit cause a computing device to perform acts comprising:
receiving an indication of a digital file; identifying a sequence of DNA nucleotides that is an identification (ID) sequence for the digital file, the ID sequence present on at least one of a 5′-end or a 3′-end of multiple DNA strands that respectively encode one of multiple portions of the digital file; receiving an indication of a sequencing technique; identifying an end sequence of a sequencing adaptor used in the sequencing technique; and designing a staple that is complementary in part to the ID sequence and complementary in part to the end sequence of the sequencing adaptor.
16 . The computer-readable storage media of claim 15 , wherein the acts further comprise sending instructions to an oligonucleotide synthesizer to synthesize multiple copies of the staple.
17 . The computer-readable storage media of claim 15 , wherein the acts further comprise sending instructions to combine a DNA pool with the staple and the sequencing adaptor, wherein the DNA pool contains the multiple DNA strands that encode a portion of the digital file and other DNA strands encoding portions of one or more different digital files.
18 . The computer-readable storage media of claim 15 , wherein the other DNA strands encoding portions of one or more different digital files contain no sequences that are complementary to the staple.
19 . The computer-readable storage media of claim 15 , wherein the acts further comprise sending instructions to ligate the sequencing adaptor to the DNA strands that encode one of the multiple portions of the digital file.
20 . The computer-readable storage media of claim 15 , wherein the acts further comprise:
receiving DNA sequence read data from a DNA sequencer; and identifying reads within the DNA sequence read data that encode one of the multiple portions of the digital file based at least in part on the presence of the ID sequence in the reads.Join the waitlist — get patent alerts
Track US2023395198A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.