US2002001804A1PendingUtilityA1
Genomic analysis of tRNA gene sets
Priority: Feb 25, 2000Filed: Feb 23, 2001Published: Jan 3, 2002
Est. expiryFeb 25, 2020(expired)· nominal 20-yr term from priority
G16B 30/10G16B 30/00
38
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods for identifying one or more positions of conserved difference in a set of similar sequence strings are provided, as well as systems and devices for identifying one or more positions of conserved difference in a set of similar sequence strings, and sets of positions of conserved differences.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for identifying one or more positions of conserved difference in a set of similar sequence strings, the method comprising:
providing a set of similar sequence strings derived from a plurality of species, wherein each similar sequence string comprises at least n sequence elements; comparing the at least n sequence elements in a first similar sequence string to the at least n sequence elements in a second similar sequence string, for a first species of the plurality of species; assigning a value to each of n positions of the at least n sequence elements, based upon whether the sequence elements are identical or different in the two similar sequence strings; repeating the comparing and assigning for each species in the plurality of species; summing the values assigned for each of the n positions across the plurality of species; and identifying which of the n positions have the greatest sum value, thereby identifying the positions of conserved difference in the set of similar sequence strings.
2 . The method of claim 1 , wherein each species in the plurality of species contributes at least two similar sequence strings to the set of similar sequence strings.
3 . The method of claim 1 , wherein each species in the plurality of species contributes more than two similar sequence strings to the set of similar sequence strings.
4 . The method of claim 1 , wherein the providing a set of similar sequence strings comprises:
providing a set of sequences; providing logical instructions for recognizing a target sequence string; and using the logical instructions to analyze the sequences and identify the target sequence strings, thereby providing a set of similar sequence strings.
5 . The method of claim 1 , wherein the set of similar sequence strings comprises sets of amino acid sequences, nucleic acid sequences, lipid-based sequences or carbohydrate sequences.
6 . The method of claim 5 , wherein the set of similar sequence strings comprises a set of tRNA molecules.
7 . The method of claim 5 , wherein the set of similar sequence strings comprises a set of alleles.
8 . The method of claim 7 , wherein the set of alleles comprises at least two alleles.
9 . The method of claim 7 , wherein the set of alleles comprises more than two alleles.
10 . The method of claim 1 , wherein the plurality of species comprises a plurality of prokaryotic species, eukaryote species, or combinations thereof.
11 . The method of claim 8 , wherein the plurality of prokaryotic species comprises a plurality of eubacteria species, archaea species, or combinations thereof.
12 . The method of claim 1 , wherein the comparing and assigning is performed in a computer.
13 . The method of claim 1 , further comprising determining whether the positions that have the greatest sum values comprise elements which interact with a protein, a peptide, a protein complex, a nucleic acid, a protein-nucleic acid complex, a carbohydrate chain, or a combination thereof.
14 . The method of claim 13 , wherein the protein comprises an enzyme.
15 . The method of claim 13 , wherein the protein-nucleic acid complex comprises a ribosome.
16 . The method of claim 1 , further comprising determining whether the positions that have the greatest sum values comprise modified elements.
17 . The method of claim 16 , wherein the modified elements comprise amino acids or nucleotides which are modified by methylation, acetylation, ubiquitination, lysinylation or glycosylation.
18 . A method for identifying one or more positions of conserved difference in a set of similar sequence strings, the method comprising:
providing a set of similar sequence strings derived from a plurality of species, wherein each similar sequence string comprises at least n sequence elements, and wherein each species in the plurality of species contributes two or more similar sequence strings to the set of similar sequence strings; simultaneously comparing the at least n sequence elements for the two or more similar sequence strings from a first species of the plurality of species; assigning a value to each of n positions of the at least n sequence elements, based upon whether the sequence elements are identical or different in the two or more similar sequence strings; repeating the comparing and assigning for each species in the plurality of species; summing the values assigned for each of the n positions across the plurality of species; and identifying which of the n positions have the greatest sum value, thereby identifying the positions of conserved difference in the set of similar sequence strings.
19 . The set of conserved differences in a set of similar sequence strings as identified by the method of claim 1 .
20 . A computer or computer-readable medium comprising one or more logical instructions for identifying at least one conserved difference in a set of similar sequence strings derived from a plurality of species,
wherein each species in the plurality of species comprises at least two similar sequence strings; and wherein the logical instructions compare at least n sequence elements in a first similar sequence string to at least n sequence elements in a second similar sequence string, for a first species of the plurality of species; assigns a value to each of n positions of the at least n sequence elements, based upon whether the sequence elements are identical or different in the two similar sequence strings; repeats the comparing and assigning for each species in the plurality of species; sums the values assigned for each of the n positions across the plurality of species; and identifies which of the n positions have the greatest sum value, thereby identifying the positions of conserved difference in the set of similar sequence strings.
21 . The computer or computer-readable medium of claim 20 , further comprising a database comprising the set of similar sequence strings derived from a plurality of species.
22 . The computer or computer-readable medium of claim 20 , comprising a neural network.
23 . The computer or computer-readable medium of claim 20 , comprising a user interface.
24 . The computer or computer-readable medium of claim 23 , wherein the user interface comprises an input field that permits data entry of the similar sequence strings.
25 . The computer or computer-readable medium of claim 23 , wherein the user interface comprises a data output file.
26 . The computer or computer-readable medium of claim 23 , wherein the user interface operates across a network.
27 . The computer or computer-readable medium of claim 23 , wherein the user interface operates across the internet.
28 . The computer or computer-readable medium of claim 23 , wherein the user interface comprises a web browser interface.Join the waitlist — get patent alerts
Track US2002001804A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.