US2014222348A1PendingUtilityA1
Mass Spectrometry
Assignee: SHIMADZU RES LAB EUROPE LTDPriority: May 30, 2002Filed: Dec 30, 2013Published: Aug 7, 2014
Est. expiryMay 30, 2022(expired)· nominal 20-yr term from priority
G16B 30/00G01N 33/6848H01J 49/0036G06F 19/22
62
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present invention is concerned with methods for the de novo sequencing of polypeptides from data obtained from mass spectrometry devices, particularly from (MS) n devices.
Claims
exact text as granted — not AI-modified1 - 22 . (canceled)
23 . A method for determining at least one amino acid sequence for a sample polypeptide, the method being implemented in a computer comprising a memory storing processor readable instructions and a processor in communication with said memory, the method comprising:
receiving, as input to the processor, a set of m/z peaks associated with a soft ionization mass spectrum obtained from said sample polypeptide, each of said m/z peaks having an associated m/z value; identifying, by the processor, a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks; processing, by the processor, said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets; discarding any identified m/z peak sets from said plurality of m/z peak sets; and outputting said at least one amino acid sequence based upon said remaining plurality of m/z peak sets.
24 . The method of claim 23 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises, for each m/z peak set:
comparing the possible amino acid sequence associated with the m/z peak set with possible amino acid sequences associated with other ones of the plurality of m/z peak sets in reverse order.
25 . The method of claim 23 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises:
for each of said m/z peak sets:
determining mass differences between each pair of neighbouring m/z peaks of the m/z peak set based upon the m/z values associated with each pair of m/z peaks;
determining a sequence of mass differences based upon the determined mass differences; and
determining a sequence of mass differences in reverse order based upon the determined mass differences; and
identifying m/z peak sets whose sequence of mass differences does not form at least part of a sequence of mass differences in reverse order of another of said m/z peak sets.
26 . The method of claim 23 , wherein identifying a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks, comprises:
determining a plurality of said received set of m/z peaks that each have an associated m/z value that differs from an m/z value associated with another of said plurality of said received set of m/z peaks by the mass of an amino acid.
27 . The method of claim 23 , further comprising:
identifying any m/z peak set which is a contiguous subsequence of another m/z peak set; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
28 . The method of claim 23 , further comprising:
identifying candidate m/z peak sets in which mass differences between each m/z peak and at least one peak in the candidate m/z peak set having a closest m/z value above and/or below the peak correspond to that of a daughter ion or an end ion cluster; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
29 . The method of claim 23 , wherein outputting said at least one amino acid sequence based upon said remaining plurality of m/z peak sets comprises:
for each of said remaining plurality of m/z peak sets, determining an amino acid sequence based upon a difference between m/z values associated with peaks of the m/z peak set.
30 . A computer program product for determining at least one amino acid sequence for a sample polypeptide, the computer program product comprising a non-transient computer usable medium having a computer readable program code, said computer readable program code comprising instructions arranged to cause a computer to:
receive, as input to the processor, a set of m/z peaks associated with a soft ionization mass spectrum obtained from said sample polypeptide, each of said m/z peaks having an associated m/z value; identify, by the processor, a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks; process, by the processor, said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets; discard any identified m/z peak sets from said plurality of m/z peak sets; and
output said at least one amino acid sequence based upon said remaining plurality of m/z peak sets.
31 . The computer program product of claim 30 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises, for each m/z peak set:
comparing the possible amino acid sequence associated with the m/z peak set with possible amino acid sequences associated with other ones of the plurality of m/z peak sets in reverse order.
32 . The computer program product of claim 30 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises:
for each of said m/z peak sets:
determining mass differences between each pair of neighbouring m/z peaks of the m/z peak set based upon the m/z values associated with each pair of m/z peaks;
determining a sequence of mass differences based upon the determined mass differences; and
determining a sequence of mass differences in reverse order based upon the determined mass differences; and
identifying m/z peak sets whose sequence of mass differences does not form at least part of a sequence of mass differences in reverse order of another of said m/z peak sets.
33 . The computer program product of claim 30 , wherein identifying a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks, comprises:
determining a plurality of said received set of m/z peaks that each have an associated m/z value that differs from an m/z value associated with another of said plurality of said received set of m/z peaks by the mass of an amino acid.
34 . The computer program product of claim 30 , wherein said computer readable program code further comprises instructions arranged to cause a computer to:
identifying any m/z peak set which is a contiguous subsequence of another m/z peak set; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
35 . The computer program product of claim 30 , further comprising:
identifying candidate m/z peak sets in which mass differences between each m/z peak and at least one peak in the candidate m/z peak set having a closest m/z value above and/or below the peak correspond to that of a daughter ion or an end ion cluster; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
36 . The computer program product of claim 30 , wherein outputting said at least one amino acid sequence based upon said remaining plurality of m/z peak sets comprises:
for each of said remaining plurality of m/z peak sets, determining an amino acid sequence based upon a difference between m/z values associated with peaks of the m/z peak set.
37 . A system for determining at least one amino acid sequence for a sample polypeptide, said system comprising:
a memory storing processor readable instructions; and a processor arranged to read and execute instructions stored in said memory; wherein said processor readable instructions contain instructions to cause a computer to: receive, as input to the processor, a set of m/z peaks associated with a soft ionization mass spectrum obtained from said sample polypeptide, each of said m/z peaks having an associated m/z value; identify, by the processor, a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks; process, by the processor, said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets; discard any identified m/z peak sets from said plurality of m/z peak sets; and
output said at least one amino acid sequence based upon said remaining plurality of m/z peak sets.
38 . The system of claim 37 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises, for each m/z peak set:
comparing the possible amino acid sequence associated with the m/z peak set with possible amino acid sequences associated with other ones of the plurality of m/z peak sets in reverse order.
39 . The system of claim 37 , wherein processing said plurality of m/z peak sets using a reflective predicate filter to identify m/z peak sets comprises:
for each of said m/z peak sets:
determining mass differences between each pair of neighbouring m/z peaks of the m/z peak set based upon the m/z values associated with each pair of m/z peaks;
determining a sequence of mass differences based upon the determined mass differences; and
determining a sequence of mass differences in reverse order based upon the determined mass differences; and
identifying m/z peak sets whose sequence of mass differences does not form at least part of a sequence of mass differences in reverse order of another of said m/z peak sets.
40 . The system of claim 37 , wherein identifying a plurality of m/z peak sets, each of said plurality of m/z peak sets corresponding to a possible amino acid sequence associated with the received set of m/z peaks, comprises:
determining a plurality of said received set of m/z peaks that each have an associated m/z value that differs from an m/z value associated with another of said plurality of said received set of m/z peaks by the mass of an amino acid.
41 . The system of claim 37 , wherein said computer readable program code further comprises instructions arranged to cause a computer to:
identifying any m/z peak set which is a contiguous subsequence of another m/z peak set; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
42 . The system of claim 37 , further comprising:
identifying candidate m/z peak sets in which mass differences between each m/z peak and at least one peak in the candidate m/z peak set having a closest m/z value above and/or below the peak correspond to that of a daughter ion or an end ion cluster; and discarding any identified m/z peak sets from the plurality of candidate m/z peak sets.
43 . The system of claim 37 , wherein outputting said at least one amino acid sequence based upon said remaining plurality of m/z peak sets comprises:
for each of said remaining plurality of m/z peak sets, determining an amino acid sequence based upon a difference between m/z values associated with peaks of the m/z peak set.Join the waitlist — get patent alerts
Track US2014222348A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.