Feature extraction and normalization algorithms for high-density oligonucleotide gene expression array data
Abstract
A characteristic intensity of a feature in image data generated by scanning a microarray probe is determined. A set of pixels of the image data that nominally represent the feature is identified. The pixels each have an value (such as an intensity value) associated therewith. For each of a plurality of subsets of the set of pixels, a variation statistic value is determined that corresponds to a variation in the values associated with the pixels of that subset. One of the subsets of pixels is chosen based on the determined variation statistic values. A method is also described to relate a first expression array of probes to a second expression array of probes. A subset of the probe for the arrays is determined based on a comparison of the ordering of the subset of the probes of the second array, according to a particular characteristic of the probes, to the ordering of corresponding probes in the first array according to the particular characteristic of the probes. A relationship of the second expression array to the first expression array is determined based on the subset of probes of the second expression array to the corresponding probes of the first array.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of determining a characteristic intensity of a feature in image data generated by scanning a microarray probe, comprising:
identifying a set of pixels of the image data that nominally represent the feature, wherein the pixels each have an intensity value associated therewith; for each of a plurality of subsets of the set of pixels, determining a variation statistic value that corresponds to a variation in the intensity values associated with the pixels of that subset; and choosing one of the subsets of pixels based on the determined variation statistic values.
2 . The method of claim 1 , and further comprising:
determining the characteristic intensity of the feature based on the intensity values associated with the pixels of the chosen subset.
3 . The method of claim 2 , wherein the subsets of the set of pixels are determined in an iterative fashion starting from a “seed” set of pixels and forming subsequent subsets by adding additional pixels.
4 . The method of claim 3 , wherein each subsequent subset is chosen at each iteration from predetermined additional subsets to have the smallest determined variation statistic value of the predetermined additional subsets.
5 . The method of claim 4 , wherein the predetermined additional subsets all have substantially a predetermined shape.
6 . A method of determining a characteristic value for a feature in image data generated by scanning a microarray probe, comprising:
identifying a set of pixels of the image data that nominally represent the feature, wherein the pixels each have a value associated therewith; for each of a plurality of subsets of the set of pixels, determining a variation statistic value that corresponds to a variation in the values associated with the pixels of that subset; choosing one of the subsets of pixels based on the determined variation statistic values.
7 . The method of claim 6 , and further comprising:
determining the characteristic value of the feature based on the intensity values associated with the pixels of the chosen subset.
8 . The method of claim 7 , wherein the subsets of the set of pixels are determined in an iterative fashion starting from a “seed” set of pixels and forming subsequent subsets by adding additional pixels.
9 . The method of claim 8 , wherein each subsequent subset is chosen at each iteration from predetermined additional subsets to have the smallest determined variation statistic value of the predetermined additional subsets.
10 . The method of claim 9 , wherein the predetermined additional subsets all have substantially a predetermined shape.
11 . A method of relating a first expression array of probes to a second expression array of probes, comprising:
determining a subset of the probes for the arrays based on a comparison of the ordering of the subset of the probes of the second array, according to a particular characteristic of the probes, to the ordering of corresponding probes in the first array according to the particular characteristic of the probes; and determining a relationship of the second expression array to the first expression array based on the subset of probes of the second expression array to the corresponding probes of the first array.
12 . The method of claim 11 , wherein the step of determining a subset of probes includes:
selecting the determined subset of probes from a plurality of subsets of the probes.
13 . The method of claim 12 , wherein the selecting step comprises:
comparing the ordering of a first of the plurality of subsets of the probes of the second array, according to a particular characteristic of the probes, to the ordering of corresponding probes in the first array according to the particular characteristic of the probes; if the comparison does not meet a particular criterion, repeating the comparing step with a second of the plurality of subsets of the probes of the second array, wherein the second subset is a subset of the first subset; using the subset of the probes for which the comparison meets the particular criterion to determine the relationship of the second expression array to the first expression array.
14 . The method of claim 13 , wherein the particular criterion is a threshold.
15 . The method of claim 14 , wherein the threshold is not identical for all the probes.
16 . The method of claim 13 , wherein the step of using the probes to determine the relationship includes applying a nonlinear regression technique to the subset of the probes.
17 . The method of claim 16 , wherein the step of applying a nonlinear regression technique includes applying a generalized cross-validation step to the subset of the probes of the second array and the corresponding probes of the first array.
18 . The method of claim 14 , wherein:
the comparison D i and threshold R i are determined by the equations: R i = [L(B i +E i )+H(2N−B i −E i )]/ 2n D i = 2|B i −E i | / (B i +E i ) where L and H are rank difference thresholds for the low and high ends of the range of characteristic values, B i and E i are the ranks for the i th characteristic value of the baseline and experiment arrays, and N is the total number of characteristic values that were ordered in the current iteration of the method.Join the waitlist — get patent alerts
Track US2004033485A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.