Computer Method and Apparatus for Processing Image Data
Abstract
A data compression method and apparatus that includes detecting a portion of a signal comprising a sequence of video frames that uses a disproportionate amount of bandwidth compared to other portions of the signal. The detected portion of the signal result in determined components of interest. Relative to certain variance, these components of interest are normalized to generate an intermediate form, which represents the components of interest reduced in complexity by the certain variance and enables a compressed form of the signal that maintains saliency. The detecting includes any of: (i) analyzing image gradients across frames where image gradient is a first derivative model and gradient flow is a second derivative, (ii) integrating finite differences of pels temporally/spatially to form a derivative model, (iii) analyzing an illumination field across frames, and (iv) predictive analysis, to determine bandwidth consumption, which is used to determine the components of interest.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A data processing system for generating an encoded form of video signal data from a plurality of video frames, the system comprising:
(a) an object detector configured to detect at least one object in two or more given video frames based on bandwidth consumption; (b) an object tracker, in communication with the object detector, configured to track the at least one object through the two or more video frames; (c) a segmenter, in communication with the object detector and the object tracker, configured to segment pel data corresponding to the at least one object from other pel data in the two or more video frames so as to generate a first intermediate form of the data, the segmenting utilizing a spatial segmentation of the pel data, the first intermediate form of the data including the segmented pel data of the at least one object and the other pel data in the two or more video frames; and (d) a normalizer, in communication with the segmenter, configured to normalize the first intermediate form of the data to generate a second intermediate form of the data by:
identifying corresponding elements of the at least one object in the given two or more video frames;
analyzing the corresponding elements to generate relationships between the corresponding elements;
generating correspondence models by using the generated relationships between the corresponding elements;
integrating the relationships between the corresponding elements into a model of global motion; and
re-sampling pel data associated with the at least one object in the two or more video frames by utilizing the correspondence models and model of global motion to generate a structural model or an appearance model representing a second intermediate form of the data.
2 . The data processing system of claim 1 comprising an encoder configured to encode the second intermediate form of the data by:
decomposing the re-sampled pel data into an encoded representation, the encoded representation representing a third intermediate form of the data;
truncating zero or more bytes of the encoded representation; and
recomposing the re-sampled pel data from the encoded representation;
wherein each of the decomposing and the recomposing uses Principal Component Analysis.
3 . The data processing system of claim 1 wherein the normalizer is configured to factor the correspondence models into local deformation models by:
generating a two dimensional mesh overlying pels corresponding to the at least one object, the mesh being based on a regular grid of vertices and edges, and
generating a model of local motion from the relationships between the corresponding elements, the relationships comprising vertex displacements based on finite differences generated from a block-based motion estimation between two or more of the video frames.
4 . The data processing system of claim 3 wherein the vertices correspond to discrete image features, the normalizer configured to identify significant image features corresponding to the object by using an analysis of the image intensity gradient.
5 . The data processing system of claim 1 further comprising an encoder configured to:
(e) restore spatial positions of the re-sampled pel data by utilizing the correspondence models, thereby generating restored pels corresponding to the at least one object; and
(f) recombine the restored pels together with the other pel data in the first intermediate form of the data to create an original video frame; and wherein the second intermediate form of the data is sufficiently reduced in complexity to enabling data compression by linear decomposition in an improved manner while maintaining saliency of the at least one object.
6 . The data processing system of claim 1 wherein the object detector and the object tracker use a face detector to facilitate the respective detection and the tracking of the at least one object.
7 . The data processing system of claim 1 wherein the normalizer configured to analyze the corresponding elements further includes the normalizer configured to use an appearance-based motion estimator configured to facilitate appearance-based motion estimation of the at least one object between two or more of the video frames.
8 . A computer-implemented method executing on one or more processors that generates an encoded form of video signal data from a plurality of video frames, the method comprising:
(a) based on bandwidth consumption, detecting at least one object in two or more given video frames; (b) tracking the at least one object through the two or more video frames; (c) segmenting pel data corresponding to the at least one object from other pel data in the two or more video frames so as to generate a first intermediate form of the data, the segmenting utilizing a spatial segmentation of the pel data, the first intermediate form of the data including the segmented pel data of the at least one object and the other pel data in the two or more video frames; (d) normalizing the first intermediate form of the data to generate a second intermediate form of the data by performing one or more of the following steps:
identifying corresponding elements of the at least one object in the given two or more video frames;
analyzing the corresponding elements to generate relationships between the corresponding elements;
generating correspondence models by using the generated relationships between the corresponding elements;
integrating the relationships between the corresponding elements into a model of global motion; and
re-sampling pel data associated with the at least one object in the two or more video frames by utilizing the correspondence models and model of global motion to generate a structural model or an appearance model representing a second intermediate form of the data.
9 . The method of claim 8 comprising encoding the second intermediate form of the data, the encoding comprising:
decomposing the re-sampled pel data into an encoded representation, the encoded representation representing a third intermediate form of the data;
truncating zero or more bytes of the encoded representation; and
recomposing the re-sampled pel data from the encoded representation;
wherein each of the decomposing and the recomposing uses Principal Component Analysis.
10 . The method of claim 8 comprising a method of factoring the correspondence models into local deformation models, the method comprising:
defining a two dimensional mesh overlying pels corresponding to the at least one object, the mesh being based on a regular grid of vertices and edges, and
generating a model of local motion from the relationships between the corresponding elements, the relationships comprising vertex displacements based on finite differences generated from a block-based motion estimation between two or more of the video frames.
11 . The method of claim 10 wherein the vertices correspond to discrete image features, the method comprising identifying significant image features corresponding to the object by using an analysis of the image intensity gradient.
12 . The method of claim 8 further comprising the steps of:
(e) restoring spatial positions of the re-sampled pel data by utilizing the correspondence models, thereby generating restored pels corresponding to the at least one object; and
(f) recombining the restored pels together with the other pel data in the first intermediate form of the data to create an original video frame; and wherein the second intermediate form of the data is sufficiently reduced in complexity to enabling data compression by linear decomposition in an improved manner while maintaining saliency of the at least one object.
13 . The method of claim 8 wherein the detecting and tracking comprise using a face detector.
14 . The method of claim 8 wherein analyzing the corresponding elements comprises using an appearance-based motion estimation between two or more of the video frames.
15 . A computer apparatus configured to generate an encoded form of video signal data from a plurality of video frames, the apparatus comprising:
one or more computer processors configured to execute: (a) an object detector configured to detect at least one object in two or more given video frames based on bandwidth consumption; (b) an object tracker, in communication with the object detector, configured to track the at least one object through the two or more video frames; (c) a segmenter, in communication with the object detector and the object tracker, configured to segment pel data corresponding to the at least one object from other pel data in the two or more video frames so as to generate a first intermediate form of the data, the segmenting utilizing a spatial segmentation of the pel data, the first intermediate form of the data including the segmented pel data of the at least one object and the other pel data in the two or more video frames; and (d) a normalizer, in communication with the segmenter, configured to normalize the first intermediate form of the data to generate a second intermediate form of the data by: identifying corresponding elements of the at least one object in the given two or more video frames; analyzing the corresponding elements to generate relationships between the corresponding elements; generating correspondence models by using the generated relationships between the corresponding elements; integrating the relationships between the corresponding elements into a model of global motion; and re-sampling pel data associated with the at least one object in the two or more video frames by utilizing the correspondence models and model of global motion to generate a structural model or an appearance model representing a second intermediate form of the data.Join the waitlist — get patent alerts
Track US2013083854A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.