Video structuring by probabilistic merging of video segments
Abstract
A method for structuring video by probabilistic merging of video segments includes the steps of obtaining a plurality of frames of unstructured video; generating video segments from the unstructured video by detecting shot boundaries based on color dissimilarity between consecutive frames; extracting a feature set by processing pairs of segments for visual dissimilarity and their temporal relationship, thereby generating an inter-segment visual dissimilarity feature and an inter-segment temporal relationship feature; and merging video segments with a merging criterion that applies a probabilistic analysis to the feature set, thereby generating a merging sequence representing the video structure. The probabilistic analysis follows a Bayesian formulation and the merging sequence is represented in a hierarchical tree structure.
Claims
exact text as granted — not AI-modified1 . A method for representing contents of a video sequence, the video sequence including individual shots, and each shot including individual frames, the method comprising the steps of:
generating a tree representation of the video sequence; and storing the tree representation in a computer readable memory, wherein the tree representation including levels of nodes, wherein the levels include a top level including only a single complete-video node, a bottom level including a plurality of leaf nodes, and an intermediate level including a plurality of nodes fewer in number than the plurality of leaf nodes, wherein each node corresponds to one or more of the shots and each node is represented in the tree representation by a frame of its corresponding shot(s), wherein each leaf node corresponds to one of the shots, wherein the complete-video node corresponds to all of the shots in the video sequence, and wherein each node in the intermediate level corresponds to a cluster of shots, but not all of the shots in the video sequence.
2 . The method of claim 1 , further comprising the step of displaying the tree representation on a display.
3 . The method of claim 1 , further comprising the step of changing the cluster of shots associated with a node in the intermediate level based at least upon user feedback.
4 . The method of claim 2 , further comprising the steps of:
changing the cluster of shots associated with a node in the intermediate level based at least upon user feedback; and displaying the tree representation on the display based at least upon results of the changing step.
5 . The method of claim 2 , further comprising the steps of:
receiving a selection of a node; and initiating playback of the shot(s) that correspond(s) to the selected node.
6 . The method of claim 5 , wherein the initiating playback step includes providing playback control functionality to a user.
7 . The method of claim 6 , wherein the playback control functionality includes pause, stop, fast-forward, and rewind.
8 . The method of claim 1 , wherein each frame representing a node in the tree representation is a key frame.
9 . The method of claim 2 , further comprising the step of editing the video sequence based at least upon user interaction with the tree representation.
10 . The method of claim 2 , further comprising the steps of:
receiving a selection of a node; retrieving the shot(s) corresponding to the selected node; and storing the retrieved shot(s) in a computer-readable memory separate from the video sequence.
11 . The method of claim 10 , further comprising the step of deleting the retrieved shot(s) from the video sequence.
12 . A processor-accessible memory system storing instructions configured to cause a data processing system to implement a method for representing contents of a video sequence, the video sequence including individual shots, and each shot including individual frames, wherein the instructions comprise:
instructions for generating a tree representation of the video sequence; and instructions for storing the tree representation in a computer readable memory, wherein the tree representation including levels of nodes, wherein the levels include a top level including only a single complete-video node, a bottom level including a plurality of leaf nodes, and an intermediate level including a plurality of nodes fewer in number than the plurality of leaf nodes, wherein each node corresponds to one or more of the shots and each node is represented in the tree representation by a frame of its corresponding shot(s), wherein each leaf node corresponds to one of the shots, wherein the complete-video node corresponds to all of the shots in the video sequence, and wherein each node in the intermediate level corresponds to a cluster of shots, but not all of the shots in the video sequence.
13 . The processor-accessible memory system of claim 12 , wherein the instructions further comprise instructions for displaying the tree representation on a display.
14 . The processor-accessible memory system of claim 12 , wherein the instructions further comprise instructions for changing the cluster of shots associated with a node in the intermediate level based at least upon user feedback.
15 . The processor-accessible memory system of claim 13 , wherein the instructions further comprise:
instructions for receiving a selection of a node; and instructions for initiating playback of the shot(s) that correspond(s) to the selected node.
16 . The processor-accessible memory system of claim 13 , wherein the instructions further comprise instructions for editing the video sequence based at least upon user interaction with the tree representation.
17 . The processor-accessible memory system of claim 13 , wherein the instructions further comprise:
instructions for receiving a selection of a node; instructions for retrieving the shot(s) corresponding to the selected node; and instructions for storing the retrieved shot(s) in a computer-readable memory separate from the video sequence.
18 . A system comprising:
a data processing system; and a memory system communicatively connected to the data processing system and storing instructions configured to cause the data processing system to implement a method for representing contents of a video sequence, the video sequence including individual shots, and each shot including individual frames, wherein the instructions comprise: instructions for generating a tree representation of the video sequence; and instructions for storing the tree representation in a computer readable memory, wherein the tree representation including levels of nodes, wherein the levels include a top level including only a single complete-video node, a bottom level including a plurality of leaf nodes, and an intermediate level including a plurality of nodes fewer in number than the plurality of leaf nodes, wherein each node corresponds to one or more of the shots and each node is represented in the tree representation by a frame of its corresponding shot(s), wherein each leaf node corresponds to one of the shots, wherein the complete-video node corresponds to all of the shots in the video sequence, and wherein each node in the intermediate level corresponds to a cluster of shots, but not all of the shots in the video sequence.
19 . The system of claim 18 , wherein the instructions further comprise:
instructions for displaying the tree representation on a display; and instructions for editing the video sequence based at least upon user interaction with the tree representation.
20 . The system of claim 13 , wherein the instructions further comprise:
instructions for displaying the tree representation on a display; instructions for receiving a selection of a node; instructions for retrieving the shot(s) corresponding to the selected node; and instructions for storing the retrieved shot(s) in a computer-readable memory separate from the video sequence.Join the waitlist — get patent alerts
Track US2008059885A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.