US2019066732A1PendingUtilityA1

Video Skimming Methods and Systems

Assignee: VID SCALE INCPriority: Aug 6, 2010Filed: Oct 25, 2018Published: Feb 28, 2019
Est. expiryAug 6, 2030(~4 yrs left)· nominal 20-yr term from priority
G11B 27/034G06K 9/00751G11B 27/031G11B 27/28G06V 20/47
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In an embodiment, an apparatus and method of creating a skimming preview of a video includes electronically receiving a plurality of video shots, analyzing each frame in a video shot from the plurality of video shots, where analyzing includes determining a saliency of each frame of the video shot. The method also includes determining a key frame of the video shot based on the saliency of each frame the video shot, extracting visual features from the key frame, performing shot clustering of the plurality of video shots to determine concept patterns based on the visual features, and generating a reconstruction reference tree based on the shot clustering. The reconstruction reference tree includes video shots categorized according to each concept pattern.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . An apparatus comprising:
 a processor;   a memory coupled to the processor;   a port coupled to the processor to electronically receive a plurality of video shots; and   a non-transitory computer-readable medium storing instructions that are operative, when executed by the processor to perform acts including:
 analyzing each frame in a video shot from the plurality of video shots, the analyzing comprising determining a saliency of each frame of the video shot, the saliency being a content attentiveness saliency providing a measurement of representative shot properties; 
 determining an effective visual saliency based on the determined saliency, the effective visual saliency based on camera motion of the plurality of video shots and a human attention model of the video shot; 
 selecting a key frame of the video shot based on the effective visual saliency of each frame the video shot; 
 extracting visual features from the key frame; performing shot clustering of the plurality of video shots to determine concept patterns based on the visual features; and 
 generating a hierarchical reconstruction based on the shot clustering, the hierarchical reconstruction enabling a skimming preview of the video. 
   
     
     
         2 . The apparatus of  claim 1 , wherein the human attention model is based on camera motion. 
     
     
         3 . The apparatus of  claim 2 , wherein:
 the human attention model comprises a camera attenuation factor based on the camera motion,   wherein the processor to performs acts further including:
 determining the camera attenuation factor; and 
 determining the effective visual saliency comprises multiplying the determined effective visual saliency with the determined camera attenuation factor. 
   
     
     
         4 . The apparatus of  claim 3 , wherein the determined camera attenuation factor is proportional to a zooming speed of video shot from the plurality of video shots. 
     
     
         5 . The apparatus of  claim 3 , wherein the determined camera attenuation factor is inversely proportional to a panning speed of video shot from the plurality of video shots. 
     
     
         6 . The apparatus of  claim 1 , wherein the hierarchical reconstruction comprises a reconstruction reference tree that includes video shots categorized according to each concept pattern. 
     
     
         7 . The apparatus of  claim 6 , wherein generating the reconstruction reference tree comprises categorizing video shots within concept categories ordered according to concept importance, and ordering video shots within each concept category according to effective visual saliency. 
     
     
         8 . The apparatus of  claim 7 , wherein the concept importance includes determining a total number of frames in a shot having a same concept. 
     
     
         9 . The apparatus of  claim 6 , further comprising wherein the hierarchical reconstruction includes generating a video skimming preview based on the reconstruction reference tree. 
     
     
         10 . The apparatus of  claim 1 , wherein the processor performs acts further including extracting audio features from the video shot. 
     
     
         11 . The apparatus of  claim 10 , wherein extracting audio features comprises:
 determining audio words from the video shot; and   performing clustering on the audio words.   
     
     
         12 . The apparatus of  claim 11 , further comprising:
 determining visual concept patterns based on the performing shot clustering; and   determining audio concept patterns based on the performing clustering on the audio words.   
     
     
         13 . The apparatus of  claim 12 , further comprising:
 calculating a number of member shots for each visual concept of the visual concept patterns and for each audio concept of the audio concept patterns;   sorting each visual concept by calculated number of shots;   sorting each audio concept by calculated number of shots; and   aligning visual concepts and audio concepts having a same number of shots.   
     
     
         14 . A method comprising:
 electronically receiving a reconstruction reference tree comprising video shots categorized within concept categories ordered according to concept importance, wherein video shots within each concept category is ordered according to saliency;   selecting shots starting from categories of highest importance and shots of highest saliency within the categories of highest importance; and   generating a preview based on the selected shots.   
     
     
         15 . The method of  claim 14 , further comprising:
 selecting frames within the selected shots having a highest saliency; and   using the frames within the selected shots having the highest saliency to generate the preview.   
     
     
         16 . The method of  claim 15 , wherein selecting frames within the selected shots having a highest saliency comprises:
 selecting a target skimming ratio;   determining a threshold according to the selected target skimming ratio; and comparing the saliency of the selected shots to the threshold.   
     
     
         17 . The method of  claim 16 , wherein the selected target skilling ratio is an arbitrary length. 
     
     
         18 . The method of  claim 14 , wherein the saliency comprises an effective visual saliency based on camera motion of the video shots. 
     
     
         19 . A non-transitory computer readable medium with an executable program stored thereon, wherein the program instructs a microprocessor to perform the following steps:
 analyzing each frame in a video shot from a plurality of video shots, the analyzing including determining a saliency of each frame of the video shot, the saliency being a content attentiveness saliency;   determining an effective visual saliency based on the determined saliency and based on camera motion of each from the video shot;   selecting a key frame of the video shot based on the effective visual saliency of each frame of the video shot;   extracting visual features from the key frame;   performing shot clustering of the plurality of video shots to determine concept patterns based on the visual features; and   generate a reconstruction reference tree based on the shot clustering, the reconstruction reference tree comprising video shots categorized according to each concept pattern.   
     
     
         20 . The non-transitory computer readable medium of  claim 19 , wherein the program instructs the microprocessor to further perform the steps of:
 determining audio features of the video shot;   determining saliency of the determined audio features;   clustering determined audio features; and   aligning audio and video concept categories.

Join the waitlist — get patent alerts

Track US2019066732A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.