US2021248377A1PendingUtilityA1

4d convolutional neural networks for video recognition

Assignee: SHENZHEN MALONG TECH CO LTDPriority: Feb 6, 2020Filed: Jun 9, 2020Published: Aug 12, 2021
Est. expiryFeb 6, 2040(~13.5 yrs left)· nominal 20-yr term from priority
G06V 20/44G06V 40/20G06V 10/82G06V 10/764G06V 20/41G06N 3/08G06N 3/045G06N 3/09G06N 3/0464G06V 20/20G06N 20/10G06T 7/0002G06T 7/10G06T 2207/10016G06T 2207/20081G06T 2207/20084G06K 9/00671G06T 3/0087G06K 9/00718G06T 3/16
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure includes technologies for video recognition in general. The disclosed system can automatically detect various types of actions in a video, including reportable actions that cause shrinkage in a practical application for loss prevention in the retail industry. The temporal evolution of spatio-temporal features in the video are used for action recognition. Such features may be learned via a 4D convolutional operation, which is adapted to model low-level features based on a residual 4D block. Further, appropriate responses may be invoked if a reportable action is recognized.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system for action recognition, comprising:
 a video camera configured to record a video in a store;   an action localizer, operationally coupled to the video camera, to divide the video to a plurality of sections and select respective random snippets from the plurality of sections as a plurality of action units; and   an action recognizer, operationally coupled to the action localizer, to recognize a concealment action to conceal a product in the store based on a video-level representation of temporal evolution features of the plurality of action units.   
     
     
         2 . The system of  claim 1 , further comprising:
 an action trainer, operationally coupled to the action recognizer, to extract respective three-dimensional (3D) features of the plurality of action units of a training video with at least one concealment action, and model a video-level representation of the training video based on a temporal evolution of the respective 3D features of the plurality of action units of the training video.   
     
     
         3 . The system of  claim 2 , wherein the action recognizer is to recognize the concealment action in the video from a prediction score generated based on a comparison of the video-level representation of temporal evolution features of the plurality of action units and the video-level representation of the training video. 
     
     
         4 . The system of  claim 1 , further comprising:
 an action responder, operationally coupled to the action recognizer, to generate a message with one or more action units of the video, and distribute the message to one or more designated devices.   
     
     
         5 . The system of  claim 1 , further comprising:
 a three-dimensional (4D) kernel with a first dimension of a number of action units, a second dimension of a temporal length of an action unit, a third dimension of a height of the action unit, and a fourth dimension of a width of the action unit; and wherein the action recognizer is to generate, based on the 4D kernel in a 4D convolutional operation, the video-level representation of temporal evolution features of the plurality of action units.   
     
     
         6 . The system of  claim 1 , further comprising:
 a plurality of 3D layers in a convolutional network to generate respective 3D features of the plurality of action units; and   a residual 4D block with a residual structure to model interactions of the respective 3D features of the plurality of action units.   
     
     
         7 . The system of  claim 1 , further comprising:
 a 4D convolution block including a plurality of 3D layers and a residual 4D block, the plurality of 3D layers to generate respective 3D features of the plurality of action units and the residual 4D block to model both short-range and long-range temporal structural features based on the respective 3D features of the plurality of action units.   
     
     
         8 . The system of  claim 7 , further comprising:
 a residual connection in the residual 4D block, to enable a short-range temporal structural feature and a long-range temporal structural feature to be learned jointly.   
     
     
         9 . A computer-implemented method for action recognition, comprising:
 dividing a video to a plurality of sections;   selecting respective random snippets from the plurality of sections to form a plurality of action units; and   recognizing an action in the video based on a video-level representation of temporal evolution features of the plurality of action units.   
     
     
         10 . The method of  claim 9 , wherein the dividing comprises dividing the video into the plurality of sections with a uniform length. 
     
     
         11 . The method of  claim 9 , wherein the selecting comprises selecting one or more action units from each of the plurality of sections, each action unit having a predetermined length. 
     
     
         12 . The method of  claim 9 , further comprising:
 constructing the video-level representation of temporal evolution features of the plurality of action units based on both local and global 3D spatio-temporal features of the plurality of action units.   
     
     
         13 . The method of  claim 9 , further comprising:
 providing the plurality of action units in parallel to respective 3D convolutional layers with weight sharing in a convolutional operation;   outputting respective 3D features from the respective 3D convolutional layers to a residual 4D block; and   modeling the video-level representation of temporal evolution features of the plurality of action units based on the residual 4D block.   
     
     
         14 . The method of  claim 13 , further comprising:
 pooling convolutional features of the plurality of action units to construct the video-level representation of temporal evolution features of the plurality of action units.   
     
     
         15 . The method of  claim 9 , further comprising:
 classifying the video into a reportable type based on the action in the video; and   generating a message in response to the video being classified into the reportable type.   
     
     
         16 . A computer-readable storage device encoded with instructions that, when executed, cause one or more processors of a computing system to perform operations of action recognition, comprising:
 dividing a video to a plurality of sections;   forming a plurality of action units from random selected snippets from the plurality of sections; and   recognizing, via a residual 4D block in a neural network, an action in the video based on both short-range temporal structural features and long-range temporal structural features of the plurality of action units selected from the plurality of sections of the video.   
     
     
         17 . The computer-readable storage device of  claim 16 , wherein the instructions that, when executed, cause the one or more processors to perform further operations comprising:
 preserving 3D spatio-temporal representations between two 3D convolutional layers in the neural network with a residual connection in the residual 4D block.   
     
     
         18 . The computer-readable storage device of  claim 16 , wherein the short-range temporal structural features comprises 3D features of an action unit, the long-range temporal structural features comprises evolution features of respective 3D features of the plurality of action units. 
     
     
         19 . The computer-readable storage device of  claim 16 , wherein the instructions that, when executed, cause the one or more processors to perform further operations comprising:
 aggregating, via the residual 4D block, respective 3D features of the plurality of action units, for video-level action recognition.   
     
     
         20 . The computer-readable storage device of  claim 16 , wherein the instructions that, when executed, further cause the one or more processors to perform operations comprising:
 implementing a 4D convolutional operation with the residual 4D block based on a summation of a plurality of 3D convolutional operations in a dimension of action unit.

Join the waitlist — get patent alerts

Track US2021248377A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.