US2021248377A1PendingUtilityA1
4d convolutional neural networks for video recognition
Est. expiryFeb 6, 2040(~13.5 yrs left)· nominal 20-yr term from priority
G06V 20/44G06V 40/20G06V 10/82G06V 10/764G06V 20/41G06N 3/08G06N 3/045G06N 3/09G06N 3/0464G06V 20/20G06N 20/10G06T 7/0002G06T 7/10G06T 2207/10016G06T 2207/20081G06T 2207/20084G06K 9/00671G06T 3/0087G06K 9/00718G06T 3/16
61
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
This disclosure includes technologies for video recognition in general. The disclosed system can automatically detect various types of actions in a video, including reportable actions that cause shrinkage in a practical application for loss prevention in the retail industry. The temporal evolution of spatio-temporal features in the video are used for action recognition. Such features may be learned via a 4D convolutional operation, which is adapted to model low-level features based on a residual 4D block. Further, appropriate responses may be invoked if a reportable action is recognized.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for action recognition, comprising:
a video camera configured to record a video in a store; an action localizer, operationally coupled to the video camera, to divide the video to a plurality of sections and select respective random snippets from the plurality of sections as a plurality of action units; and an action recognizer, operationally coupled to the action localizer, to recognize a concealment action to conceal a product in the store based on a video-level representation of temporal evolution features of the plurality of action units.
2 . The system of claim 1 , further comprising:
an action trainer, operationally coupled to the action recognizer, to extract respective three-dimensional (3D) features of the plurality of action units of a training video with at least one concealment action, and model a video-level representation of the training video based on a temporal evolution of the respective 3D features of the plurality of action units of the training video.
3 . The system of claim 2 , wherein the action recognizer is to recognize the concealment action in the video from a prediction score generated based on a comparison of the video-level representation of temporal evolution features of the plurality of action units and the video-level representation of the training video.
4 . The system of claim 1 , further comprising:
an action responder, operationally coupled to the action recognizer, to generate a message with one or more action units of the video, and distribute the message to one or more designated devices.
5 . The system of claim 1 , further comprising:
a three-dimensional (4D) kernel with a first dimension of a number of action units, a second dimension of a temporal length of an action unit, a third dimension of a height of the action unit, and a fourth dimension of a width of the action unit; and wherein the action recognizer is to generate, based on the 4D kernel in a 4D convolutional operation, the video-level representation of temporal evolution features of the plurality of action units.
6 . The system of claim 1 , further comprising:
a plurality of 3D layers in a convolutional network to generate respective 3D features of the plurality of action units; and a residual 4D block with a residual structure to model interactions of the respective 3D features of the plurality of action units.
7 . The system of claim 1 , further comprising:
a 4D convolution block including a plurality of 3D layers and a residual 4D block, the plurality of 3D layers to generate respective 3D features of the plurality of action units and the residual 4D block to model both short-range and long-range temporal structural features based on the respective 3D features of the plurality of action units.
8 . The system of claim 7 , further comprising:
a residual connection in the residual 4D block, to enable a short-range temporal structural feature and a long-range temporal structural feature to be learned jointly.
9 . A computer-implemented method for action recognition, comprising:
dividing a video to a plurality of sections; selecting respective random snippets from the plurality of sections to form a plurality of action units; and recognizing an action in the video based on a video-level representation of temporal evolution features of the plurality of action units.
10 . The method of claim 9 , wherein the dividing comprises dividing the video into the plurality of sections with a uniform length.
11 . The method of claim 9 , wherein the selecting comprises selecting one or more action units from each of the plurality of sections, each action unit having a predetermined length.
12 . The method of claim 9 , further comprising:
constructing the video-level representation of temporal evolution features of the plurality of action units based on both local and global 3D spatio-temporal features of the plurality of action units.
13 . The method of claim 9 , further comprising:
providing the plurality of action units in parallel to respective 3D convolutional layers with weight sharing in a convolutional operation; outputting respective 3D features from the respective 3D convolutional layers to a residual 4D block; and modeling the video-level representation of temporal evolution features of the plurality of action units based on the residual 4D block.
14 . The method of claim 13 , further comprising:
pooling convolutional features of the plurality of action units to construct the video-level representation of temporal evolution features of the plurality of action units.
15 . The method of claim 9 , further comprising:
classifying the video into a reportable type based on the action in the video; and generating a message in response to the video being classified into the reportable type.
16 . A computer-readable storage device encoded with instructions that, when executed, cause one or more processors of a computing system to perform operations of action recognition, comprising:
dividing a video to a plurality of sections; forming a plurality of action units from random selected snippets from the plurality of sections; and recognizing, via a residual 4D block in a neural network, an action in the video based on both short-range temporal structural features and long-range temporal structural features of the plurality of action units selected from the plurality of sections of the video.
17 . The computer-readable storage device of claim 16 , wherein the instructions that, when executed, cause the one or more processors to perform further operations comprising:
preserving 3D spatio-temporal representations between two 3D convolutional layers in the neural network with a residual connection in the residual 4D block.
18 . The computer-readable storage device of claim 16 , wherein the short-range temporal structural features comprises 3D features of an action unit, the long-range temporal structural features comprises evolution features of respective 3D features of the plurality of action units.
19 . The computer-readable storage device of claim 16 , wherein the instructions that, when executed, cause the one or more processors to perform further operations comprising:
aggregating, via the residual 4D block, respective 3D features of the plurality of action units, for video-level action recognition.
20 . The computer-readable storage device of claim 16 , wherein the instructions that, when executed, further cause the one or more processors to perform operations comprising:
implementing a 4D convolutional operation with the residual 4D block based on a summation of a plurality of 3D convolutional operations in a dimension of action unit.Join the waitlist — get patent alerts
Track US2021248377A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.