Neural network accelerator
Abstract
An embodiment of the present application discloses a neural network accelerator, including: a convolution calculation module, which is used to perform a convolution operation on an input data input into a preset neural network to obtain a first output data; a tail calculation module, which is used to perform a calculation on the first output data to obtain a second output data; a storage module, which is used to cache the input data and the second output data; and a first control module, which is used to transmit the first output data to the tail calculation module. The convolution calculation module includes a plurality of convolution calculation units, the tail calculation module includes a plurality of tail calculation units, the first control module includes a plurality of first control units, and at least two convolution calculation units are connected to one tail calculation unit through one first control unit.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A neural network accelerator, comprising:
a convolution calculation module used to perform a convolution operation on an input data input into a preset neural network to obtain a first output data; a tail calculation module used to perform a calculation on the first output data to obtain a second output data; a storage module used to cache the input data and the second output data; and a first control module used to transmit the first output data to the tail calculation module; wherein the convolution calculation module includes a plurality of convolution calculation units, the tail calculation module includes a plurality of tail calculation units, the first control module includes a plurality of first control units, and at least two convolution calculation units are connected to one tail calculation unit through one first control unit.
2 . The neural network accelerator according to claim 1 , further comprising a second control module used to transmit the output data calculated by the neural network to the storage module, the second control module comprising a plurality of second control units, and at least one tail calculation unit being connected to the storage module through one second control unit.
3 . The neural network accelerator according to claim 1 , wherein a data flow rate of the convolution calculation module is less than or equal to a data flow rate of the tail calculation module.
4 . The neural network accelerator according to claim 1 , wherein a sum of on-chip resources consumed by the convolution calculation module and the tail calculation module is less than or equal to a total on-chip resource.
5 . The neural network accelerator according to claim 1 , further comprising: a preset parameter configuration module used to configure preset parameters, the preset parameters comprising a convolution kernel size, an input feature map size, an input data storage location and a second output data storage location.
6 . The neural network accelerator according to claim 5 , wherein each convolution calculation unit comprises a weight value unit, an input feature map unit, and a convolution kernel;
the weight value unit is used to form a corresponding weight value according to the convolution kernel size; the input feature map unit is used to obtain the input data from the storage module according to the input feature map size and the input data storage location to form a corresponding input feature map; the convolution kernel is used to perform a calculation on the weight value and the input feature map.
7 . The neural network accelerator according to claim 6 , wherein each convolution calculation unit is used to perform the calculation on the weight value and the input feature map to obtain the first output data.
8 . The neural network accelerator according to claim 5 , wherein the storage module comprises an on-chip memory and/or an off-chip memory.
9 . The neural network accelerator according to claim 8 , wherein when the input data storage location is an off-chip memory, the input data in the off-chip memory is transmitted to the on-chip memory through a DMA.
10 . The neural network accelerator according to claim 8 , wherein when the second output data storage location is an off-chip memory, the second output data is transmitted to the off-chip memory by a DMA.Join the waitlist — get patent alerts
Track US2023128421A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.