US2023297819A1PendingUtilityA1
Processor array for processing sparse binary neural networks
Est. expirySep 25, 2039(~13.2 yrs left)· nominal 20-yr term from priority
Inventors:Ram KrishnamurthyGregory K. ChenRaghavan KumarPhil KnagHuseyin Ekin SumbulDeepak Kadetotad
G06N 3/09G06N 3/0499G06N 3/0495G06N 3/063G06N 3/084G06N 3/048G06F 17/16G06N 3/04G06F 7/5443
72
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
An apparatus is described. The apparatus includes a circuit to process a binary neural network. The circuit includes an array of processing cores, wherein, processing cores of the array of processing cores are to process different respective areas of a weight matrix of the binary neural network. The processing cores each include add circuitry to add only those weights of an i layer of the binary neural network that are to be effectively multiplied by a non zero nodal output of an i−1 layer of the binary neural network.
Claims
exact text as granted — not AI-modified1 . An apparatus, comprising:
a plurality of processing cores, each processing core of the plurality of processing cores comprising multiply circuitry, the multiply circuitry to multiply respective output values of a first respective subset of nodes of an (i−1)th layer of a neural network with respective weights of connections that flow from the first respective subset of nodes of the (i−1)th layer of the neural network to a second respective subset of nodes of an ith layer of the neural network, each processing core of the plurality of processing cores comprising accumulate circuitry, the accumulate circuitry to accumulate respective product terms generated by the multiply circuitry for a same node of the second respective subset of nodes of the ith layer of the neural network, wherein, the combination of first respective subset and second respective subset are unique for each of the processing cores, wherein, each of the processing cores comprise a buffer to store output values from the core's respective first subset of nodes of the (i−1)th layer and read circuitry to selectively read weights of connections for those of the output values in the buffer that are non-zero.
2 . The apparatus of claim 1 wherein the neural network is a binary neural network.
3 . The apparatus of claim 2 wherein the weights of connections that are to be read by the read circuitry are to have a value of 1 or −1.
4 . The apparatus of claim 1 wherein the plurality of processors are components of an accelerator.
5 . The apparatus of claim 4 wherein the accelerator is integrated on a semiconductor chip comprising general purpose CPU processing cores, the accelerator to be invoked by software that executes on at least one of the general purpose CPU processing cores.
6 . The apparatus of claim 1 wherein the plurality of processing cores are integrated into an execution unit of an instruction execution pipeline of a general purpose CPU processing core.
7 . The apparatus of claim 1 wherein the read circuitry is to selectively read the weights of connections from a memory.
8 . A computing system, comprising:
a plurality of general purpose CPU processing cores; a network, the plurality of general purpose CPU processing cores coupled to the network; an accelerator coupled to the network, the accelerator comprising a plurality of processing cores, each processing core of the plurality of processing cores comprising multiply circuitry, the multiply circuitry to multiply respective output values of a first respective subset of nodes of an (i−1)th layer of a neural network with respective weights of connections that flow from the first respective subset of nodes of the (i−1)th layer of the neural network to a second respective subset of nodes of an ith layer of the neural network, each processing core of the plurality of processing cores comprising accumulate circuitry, the accumulate circuitry to accumulate respective product terms generated by the multiply circuitry for a same node of the second respective subset of nodes of the ith layer of the neural network, wherein, the combination of first respective subset and second respective subset are unique for each of the processing cores, wherein, each of the processing cores comprise a buffer to store output values from the core's respective first subset of nodes of the (i−1)th layer and read circuitry to selectively read weights of connections for those of the output values in the buffer that are non-zero.
9 . The computing system of claim 8 wherein the neural network is a binary neural network.
10 . The computing system of claim 9 wherein the weights of connections that are to be read by the read circuitry have a value of 1 or −1.
11 . The computing system of claim 8 wherein the plurality of general purpose CPU processing cores, the network and the accelerator are integrated on a same semiconductor chip.
12 . The computing system of claim 11 wherein the accelerator is to be invoked by software that executes on at least one of the general purpose CPU processing cores.
13 . The computing system of claim 8 wherein the read circuitry is to selectively read the weights of connections from a memory.
14 . A method, comprising:
storing respective output values of a first respective subset of nodes of an (i−1)th layer of a neural network into a buffer; selectively reading respective weights for those of the output values in the buffer that are non-zero, the weights for respective connections of the neural network that flow into a second respective subset of nodes of an ith layer of the neural network; multiplying the respective output values with their respective weights to generate product terms; accumulating those of the product terms for a same one of the second respective subset of nodes of an ith layer of the neural network; and, performing the above for different first respective subset of nodes of the (i−1)th layer and second respective subset of nodes of the ith layer combinations.Join the waitlist — get patent alerts
Track US2023297819A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.