Processing element and neural processing device including same
Abstract
The present disclosure discloses a processing element and a neural processing device including the processing element. The processing element includes a weight register configured to store a weight, an input activation register configured to store an input activation, a flexible multiplier configured to receive a first sub-weight of a first precision included in the weight, receive a first sub-input activation of the first precision included in the input activation, and generate result data by performing multiplication calculation of the first sub-weight and the first sub-input activation as the first precision or a second precision different from the first precision according to the first sub-weight and the first sub-input activation and a saturating adder configured to generate a partial sum by using the result data.
Claims
exact text as granted — not AI-modified1 - 4 . (canceled)
5 . A neural processing device comprising:
at least one neural core, wherein the at least one neural core includes a processing unit configured to perform calculation, and a L0 memory configured to store input/output data of the processing unit, the processing unit includes a processing element (PE) array including at least one processing element, and the PE array includes a flexible multiplier, including a first multiplier and a second multiplier, configured to receive a weight and an input activation, provide one of the first multiplier of a first precision and the second multiplier of a second precision different from the first precision, separate from the first multiplier with the weight and the input activation based on sizes of the weight and the input activation and perform multiplication calculation of the weight and the input activation, and a saturating adder configured to receive a result data for at least a part of the multiplication calculation of the weight and the input activation and generate a partial sum.
6 . The neural processing device of claim 5 , wherein the flexible multiplier performs the multiplication calculation of the weight and the input activation as the first precision using one of the first multiplier and the second multiplier, and
wherein the first precision is larger than the second precision.
7 . The neural processing device of claim 6 , wherein the flexible multiplier performs multiplication calculation of the weight and the input activation as the first precision using the first multiplier if a size of at least one of the weight and the input activation is greater than a greatest value of the second precision, and performs multiplication calculation of the weight and the input activation as the second precision using the second multiplier if a size of each of the weight and the input activation is less than or equal to the greatest value of the second precision.
8 . The neural processing device of claim 7 , wherein the weight includes a first sub-weight and a second sub-weight, the input activation includes a first sub-input activation and a second sub-input activation, and the flexible multiplier performs multiplication calculation of the first sub-weight and the first sub-input activation as one of the first precision using the first multiplier and the second precision using the second multiplier based on respective sizes of the first sub-weight and the first sub-input activation and performs multiplication calculation of the second sub-weight and the second sub-input activation as one of the first precision and the second precision based on respective sizes of the second sub-weight and the second sub-input activation.
9 . The neural processing device of claim 7 , wherein the weight includes a first sub-weight and a second sub-weight, the input activation includes a first sub-input activation and a second sub-input activation, and the flexible multiplier performs multiplication calculation of the weight and the input activation as one of the first precision using the first multiplier and the second precision using the second multiplier based on respective sizes of the first sub-weight, the second sub-weight, the first sub-input activation, and the second sub-input activation.
10 . The neural processing device of claim 5 , wherein the number of the first multiplier is k, and the number of the second multiplier is 2 k, where k is a natural number.
11 . The neural processing device of claim 5 , wherein the flexible multiplier further includes a path determination unit configured to generate a path determination signal to determine which multiplier performs the multiplication calculation of the weight and the input activation among the first multiplier and the second multiplier, based on the weight and the input activation.
12 . The neural processing device of claim 11 , wherein the flexible multiplier further includes a demultiplexer configured to provide one of the first multiplier and the second multiplier with the weight and the input activation in response to the path determination signal.
13 . A processing element comprising:
a weight register configured to store a weight; an input activation register configured to store an input activation; a flexible multiplier, including a first multiplier with a first precision and a second multiplier with a second precision, separated from the first multiplier, configured to receive a first sub-weight of a first precision included in the weight, receive a first sub-input activation of the first precision included in the input activation, provide the first sub-weight and the first sub-input activation to one of the first multiplier and the second multiplier, and generate result data of the first precision by performing multiplication calculation of the first sub-weight and the first sub-input activation using one of the first multiplier and the second multiplier; and a saturating adder configured to generate a partial sum by using the result data.
14 . The processing element of claim 13 , wherein the flexible multiplier further includes a path determination unit configured to generate a path determination signal to determine which multiplier performs the multiplication calculation of the first sub-weight and the first sub-input activation among the first multiplier and the second multiplier, based on the first sub-weight and the first sub-input activation; and
a demultiplexer configured to provide any one of the first multiplier and the second multiplier with the first sub-weight and the first sub-input activation in response to the path determination signal.
15 . The processing element of claim 14 , wherein the path determination unit generates the path determination signal as a first signal for providing the first sub-weight and the first sub-input activation to the first multiplier if a size of at least one of the first sub-weight and the first sub-input activation is greater than a predetermined first size, and
generates the path determination signal as a second signal for providing the first sub-weight and the first sub-input activation to the second multiplier if a size of each of the first sub-weight and the first sub-input activation is less than or equal to the first size.
16 . The processing element of claim 14 , wherein the path determination unit includes
a bit division logic configured to generate the first sub-weight by dividing the weight into a unit of the first precision or the second precision and generate the first sub-input activation by dividing the input activation into a unit of the first precision or the second precision in response to the calculation mode signal, a path selection logic configured to generate the path determination signal based on the calculation mode signal, the first sub-weight, and the first sub-input activation, and a conversion logic configured to convert precisions of the first sub-weight and the first sub-input activation.
17 . The processing element of claim 13 , wherein the number of the first multipliers is k, and the number of the second multipliers is 2k, where k is a natural number.
18 . The processing element of claim 13 , wherein the first precision has 2N bits, and the second precision has N bits, where N is a natural number.
19 . The processing element of claim 13 , wherein
the weight includes the first sub-weight and a second sub-weight, the input activation includes the first sub-input activation and a second sub-input activation, the flexible multiplier generates a first path determination signal based on the first sub-weight and the first sub-input activation, and generates a second path determination signal based on the second sub-weight and the second sub-input activation, and the first path determination signal and the second path determination signal are independently generated.
20 . The processing element of claim 13 , wherein
the weight includes the first sub-weight and a second sub-weight, the input activation includes the first sub-input activation and a second sub-input activation, and the flexible multiplier generates the path determination signal based on sizes of the first sub-weight, the second sub-weight, the first sub-input activation, and the second sub-input activation.
21 . The processing element of claim 13 , wherein the flexible multiplier further includes a control pipeline configured to synchronize reception of the first sub-weight and the first sub-input activation with generation of the result data.
22 . A processing element comprising:
a weight register configured to store a weight; an input activation register configured to store an input activation; a flexible multiplier, including a first multiplier with a first precision and a second multiplier with a second precision smaller than the first precision, configured to generate result data by performing multiplication calculation of the weight and the input activation as the first precision or the second precision based on a calculation mode signal; and a saturating adder configured to generate a partial sum by using the result data, wherein the flexible multiplier performs the multiplication calculation of the weight and the input activation using the second multiplier if the calculation mode signal is a first mode signal for calculation with the second precision, and wherein the flexible multiplier performs the multiplication calculation of the weight and the input activation using one of the first multiplier and the second multiplier if the calculation mode signal is a second mode signal for calculation with the first precision based on sizes of the weight and the input activation.
23 . The processing element of claim 22 , wherein the flexible multiplier further includes
an error detection logic configured to generate a detection result by checking whether overflow or underflow occurs according to the multiplication calculation of the weight and the input activation.
24 . The processing element of claim 22 , wherein the flexible multiplier provides the weight and the input activation to one of the first multiplier and the second multiplier based on whether at least one of the weight and the input activation is greater than a greatest value of the second precision, if the calculation mode signal is the second mode signal.Join the waitlist — get patent alerts
Track US2023143798A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.