Training neural network accelerators using mixed precision data formats
Abstract
Technology related to training a neural network accelerator using mixed precision data formats is disclosed. In one example of the disclosed technology, a neural network accelerator is configured to accelerate a given layer of a multi-layer neural network. An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor. A result of the tensor operation can be converted from the block floating-point format to the normal-precision floating-point format. The converted result can be used to generate an output tensor of the layer of the neural network, where the output tensor is in normal-precision floating-point format.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A computing system comprising:
a computer-readable memory; and a hardware accelerator in communication with the computer-readable memory, the hardware accelerator configured, during processing using a multi-layer neural network, to:
receive an input tensor for a given layer of the multi-layer neural network;
convert the input tensor from a normal-precision floating-point format to a quantized-precision floating-point format, the quantized-precision floating-point format being a block floating-point format, wherein a first converted input tensor portion corresponding to a first portion of the input tensor comprises a first common exponent for values in the first portion of the input tensor and a first plurality of mantissa values and a second converted tensor portion corresponding to a second portion of the input tensor comprises a second common exponent value for values in the second portion of the input tensor and a second plurality of mantissa values, wherein the first common exponent is different than the second common exponent; and
perform a tensor operation using the input tensor converted to the quantized-precision floating-point format.
2 . The computing system of claim 1 , wherein the hardware accelerator is further configured to convert a result of the tensor operation from the quantized-precision floating-point format to the normal-precision floating-point format to provide a converted result in the normal-precision floating-point format.
3 . The computing system of claim 2 , wherein the hardware accelerator is further configured to perform an operation using the converted result in the normal-precision floating-point format.
4 . The computing system of claim 2 , wherein the hardware accelerator is further configured to generate an output tensor using the converted result in the normal-precision floating-point format.
5 . The computing system of claim 1 , wherein the input tensor is a two-dimensional matrix, and the quantized-precision floating-point format is a block floating-point format where a plurality of mantissa values within a given row share a common exponent, and mantissa values in different rows have different respective exponents.
6 . The computing system of claim 1 , wherein the input tensor is a convolution filter, and the quantized-precision floating-point format is a block floating-point format where a plurality of mantissa values within a spatial pixel share a common exponent.
7 . The computing system of claim 1 , wherein the tensor operation is a dot product computation.
8 . The computing system of claim 1 , wherein the tensor operation is a convolution.
9 . The computing system of claim 1 , wherein the converting the input tensor from a normal-precision floating-point format to a quantized-precision floating-point format comprises:
selecting a first bounding box, the first bounding box defining the first portion of the input tensor; and selecting a second bounding box, the second bounding box defining the second portion of the input tensor.
10 . The method of claim 9 , wherein the first bounding box is a row of a matrix of the input tensor.
11 . The method of claim 9 , wherein the first bounding box is a column of a matrix of the input tensor.
12 . A method, implemented in a computing system, comprising:
converting an input tensor for a given layer of a multi-layer neural network from a normal-precision floating-point format to converted values represented in a block floating-point format by (1) for a first portion on the input tensor, selecting a first bounding box including a first set of values expressed in the normal-precision floating-point format and where the block floating-point format uses a first common exponent for converted values of the first set of values; and (2) for a second portion of the input tensor, selecting a second bounding box comprising a second set of values expressed in the normal-precision floating point format and where the block-floating point format uses a second common exponent for converted values of the second set of values, where the second set of values is different from the first set of values and the second common exponent is different from the first common exponent; performing a tensor operation using the converted values in the input tensor converted to the block floating-point format; converting a result of the tensor operation from the block floating-point format to the normal-precision floating-point format; and using the converted result in the normal-precision floating-point format to generate an output tensor of the layer of the neural network, where the output tensor is in normal-precision floating-point format.
13 . The method of claim 12 , wherein the first bounding box is a row of a matrix of the input tensor.
14 . The method of claim 12 , wherein the first bounding box is a column of a matrix of the input tensor.
15 . The method of claim 12 , wherein converting the input tensor for the given layer from the normal-precision floating-point format to the block floating-point format comprises:
scaling mantissa values of elements of the input tensor so that integer portions of the scaled mantissas have a selected number of bits for the block floating-point format; removing fractional bits from the scaled integer portions of the mantissas; and rounding the mantissas to produce block floating-point values.
16 . One or more non-transitory computer-readable media comprising:
computer-executable instructions that, when executed by a computing device, cause the computing device to convert an input tensor for a given layer of a multi-layer neural network from a normal-precision floating-point format to a block floating-point format, by (1) for a first portion of the input tensor, selecting a first bounding box around a first set of values expressed with the normal-precision floating-point format and where the block floating-point format uses a first common exponent for converted values of the first set of values; and (2) for a second portion of the input tensor, selecting a second bounding box comprising a second set of values expressed in the normal-precision floating point format and where the block-floating point format uses a second common exponent for converted values of the second set of values, where the second set of values is different from the first set of values and the second common exponent is different from the first common exponent; and computer-executable instructions that, when executed by the computing device, cause the computing device to perform a tensor operation using the input tensor converted to the block floating-point format.
17 . The one or more non-transitory computer-readable media of claim 16 , wherein the first bounding box is a row of a matrix of the input tensor.
18 . The one or more non-transitory computer-readable media of claim 16 , wherein the first bounding box is a column of a matrix of the input tensor.
19 . The one or more non-transitory computer-readable media of claim 16 , wherein the input tensor is a two-dimensional matrix, and in the block floating-point format a plurality of mantissa values within a given row share a common exponent, and mantissa values in different rows have different respective exponents.
20 . The one or more non-transitory computer-readable media of claim 16 , wherein the input tensor is a convolution filter, and the block floating-point format is a block floating-point format where a plurality of mantissa values within a spatial pixel share a common exponent.Join the waitlist — get patent alerts
Track US2023267319A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.