US2022121917A1PendingUtilityA1

Hardware accelerator template and design framework for implementing recurrent neural networks

Assignee: INTEL CORPPriority: Dec 31, 2016Filed: Dec 28, 2021Published: Apr 21, 2022
Est. expiryDec 31, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06N 3/044G06N 3/0495G06N 3/0442G06N 3/063G06N 3/0445
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Hardware accelerator templates and design frameworks for implementing recurrent neural networks (RNNs) and variants thereof are described. A design framework module obtains a flow graph for an RNN algorithm. The flow graph identifies operations to be performed to implement the RNN algorithm and further identifies data dependencies between ones of the operations. The operations include matrix operations and vector operations. The design framework module maps the operations of the flow graph to an accelerator hardware template, yielding an accelerator instance comprising register transfer language code that describes how one or more matrix processing units and one or more vector processing units are to be arranged to perform the RNN algorithm. At least one of the one or more MPUs, as part of implementing the RNN algorithm, is to directly provide or directly receive a value from one of the one or more VPUs.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for implementing a recurrent neural network (RNN) algorithm by a computer device, comprising:
 obtaining, by an automated framework for implementing a long-short term memory (LSTM) RNN, a flow graph for the LSTM RNN, the flow graph representing a plurality of matrix and vector operations to be performed to implement the LSTM RNN and having data dependencies among the plurality of operations;   tuning, by the automated framework, design parameters to optimize a hardware template for the flow graph;   validating, by the automated framework, performance of a target field-programmable gate array (FPGA) based on the optimized design parameters; and   mapping, by the automated framework, the flow graph to the hardware template based on the optimized design parameters to generate register transfer language (RTL) code for the LSTM RNN on the target FPGA.   
     
     
         2 . The method of  claim 1 , wherein the obtaining comprises computing, by the automated framework, the flow graph based upon a plurality of equations corresponding to the RNN algorithm. 
     
     
         3 . The method of  claim 1 , wherein the mapping is based upon hardware design constraints indicating amounts or capabilities of hardware elements that can be utilized in the FPGA. 
     
     
         4 . The method of  claim 1 , wherein the mapping is based upon optimization goals indicating properties of the FPGA that should be optimized for. 
     
     
         5 . The method of  claim 1 , wherein the mapping is based upon one or more dataset properties identifying properties of input data for the RNN algorithm to be used with the FPGA. 
     
     
         6 . The method of  claim 1 , wherein the mapping further yields a compiler that is executable to program the FPGA to execute micro-code to implement the RNN algorithm. 
     
     
         7 . The method of  claim 6 , wherein the compiler is to program the FPGA by causing a control unit of the FPGA to execute at least some of the micro-code. 
     
     
         8 . The method of  claim 1 , further comprising validating a performance of and functionalities of the generated FPGA against one or more performance and functional models derived from hardware design constraints and optimization goals. 
     
     
         9 . The method of  claim 1 , further comprising providing the RTL code to be used as an input to a logic synthesis tool to yield a circuit design for an Application-Specific Integrated Circuit (ASIC). 
     
     
         10 . A non-transitory machine-readable storage medium having instructions which, when executed by one or more processors of a device, cause the device to implement a recurrent neural network (RNN) algorithm, comprising:
 obtaining, by an automated framework for implementing a long-short term memory (LSTM) RNN, a flow graph for the LSTM RNN, the flow graph representing a plurality of matrix and vector operations to be performed to implement the LSTM RNN and having data dependencies among the plurality of operations;   tuning, by the automated framework, design parameters to optimize a hardware template for the flow graph;   validating, by the automated framework, performance of a target field-programmable gate array (FPGA) based on the optimized design parameters; and   mapping, by the automated framework, the flow graph to the hardware template based on the optimized design parameters to generate register transfer language (RTL) code for the LSTM RNN on the target FPGA.   
     
     
         11 . The non-transitory machine-readable storage medium of  claim 10 , wherein the obtaining comprises computing the flow graph based upon a plurality of equations corresponding to the RNN algorithm. 
     
     
         12 . The non-transitory machine-readable storage medium of  claim 10 , wherein the mapping is based upon hardware design constraints indicating amounts or capabilities of hardware elements that can be utilized in the FPGA. 
     
     
         13 . The non-transitory machine-readable storage medium of  claim 10 , wherein the mapping is based upon optimization goals indicating properties of the FPGA that should be optimized for. 
     
     
         14 . The non-transitory machine-readable storage medium of  claim 10 , wherein the mapping is based upon one or more dataset properties identifying properties of input data for the RNN algorithm to be used with the FPGA. 
     
     
         15 . The non-transitory machine-readable storage medium of  claim 10 , wherein the mapping further yields a compiler that is executable to program the FPGA to execute micro-code to implement the RNN algorithm. 
     
     
         16 . The non-transitory machine-readable storage medium of  claim 15 , wherein the compiler is to program the FPGA by causing a control unit of the FPGA to execute at least some of the micro-code. 
     
     
         17 . The non-transitory machine-readable storage medium of  claim 10 , wherein the operations further comprise validating a performance of and functionalities of the FPGA against one or more performance and functional models derived from hardware design constraints and optimization goals. 
     
     
         18 . The non-transitory machine-readable storage medium of  claim 17 , wherein the operations further comprise providing the RTL code to be used as an input to a logic synthesis tool to yield a circuit design for an Application-Specific Integrated Circuit (ASIC). 
     
     
         19 . A device comprising:
 one or more processors; and   one or more non-transitory machine-readable storage media having instructions which, when executed by the one or more processors, cause the device to implement to implement a recurrent neural network (RNN) algorithm, comprising:
 obtaining, by an automated framework for implementing a long-short term memory (LSTM) RNN, a flow graph for the LSTM RNN, the flow graph representing a plurality of matrix and vector operations to be performed to implement the LSTM RNN and having data dependencies among the plurality of operations; 
 tuning, by the automated framework, design parameters to optimize a hardware template for the flow graph; 
 validating, by the automated framework, performance of a target field-programmable gate array (FPGA) based on the optimized design parameters; and 
 mapping, by the automated framework, the flow graph to the hardware template based on the optimized design parameters to generate register transfer language (RTL) code for the LSTM RNN on the target FPGA. 
   
     
     
         20 . The device of  claim 19 , wherein obtaining comprises computing, by the automated framework, the flow graph based upon a plurality of equations corresponding to the RNN algorithm.

Join the waitlist — get patent alerts

Track US2022121917A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.