US2022245438A1PendingUtilityA1

Deep learning hardware

Assignee: INTEL CORPPriority: Dec 30, 2016Filed: Apr 25, 2022Published: Aug 4, 2022
Est. expiryDec 30, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/098G06N 3/0464G06N 3/0455G06N 3/08G06N 3/063G06N 3/04G06F 17/16G06F 9/5027
72
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A network of matrix processing units (MPUs) is provided on a device, where each MPU is connected to at least one other MPU in the network, and each MPU is to perform matrix multiplication operations. Computer memory stores tensor data and a master control central processing unit (MCC) is provided on the device to receive an instruction from a host device, where the instruction includes one or more tensor operands based on the tensor data. The MCC invokes a set of operations on one or more of the MPUs based on the instruction, where the set of operations includes operations on the tensor operands. A result is generated from the set of operations, the result embodied as a tensor value.

Claims

exact text as granted — not AI-modified
1 . An apparatus comprising:
 a plurality of matrix processing units (MPUs), wherein each MPU is to perform matrix multiplication operations;   a memory to store tensor data including matrix data;   at least one processor to:
 cause the matrix data of the tensor data to be partitioned into a plurality of partitions, wherein the matrix data is partitioned based on a hardware size of the apparatus; 
 cause the MPUs to operate on the partitioned matrix data to generate output data; 
 store the output data. 
   
     
     
         2 . The apparatus of  claim 1 , wherein the memory comprises a memory resource block to be shared by two or more MPUs in the plurality of MPUs. 
     
     
         3 . The apparatus of  claim 1 , wherein the output data includes a tensor value. 
     
     
         4 . The apparatus of  claim 1 , wherein the MPU implements a recurrent neural network. 
     
     
         5 . The apparatus of  claim 1 , wherein the MPU is implemented by a field programmable gate array. 
     
     
         6 . The apparatus of  claim 1 , further including a control processor to manage plurality of MPUs. 
     
     
         7 . The apparatus of  claim 1 , wherein the tensor data includes a single exponent value for values in the tensor data. 
     
     
         8 . A non-transitory computer readable medium comprising instructions that, when executed, cause a machine to at least:
 cause matrix data of tensor data to be partitioned into a plurality of partitions, wherein the matrix data is partitioned based on a hardware size of system including matrix processing units (MPUs);   cause a plurality of the MPUs to operate on the partitioned matrix data to generate output data, wherein each MPU is to perform matrix multiplication operations;   store the output data.   
     
     
         9 . The non-transitory computer readable medium of  claim 8 , wherein two or more MPUs in the plurality of MPUs share a memory resource block. 
     
     
         10 . The non-transitory computer readable medium of  claim 8 , wherein the output data includes a tensor value. 
     
     
         11 . The non-transitory computer readable medium of  claim 8 , wherein the MPU implements a recurrent neural network. 
     
     
         12 . The non-transitory computer readable medium of  claim 8 , wherein the MPU is implemented by a field programmable gate array. 
     
     
         13 . The non-transitory computer readable medium of  claim 8 , wherein the tensor data includes a single exponent value for values in the tensor data. 
     
     
         14 . The non-transitory computer readable medium of  claim 8 , wherein the machine is a component in a cloud computing system. 
     
     
         15 . A method comprising:
 causing matrix data of tensor data to be partitioned into a plurality of partitions, wherein the matrix data is partitioned based on a hardware size of a system including a plurality of matrix processing units (MPUs);   causing a plurality of the MPUs to operate on the partitioned matrix data to generate output data, wherein each MPU is to perform matrix multiplication operations;   storing the output data.   
     
     
         16 . The method of  claim 15 , wherein two or more MPUs in the plurality of MPUs share a memory resource block. 
     
     
         17 . The method of  claim 15 , wherein the output data includes a tensor value. 
     
     
         18 . The method of  claim 15 , wherein the MPU implements a recurrent neural network. 
     
     
         19 . The method of  claim 15 , wherein the MPU is implemented by a field programmable gate array. 
     
     
         20 . The method of  claim 15 , wherein the tensor data includes a single exponent value for values in the tensor data.

Join the waitlist — get patent alerts

Track US2022245438A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.