US2023177108A1PendingUtilityA1

Data path for scalable matrix node engine with mixed data formats

Assignee: TESLA INCPriority: May 3, 2019Filed: Jan 13, 2023Published: Jun 8, 2023
Est. expiryMay 3, 2039(~12.8 yrs left)· nominal 20-yr term from priority
G06F 17/16G06F 7/483G06N 3/063G06F 9/30032G06F 2207/4824G06N 3/08G06F 9/3877G06N 20/00G06F 7/5443G06N 3/045G06F 9/30014G06F 7/49915G06N 3/0495G06N 3/0464G06F 9/30007G06F 9/30036
70
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A microprocessor system comprises a matrix computational unit and a control unit. The matrix computational unit includes a plurality of processing elements. The control unit is configured to provide a matrix processor instruction to the matrix computational unit. The matrix processor instruction specifies a floating-point operand formatted using a first floating-point representation format. The matrix computational unit accumulates an intermediate result value calculated using the floating-point operand. The intermediate result value is in a second floating-point representation format.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A microprocessor system, comprising:
 a plurality of matrix processors which form a node engine;   a control unit configured to provide matrix processors instructions to the matrix processors; and   a plurality of processor registers allocated to store a plurality of bit depth formats,   wherein the node engine is configured to execute a plurality of matrix processor instructions, wherein operations associated with the matrix processor instructions are interleaved for execution by the matrix processors which form the node engine,   and wherein the matrix processors are configured to execute the operations in different bit depth formats based on the associated matrix processor instructions.   
     
     
         22 . The microprocessor system of  claim 21 , wherein a first bit depth format is an 8-bit floating point format, and wherein input data associated with a particular matrix processor instruction is formatted according to the first bit depth format. 
     
     
         23 . The microprocessor system of  claim 22 , wherein a second bit depth format is a 21-bit floating point format, and wherein intermediate results associated with the particular matrix processor instruction are formatted according to the second bit depth format. 
     
     
         24 . The microprocessor system of  claim 21 , wherein a particular matrix processor instruction specifies a configurable exponent bias. 
     
     
         25 . The microprocessor system of  claim 21 , wherein the matrix processors interleave the operations based on data availability. 
     
     
         26 . The microprocessor system of  claim 21 , wherein each of the plurality of matrix processors includes a plurality of accumulators. 
     
     
         27 . The microprocessor system of  claim 26 , wherein each matrix processor uses different accumulators for operations associated with different matrix processor instructions. 
     
     
         28 . The microprocessor system of  claim 27 , wherein for a particular matrix processor, the different accumulators are designated to accumulate and store intermediate results associated with the different matrix processor instructions. 
     
     
         29 . The microprocessor system of  claim 28 , wherein the different accumulators store the intermediate results using a first bit depth format with a higher-bit floating point format than a second bit depth format used for input operands. 
     
     
         30 . The microprocessor system of  claim 26 , wherein a particular matrix processor instruction specifies a particular accumulator for operations associated with the particular matrix processor instruction. 
     
     
         31 . A microprocessor system, comprising:
 a matrix processor which forms at least part of a node engine;   a control unit configured to provide matrix processors instructions to the matrix processor; and   a plurality of processor registers allocated to store a plurality of bit depth formats,   wherein the matrix processor is configured to execute, at least in part, a plurality of matrix processor instructions, wherein the matrix processor interleaves execution of a subset of operations associated with the matrix processor instructions.   and wherein the matrix processor is configured to execute the subset of the operations in different bit depth formats based on the associated matrix processor instructions.   
     
     
         32 . The microprocessor system of  claim 31 , wherein the matrix processor interleaves the operations based on data availability. 
     
     
         33 . The microprocessor system of  claim 31 , wherein the matrix processor interleaves execution of the subset based on use of a plurality of accumulators. 
     
     
         34 . The microprocessor system of  claim 31 , wherein a particular matrix processor instruction specifies use of a particular accumulator of a plurality of accumulators included in the matrix processor for operations associated with the particular matrix processor instruction. 
     
     
         35 . The microprocessor system of  claim 34 , wherein the particular accumulator stores intermediate results using a first bit depth format with a higher-bit floating point format than a second bit depth format used for input operands. 
     
     
         36 . The microprocessor system of  claim 31 , wherein a particular matrix processor instruction specifies a first floating-point matrix operand and a second floating-point matrix operand, and wherein the first and second floating-point matrix operands are formatted using a first bit depth format. 
     
     
         37 . The microprocessor system of  claim 31 , wherein a particular matrix processor instruction specifies a configurable exponent bias. 
     
     
         38 . A method implemented by a node engine comprising a plurality of matrix processors, the method comprising:
 receiving, from a control unit of the node engine, a plurality of matrix processor instructions; and   executing, via the node engine, the matrix processor instructions, wherein operations associated with the matrix processor instructions are interleaved for execution by the matrix processors which form the node engine,   and wherein the matrix processors are configured to execute the operations in different bit depth formats based on the associated matrix processor instructions.   
     
     
         39 . The method of  claim 38 , wherein each matrix processor includes a plurality of accumulators, and wherein the matrix processors interleave execution of the subset based on use of accumulators for respective matrix processor instructions. 
     
     
         40 . The method of  claim 39 , wherein for each matrix processor, a particular accumulator of the plurality of accumulators stores intermediate results using a first bit depth format with a higher-bit floating point format than a second bit depth format used for input operands.

Join the waitlist — get patent alerts

Track US2023177108A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.