US2009144745A1PendingUtilityA1

Performance Evaluation of Algorithmic Tasks and Dynamic Parameterization on Multi-Core Processing Systems

Individually held — no corporate assignee on recordPriority: Nov 29, 2007Filed: Nov 29, 2007Published: Jun 4, 2009
Est. expiryNov 29, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G06F 11/3428G06F 11/3447G06F 11/3404G06F 11/3433
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Apparatus for evaluating the performance of DMA-based algorithmic tasks on a target multi-core processing system includes a memory and at least one processor coupled to the memory. The processor is operative: to input a template for a specified task, the template including DMA-related parameters specifying DMA operations and computational operations to be performed; to evaluate performance for the specified task by running a benchmark on the target multi-core processing system, the benchmark being operative to generate data access patterns using DMA operations and invoking prescribed computation routines as specified by the input template; and to provide results of the benchmark indicative of a measure of performance of the specified task corresponding to the target multi-core processing system.

Claims

exact text as granted — not AI-modified
1 - 4 . (canceled) 
     
     
         5 . Apparatus for evaluating performance of direct memory access (DMA)-based algorithmic tasks on a target multi-core processing system, the apparatus comprising:
 a memory; and   at least one processor coupled to the memory and operative: (i) to input a template for a specified task, the template including DMA-related parameters specifying DMA operations and computational operations to be performed, the template comprising information specifying at least one of: a number of processing cores to use for executing the specified task; a number of iterations of the specified task; a list of work-items corresponding to the specified task; and dependencies between the work-items; (ii) to evaluate performance for the specified task by running a benchmark on the target multi-core processing system, the benchmark being operative to generate data access patterns using DMA operations and invoking prescribed computation routines as specified by the template; (iii) to provide results of the benchmark indicative of a measure of performance of the specified task corresponding to the target multi-core processing system, the step of providing results further comprising at least one of measuring an execution time of the benchmark run on the target multi-core processing system and measuring a computation rate of the benchmark; (iv) for a plurality of templates, to repeat the steps of inputting a template, evaluating performance and providing results, each of the plurality of templates including a unique set of parameter values characterizing the specified task; (v) to iteratively refine the target multi-core processing system as a function of the results of the benchmark; and (vi) to perform algorithmic DMA throttling by utilizing the results of the benchmark in a feedback configuration to so as to reduce a likelihood of power throttling in the target multi-processing system;   wherein each of the work-items corresponding to one of a DMA operation, a DMA wait operation, and a compute operation, the DMA operation work-item comprising at least one parameter specifying at least one of: a type of DMA operation to be performed; an identifier uniquely identifying each of the work-items corresponding to the DMA operation; a starting global address of a remote or shared memory unit utilized by the DMA operation work-item; a starting local address of a local memory unit utilized by the DMA operation work-item; a number of outer-block iterations to be performed; a number of middle-block iterations to be performed; a number of inner-block iterations to be performed; a jump size by which to increment an address for performing outer-loop iterations; a jump size by which to increment the address for performing middle-loop iterations; a jump size by which to increment the address for performing inner-loop iterations; a number of list entries in a DMA list; a DMA size of a given list entry in the DMA list; a size by which to increment the address between list entries in a DMA list; a number of times the DMA operation work-item is to be performed; and when the DMA operation work-item is to be performed first.

Join the waitlist — get patent alerts

Track US2009144745A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.