Performance Evaluation of Algorithmic Tasks and Dynamic Parameterization on Multi-Core Processing Systems
Abstract
A method for evaluating performance of DMA-based algorithmic tasks on a target multi-core processing system includes the steps of: inputting a template for a specified task, the template including DMA-related parameters specifying DMA operations and computational operations to be performed; evaluating performance for the specified task by running a benchmark on the target multi-core processing system, the benchmark being operative to generate data access patterns using DMA operations and invoking prescribed computation routines as specified by the input template; and providing results of the benchmark indicative of a measure of performance of the specified task corresponding to the target multi-core processing system.
Claims
exact text as granted — not AI-modified1 - 21 . (canceled)
22 . A method for evaluating performance of direct memory access (DMA)-based algorithmic tasks on a target multi-core processing system, the method comprising the steps of:
inputting a template for a specified task, the template including DMA-related parameters specifying DMA operations and computational operations to be performed, the template comprising information specifying at least one of: a number of processing cores to use for executing the specified task; a number of iterations of the specified task; a list of work-items corresponding to the specified task; and dependencies between the work-items; evaluating performance for the specified task by running a benchmark on the target multi-core processing system, the benchmark being operative to generate data access patterns using DMA operations and invoking prescribed computation routines as specified by the template; providing results of the benchmark indicative of a measure of performance of the specified task corresponding to the target multi-core processing system, the step of providing results further comprising at least one of measuring an execution time of the benchmark run on the target multi-core processing system and measuring a computation rate of the benchmark; for a plurality of templates, repeating the steps of inputting a template, evaluating performance and providing results, each of the plurality of templates including a unique set of parameter values characterizing the specified task; iteratively refining the target multi-core processing system as a function of the results of the benchmark; and performing algorithmic DMA throttling by utilizing the results of the benchmark in a feedback configuration to so as to reduce a likelihood of power throttling in the target multi-processing system; wherein each of the work-items corresponding to one of a DMA operation, a DMA wait operation, and a compute operation, the DMA operation work-item comprising at least one parameter specifying at least one of: a type of DMA operation to be performed; an identifier uniquely identifying each of the work-items corresponding to the DMA operation; a starting global address of a remote or shared memory unit utilized by the DMA operation work-item; a starting local address of a local memory unit utilized by the DMA operation work-item; a number of outer-block iterations to be performed; a number of middle-block iterations to be performed; a number of inner-block iterations to be performed; a jump size by which to increment an address for performing outer-loop iterations; a jump size by which to increment the address for performing middle-loop iterations; a jump size by which to increment the address for performing inner-loop iterations; a number of list entries in a DMA list; a DMA size of a given list entry in the DMA list; a size by which to increment the address between list entries in a DMA list; a number of times the DMA operation work-item is to be performed; and when the DMA operation work-item is to be performed first.Join the waitlist — get patent alerts
Track US2009144744A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.