Method and apparatus for performing reduction operations on a plurality of associated data element values
Abstract
Embodiments detailed herein relate to reduction operations on a plurality of data element values. In one embodiment, a process comprises decoding circuitry to decode an instruction and execution circuitry to execute the decoded instruction. The instruction specifies a first input register containing a plurality of data element values, a first index register containing a plurality of indices, and an output register, where each index of the plurality of indices maps to one unique data element position of the first input register. The execution includes to identify data element values that are associated with one another based on the indices, perform one or more reduction operations on the associated data element values based on the identification, and store results of the one or more reduction operations in the output register.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor comprising:
an instruction cache to store instructions; cache memory to share data among a plurality of processor cores; and the plurality of processor cores, a processor core of the plurality of processor cores comprising execution circuitry to execute an instruction of the instructions, the execution including:
identifying data element values indicated by a first operand of the instruction, the data element values being associated with one another based on indices indicated by a second operand of the instruction,
performing one or more same reduction operations on the associated data element values based on the identification, and
storing results of the one or more same reduction operations in a storage location indicated by a third operand of the instruction.
2 . The processor of claim 1 , wherein operation code (opcode) of the instruction specifies the one or more same reduction operations.
3 . The processor of claim 1 , wherein a group of data element values being associated with one another when the group of data element values have a same index value, and wherein to perform the one or more same reduction operations is to, for the group of data element values sharing the same index value, combine the group of data element values to generate an arithmetic combination as a result.
4 . The processor of claim 1 , wherein the one or more same reduction operations comprises one or more of:
selection of a maximum value or minimum value of the associated data element values, and computation of a mean or median value of the associated data element values.
5 . The processor of claim 1 , wherein the storage location comprises an output register, each data element position of the output register corresponding to one of corresponding associated data element values.
6 . The processor of claim 1 , wherein the instruction includes a fourth operand indicating a plurality of mask values, wherein each mask value indicates a data element position of the storage location being active or inactive, and wherein the results do not write to the data element position that is inactive.
7 . The processor of claim 1 , wherein the processor core comprises a plurality of computing units that are synchronized in performing the same one or more reduction operations.
8 . The processor of claim 1 , wherein the processor is a graphics processing unit (GPU).
9 . A method comprising:
storing instructions in an instruction cache of a processor; sharing data in a cache memory among a plurality of processor cores within the processor; and executing an instruction of the instructions in a processor core of the plurality of processor cores, the execution including:
identifying data element values indicated by a first operand of the instruction, the data element values being associated with one another based on indices indicated by a second operand of the instruction,
performing one or more same reduction operations on the associated data element values based on the identification, and
storing results of the one or more same reduction operations in a storage location indicated by a third operand of the instruction.
10 . The method of claim 9 , wherein operation code (opcode) of the instruction specifies the one or more same reduction operations.
11 . The method of claim 9 , wherein a group of data element values being associated with one another when the group of data element values have a same index value, and wherein to perform the one or more same reduction operations is to, for the group of data element values sharing the same index value, combine the group of data element values to generate an arithmetic combination as a result.
12 . The method of claim 9 , wherein the one or more same reduction operations comprises one or more of:
selection of a maximum value or minimum value of the associated data element values, and computation of a mean or median value of the associated data element values.
13 . The method of claim 9 , wherein the storage location comprises an output register, each data element position of the output register corresponding to one of corresponding associated data element values.
14 . The method of claim 9 , wherein the instruction includes a fourth operand indicating a plurality of mask values, wherein each mask value indicates a data element position of the storage location being active or inactive, and wherein the results do not write to the data element position that is inactive.
15 . The method of claim 9 , wherein the processor core comprises a plurality of computing units that are synchronized in performing the same one or more reduction operations.
16 . A non-transitory machine-readable medium storing an instruction, which when executed by a processor causes the processor to perform operations, the operations comprising:
storing instructions in an instruction cache of a processor; sharing data in a cache memory among a plurality of processor cores within the processor; and executing an instruction of the instructions in a processor core of the plurality of processor cores, the execution including:
identifying data element values indicated by a first operand of the instruction, the data element values being associated with one another based on indices indicated by a second operand of the instruction,
performing one or more same reduction operations on the associated data element values based on the identification, and
storing results of the one or more same reduction operations in a storage location indicated by a third operand of the instruction.
17 . The non-transitory machine-readable medium of claim 16 , wherein operation code (opcode) of the instruction specifies the one or more same reduction operations.
18 . The non-transitory machine-readable medium of claim 16 , wherein a group of data element values being associated with one another when the group of data element values have a same index value, and wherein to perform the one or more same reduction operations is to, for the group of data element values sharing the same index value, combine the group of data element values to generate an arithmetic combination as a result.
19 . The non-transitory machine-readable medium of claim 16 , wherein the storage location comprises an output register, each data element position of the output register corresponding to one of corresponding associated data element values.
20 . The non-transitory machine-readable medium of claim 16 , wherein the instruction includes a fourth operand indicating a plurality of mask values, wherein each mask value indicates a data element position of the storage location being active or inactive, and wherein the results do not write to the data element position that is inactive.Join the waitlist — get patent alerts
Track US2022229661A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.