General-purpose parallel computing architecture
Abstract
An apparatus includes multiple parallel computing cores and multiple parallel coprocessor/reducer cores associated with each computing core. Each computing core is configured to perform one or more processing operations, generate input data, and provide the input data to designated coprocessor/reducer cores associated with at least some of the computing cores. Each coprocessor/reducer core associated with a respective computing core is configured to generate output data. Some of the coprocessor/reducer cores associated with the respective computing core are configured to perform part of a distributed operation using the output data to generate intermediate results. A designated one of the coprocessor/reducer cores associated with the respective computing core is configured to provide one or more final results to the computing core. The coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different computing core, and each computing core is communicatively coupled to its designated coprocessor/reducer cores in the columns.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An apparatus comprising:
multiple parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and multiple parallel coprocessor/reducer cores associated with each computing core, each computing core configured to provide the input data generated by that computing core to designated coprocessor/reducer cores associated with at least some of the computing cores; wherein the coprocessor/reducer cores are functional units, each of the coprocessor/reducer cores associated with a respective computing core configured to generate output data, some of the coprocessor/reducer cores associated with the respective computing core configured to perform part of a distributed operation using the output data to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core configured to provide one or more final results to the computing core; and wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.
2 . The apparatus of claim 1 , further comprising:
signal lines that communicatively couple the computing cores to the coprocessor/reducer cores in all of the columns.
3 . The apparatus of claim 1 , wherein:
the parallel computing cores comprise N computing cores; and each computing core is associated with N parallel coprocessor/reducer cores.
4 . The apparatus of claim 1 , wherein:
the computing cores reside in a first integrated circuit chip; and the coprocessor/reducer cores reside in a second integrated circuit chip.
5 . The apparatus of claim 4 , wherein at least one of:
the computing cores in the first integrated circuit chip are configured to communicate with different numbers or types of coprocessor/reducer cores in different second integrated circuit chips; and the coprocessor/reducer cores in the second integrated circuit chip are configured to communicate with different numbers or types of computing cores in different first integrated circuit chips.
6 . The apparatus of claim 1 , wherein each coprocessor/reducer core comprises processing circuitry and a memory.
7 . An apparatus comprising:
multiple parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and multiple parallel coprocessor/reducer cores associated with each computing core, each computing core configured to provide the input data generated by that computing core to designated coprocessor/reducer cores associated with at least some of the computing cores, each coprocessor/reducer core configured to generate output data; wherein subsets of the coprocessor/reducer cores associated with the computing cores are configured to distributively apply one or more operations to the output data, one of the coprocessor/reducer cores in each subset further configured to provide one or more final results to the associated computing core; and wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.
8 . The apparatus of claim 7 , further comprising:
signal lines that communicatively couple all of the computing cores to coprocessor/reducer cores in all of the columns.
9 . The apparatus of claim 7 , wherein:
the parallel computing cores comprise N computing cores; and each computing core is associated with N parallel coprocessor/reducer cores.
10 . The apparatus of claim 7 , wherein:
the computing cores reside in a first integrated circuit chip; and the coprocessor/reducer cores reside in a second integrated circuit chip.
11 . An apparatus comprising:
N parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and N×N coprocessor/reducer cores, wherein each computing core is associated with N parallel coprocessor/reducer cores, each computing core configured to provide the input data generated by that computing core to a designated one of the coprocessor/reducer cores associated with at least some of the computing cores, each coprocessor/reducer core configured to generate output data; wherein some of the coprocessor/reducer cores associated with a respective computing core are configured to perform part of a distributed operation using the output data generated by the coprocessor/reducer cores associated with the respective computing core to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core is configured to provide one or more final results to the computing core; wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns; wherein the computing cores and the coprocessor/reducer cores are arranged laterally side-by-side in a two-dimensional layout; and wherein N is an integer having a value of at least sixteen.
12 . An apparatus comprising:
multiple computing cores, each computing core configured to perform one or more processing operations and generate input data; multiple coprocessor/reducer cores associated with each computing core, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores and process the input data; and multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores; wherein the coprocessor/reducer cores are functional units, each of the coprocessor/reducer cores associated with a respective computing core configured to generate output data, some of the coprocessor/reducer cores associated with the respective computing core configured to perform part of a distributed operation using the output data to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core configured to provide one or more final results to the computing core; and wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.
13 . An apparatus comprising:
multiple computing cores, each computing core configured to perform one or more processing operations and generate input data; multiple coprocessor/reducer cores associated with each computing core, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores, process the input data, and generate output data; and multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores; wherein the coprocessor/reducer cores in a subset of the coprocessor/reducer cores for each computing core are also configured to collectively apply one or more functions to the output data, one of the coprocessor/reducer cores in the subset further configured to provide one or more results to the associated computing core; and wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.
14 . The apparatus of claim 13 , wherein, for each computing core, the communication links comprise direct connections between that computing core and its associated coprocessor/reducer cores.
15 . The apparatus of claim 13 , wherein, for each computing core:
the coprocessor/reducer cores associated with that computing core are linked together in one or more chains; and the communication links comprise one or more direct connections between that computing core and one or more coprocessor/reducer cores at a beginning of the one or more chains.
16 . The apparatus of claim 13 , wherein, for each computing core:
the communication links comprise a direct connection between that computing core and one of the coprocessor/reducer cores associated with that computing core; and the one of the coprocessor/reducer cores is coupled to multiple additional coprocessor/reducer cores associated with that computing core.
17 . The apparatus of claim 13 , wherein the coprocessor/reducer cores associated with each of the computing cores are arranged in a tree.
18 . The apparatus of claim 13 , wherein the communication links comprise links to a shared resource, the shared resource configured to store the input data from the computing cores and to provide the input data to the coprocessor/reducer cores.
19 . The apparatus of claim 18 , wherein the shared resource comprises a shared memory.
20 . The apparatus of claim 19 , wherein:
the shared memory comprises multiple memory locations having multiple memory addresses; the computing cores are configured to write the input data to different memory addresses; and the coprocessor/reducer cores are configured to read the input data from the different memory addresses.
21 . An apparatus comprising:
N parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; N×N coprocessor/reducer cores, wherein each computing core is associated with N parallel coprocessor/reducer cores, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores, process the input data, and generate output data; and multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores; wherein the communication links comprise links to a shared memory, the shared memory configured to store the input data from the computing cores and to provide the input data to the coprocessor/reducer cores; wherein the shared memory comprises multiple memory locations having multiple memory addresses; wherein the computing cores are configured to write the input data to different memory addresses; wherein the coprocessor/reducer cores are configured to read the input data from the different memory addresses; and wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.Join the waitlist — get patent alerts
Track US2022269637A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.