US2022269637A1PendingUtilityA1

General-purpose parallel computing architecture

Assignee: GOLDMAN SACHS & CO LLCPriority: May 21, 2015Filed: May 11, 2022Published: Aug 25, 2022
Est. expiryMay 21, 2035(~8.8 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 3/063G06N 20/00G06F 13/4068G06F 13/40G06F 15/80G06F 15/8023G06F 13/16G06N 7/005
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An apparatus includes multiple parallel computing cores and multiple parallel coprocessor/reducer cores associated with each computing core. Each computing core is configured to perform one or more processing operations, generate input data, and provide the input data to designated coprocessor/reducer cores associated with at least some of the computing cores. Each coprocessor/reducer core associated with a respective computing core is configured to generate output data. Some of the coprocessor/reducer cores associated with the respective computing core are configured to perform part of a distributed operation using the output data to generate intermediate results. A designated one of the coprocessor/reducer cores associated with the respective computing core is configured to provide one or more final results to the computing core. The coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different computing core, and each computing core is communicatively coupled to its designated coprocessor/reducer cores in the columns.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An apparatus comprising:
 multiple parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and   multiple parallel coprocessor/reducer cores associated with each computing core, each computing core configured to provide the input data generated by that computing core to designated coprocessor/reducer cores associated with at least some of the computing cores;   wherein the coprocessor/reducer cores are functional units, each of the coprocessor/reducer cores associated with a respective computing core configured to generate output data, some of the coprocessor/reducer cores associated with the respective computing core configured to perform part of a distributed operation using the output data to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core configured to provide one or more final results to the computing core; and   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.   
     
     
         2 . The apparatus of  claim 1 , further comprising:
 signal lines that communicatively couple the computing cores to the coprocessor/reducer cores in all of the columns.   
     
     
         3 . The apparatus of  claim 1 , wherein:
 the parallel computing cores comprise N computing cores; and   each computing core is associated with N parallel coprocessor/reducer cores.   
     
     
         4 . The apparatus of  claim 1 , wherein:
 the computing cores reside in a first integrated circuit chip; and   the coprocessor/reducer cores reside in a second integrated circuit chip.   
     
     
         5 . The apparatus of  claim 4 , wherein at least one of:
 the computing cores in the first integrated circuit chip are configured to communicate with different numbers or types of coprocessor/reducer cores in different second integrated circuit chips; and   the coprocessor/reducer cores in the second integrated circuit chip are configured to communicate with different numbers or types of computing cores in different first integrated circuit chips.   
     
     
         6 . The apparatus of  claim 1 , wherein each coprocessor/reducer core comprises processing circuitry and a memory. 
     
     
         7 . An apparatus comprising:
 multiple parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and   multiple parallel coprocessor/reducer cores associated with each computing core, each computing core configured to provide the input data generated by that computing core to designated coprocessor/reducer cores associated with at least some of the computing cores, each coprocessor/reducer core configured to generate output data;   wherein subsets of the coprocessor/reducer cores associated with the computing cores are configured to distributively apply one or more operations to the output data, one of the coprocessor/reducer cores in each subset further configured to provide one or more final results to the associated computing core; and   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.   
     
     
         8 . The apparatus of  claim 7 , further comprising:
 signal lines that communicatively couple all of the computing cores to coprocessor/reducer cores in all of the columns.   
     
     
         9 . The apparatus of  claim 7 , wherein:
 the parallel computing cores comprise N computing cores; and   each computing core is associated with N parallel coprocessor/reducer cores.   
     
     
         10 . The apparatus of  claim 7 , wherein:
 the computing cores reside in a first integrated circuit chip; and   the coprocessor/reducer cores reside in a second integrated circuit chip.   
     
     
         11 . An apparatus comprising:
 N parallel computing cores, each computing core configured to perform one or more processing operations and generate input data; and   N×N coprocessor/reducer cores, wherein each computing core is associated with N parallel coprocessor/reducer cores, each computing core configured to provide the input data generated by that computing core to a designated one of the coprocessor/reducer cores associated with at least some of the computing cores, each coprocessor/reducer core configured to generate output data;   wherein some of the coprocessor/reducer cores associated with a respective computing core are configured to perform part of a distributed operation using the output data generated by the coprocessor/reducer cores associated with the respective computing core to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core is configured to provide one or more final results to the computing core;   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns;   wherein the computing cores and the coprocessor/reducer cores are arranged laterally side-by-side in a two-dimensional layout; and   wherein N is an integer having a value of at least sixteen.   
     
     
         12 . An apparatus comprising:
 multiple computing cores, each computing core configured to perform one or more processing operations and generate input data;   multiple coprocessor/reducer cores associated with each computing core, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores and process the input data; and   multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores;   wherein the coprocessor/reducer cores are functional units, each of the coprocessor/reducer cores associated with a respective computing core configured to generate output data, some of the coprocessor/reducer cores associated with the respective computing core configured to perform part of a distributed operation using the output data to generate intermediate results and a designated one of the coprocessor/reducer cores associated with the respective computing core configured to provide one or more final results to the computing core; and   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.   
     
     
         13 . An apparatus comprising:
 multiple computing cores, each computing core configured to perform one or more processing operations and generate input data;   multiple coprocessor/reducer cores associated with each computing core, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores, process the input data, and generate output data; and   multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores;   wherein the coprocessor/reducer cores in a subset of the coprocessor/reducer cores for each computing core are also configured to collectively apply one or more functions to the output data, one of the coprocessor/reducer cores in the subset further configured to provide one or more results to the associated computing core; and   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.   
     
     
         14 . The apparatus of  claim 13 , wherein, for each computing core, the communication links comprise direct connections between that computing core and its associated coprocessor/reducer cores. 
     
     
         15 . The apparatus of  claim 13 , wherein, for each computing core:
 the coprocessor/reducer cores associated with that computing core are linked together in one or more chains; and   the communication links comprise one or more direct connections between that computing core and one or more coprocessor/reducer cores at a beginning of the one or more chains.   
     
     
         16 . The apparatus of  claim 13 , wherein, for each computing core:
 the communication links comprise a direct connection between that computing core and one of the coprocessor/reducer cores associated with that computing core; and   the one of the coprocessor/reducer cores is coupled to multiple additional coprocessor/reducer cores associated with that computing core.   
     
     
         17 . The apparatus of  claim 13 , wherein the coprocessor/reducer cores associated with each of the computing cores are arranged in a tree. 
     
     
         18 . The apparatus of  claim 13 , wherein the communication links comprise links to a shared resource, the shared resource configured to store the input data from the computing cores and to provide the input data to the coprocessor/reducer cores. 
     
     
         19 . The apparatus of  claim 18 , wherein the shared resource comprises a shared memory. 
     
     
         20 . The apparatus of  claim 19 , wherein:
 the shared memory comprises multiple memory locations having multiple memory addresses;   the computing cores are configured to write the input data to different memory addresses; and   the coprocessor/reducer cores are configured to read the input data from the different memory addresses.   
     
     
         21 . An apparatus comprising:
 N parallel computing cores, each computing core configured to perform one or more processing operations and generate input data;   N×N coprocessor/reducer cores, wherein each computing core is associated with N parallel coprocessor/reducer cores, each coprocessor/reducer core configured to receive the input data from at least one of the computing cores, process the input data, and generate output data; and   multiple communication links communicatively coupling the computing cores and the coprocessor/reducer cores associated with the computing cores;   wherein the communication links comprise links to a shared memory, the shared memory configured to store the input data from the computing cores and to provide the input data to the coprocessor/reducer cores;   wherein the shared memory comprises multiple memory locations having multiple memory addresses;   wherein the computing cores are configured to write the input data to different memory addresses;   wherein the coprocessor/reducer cores are configured to read the input data from the different memory addresses; and   wherein the coprocessor/reducer cores are arranged in rows and columns, each column is associated with a different one of the computing cores, and each of the computing cores is communicatively coupled to its designated coprocessor/reducer cores in the columns.

Join the waitlist — get patent alerts

Track US2022269637A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.