US2013185496A1PendingUtilityA1

Vector Processing System

Assignee: HESSEL CONNIEPriority: Feb 10, 2005Filed: Dec 17, 2012Published: Jul 18, 2013
Est. expiryFeb 10, 2025(expired)· nominal 20-yr term from priority
G06F 9/30038G06F 9/3888G06F 9/30036G06F 9/3887G11C 7/1072G06F 9/3834G06F 15/8061G06F 9/3836G06F 9/3004G06F 9/3867G06F 9/3885G06F 9/30087
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A vector processing system provides high performance vector processing using a System-On-a-Chip (SOC) implementation technique. One or more scalar processors (or cores) operate in conjunction with a vector processor, and the processors collectively share access to a plurality of memory interfaces coupled to Dynamic Random Access read/write Memories (DRAMs). In typical embodiments the vector processor operates as a slave to the scalar processors, executing computationally intensive Single Instruction Multiple Data (SIMD) codes in response to commands received from the scalar processors. The vector processor implements a vector processing Instruction Set Architecture (ISA) including machine state, instruction set, exception model, and memory model.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A system comprising:
 a plurality of floating point execution units compatible with operation according to a plurality of execution threads;   a plurality of memory channels each coupled to at least one memory element;   a memory buffer switch unit that couples the floating point execution units to the memory channels;   an instruction control that controls the floating point execution units according to a stream of vector instructions executed in accordance with the execution threads;   a processor interface receiving the stream of vector instructions from a processor; and   wherein the memory buffer switch unit consolidates at least two memory requests from the floating point execution units into a single memory access operation directed to one of the memory channels and the at least two memory requests are processed according to a coherency domain implemented by the processor.   
     
     
         3 . The system of  claim 2 , wherein respective parts of multiple ones of the vector instructions are executed independently by respective ones of the floating point execution units. 
     
     
         4 . The system of  claim 2 , wherein at least some of the floating point execution units operate concurrently on parts of a same one of the vector instructions. 
     
     
         5 . The system of  claim 2 , wherein at least one of the memory elements comprises one or more Dynamic Random Accessible read/write Memory (DRAM) memories. 
     
     
         6 . The system of  claim 5 , further comprising the DRAM memories. 
     
     
         7 . The system of  claim 5 , wherein the DRAM memories comprise at least one single DIMM of ×4 DRAMs, and a Reed-Solomon ECC provides for ChipKill operation with the at least one single DIMM of ×4 DRAMs. 
     
     
         8 . The system of  claim 2 , wherein the processor comprises an x 86  processor. 
     
     
         9 . The system of  claim 8 , further comprising the x86 processor. 
     
     
         10 . The system of  claim 9 , wherein the floating point execution units and the x86 processor are implemented in a single integrated circuit die. 
     
     
         11 . The system of  claim 2 , further comprising a page-walking block that fills an Instruction TLB without assistance of the processor. 
     
     
         12 . A method comprising:
 operating a plurality of execution threads on a plurality of floating point execution units;   accessing respective memory elements via a plurality of memory channels;   coupling the floating point execution units to the memory channels via a memory buffer switch unit;   controlling the floating point execution units according to a stream of vector instructions executing, in accordance with the execution threads;   receiving the stream of vector instructions from a processor via a processor interface; and   consolidating, at least two memory requests from the floating point execution units into a single ‘memory access operation directed to one of the memory channels, and processing the at least two memory requests according to a coherency domain implemented by the processor.   
     
     
         13 . The method, of  claim 12 , further comprising independently executing respective parts of multiple ones of the vector instructions by respective ones of the floating point execution units. 
     
     
         14 . The method of  claim 12 , further comprising operating at least some of the floating point execution units concurrently on parts of a same one of the vector instructions. 
     
     
         15 . The method of  claim 12 , wherein at least one of the memory elements comprises one or more Dynamic Random Accessible read/write Memory (DRAM) memories. 
     
     
         16 . The method of  claim 15 , wherein the DRAM memories comprise at least one single DIMM of ×4 DRAMs, and further comprising providing ChipKill operation with the at least one single DTMM of ×4 DRAMs. 
     
     
         17 . The method of  claim 12 , wherein the floating point execution units and the processor are implemented in a single integrated circuit die. 
     
     
         18 . The method of  claim 12 , further comprising filling an Instruction TLB without assistance of the processor. 
     
     
         19 . A system comprising:
 a plurality of floating point means for performing floating point operations according to a plurality of execution threads;   a plurality of memory channels each enabled to access at least one memory element;   memory buffer switch means for coupling the plurality of floating point means to the memory channels;   instruction control means for controlling the plurality of floating point means according to a stream of vector instructions executed in accordance with the execution threads;   processor interface means for receiving the stream of vector instructions from a processor; and   wherein the memory buffer switch means operates at least in part via consolidating at least two memory requests from the plurality of floating point means into a single memory access operation directed to one of the memory channels and the at least two memory requests are processed according to a coherency domain implemented by the processor.   
     
     
         20 . The system of  claim 19 , wherein respective parts of multiple ones of the vector instructions are executed independently by respective ones of the plurality of floating point means. 
     
     
         21 . The system of  claim 19 , wherein at least some of the plurality of floating point means operate concurrently on parts of a same one of the vector instructions.

Join the waitlist — get patent alerts

Track US2013185496A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.