US2024036932A1PendingUtilityA1

Graphics processors

Assignee: ADVANCED RISC MACH LTDPriority: Aug 1, 2022Filed: Jul 26, 2023Published: Feb 1, 2024
Est. expiryAug 1, 2042(~16 yrs left)· nominal 20-yr term from priority
G06T 2200/28G06F 9/505G06T 15/005G06F 9/5044G06T 1/20G06F 9/4881G06F 2209/509G06F 9/5038G06F 9/544G06T 1/60G06F 9/4843G06F 9/542G06F 12/0842G06F 2209/543G06F 2212/60G06F 2212/62G06F 9/4806G06F 9/5077G06F 9/5066G06F 9/30047G06F 9/3836G06F 9/3867G06N 3/063
77
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed herein is a graphics processor that comprises a programmable execution unit operable to execute programs to perform graphics processing operations. The graphics processor further comprises a dedicated machine learning processing circuit operable to perform processing operations for machine learning processing tasks. The machine learning processing circuit is in communication with the programmable execution unit internally to the graphics processor. In this way, the graphics processor can be configured such that machine learning processing tasks can be performed by the programmable execution unit, the machine learning processing circuit, or a combination of both, with the different units being able to message each other accordingly to control the processing.

Claims

exact text as granted — not AI-modified
1 . A graphics processor comprising:
 a programmable execution unit operable to execute programs to perform graphics processing operations; and   a machine learning processing circuit operable to perform processing operations for machine learning processing tasks and in communication with the programmable execution unit internally to the graphics processor,   the graphics processor configured such that machine learning processing tasks can be performed by the programmable execution unit, the machine learning processing circuit, or a combination of both.   
     
     
         2 . The graphics processor of  claim 1 , being configured such that, when the execution unit is executing a program including an instruction that relates to a set of machine learning operations to be performed by the machine learning processing circuit: in response to the execution unit executing the instruction, the programmable execution unit is caused to message the machine learning processing circuit to cause the machine learning processing circuit to perform the set of machine learning processing operations. 
     
     
         3 . The graphics processor of  claim 2 , wherein the machine learning processing circuit is configured to return a result of its processing to the execution unit for further processing. 
     
     
         4 . The graphics processor of  claim 1 , wherein the machine learning processing circuit, when performing a machine learning processing task, is operable to cause the execution unit to perform one or more processing operations for the machine learning processing task being performed by the machine learning processing circuit. 
     
     
         5 . The graphics processor of  claim 4 , wherein the machine learning processing circuit is operable to trigger the generation of threads for execution by the programmable execution unit to cause the execution unit to perform the one or more processing operations for the machine learning processing task being performed by the machine learning processing circuit. 
     
     
         6 . The graphics processor of  claim 1 , wherein the machine learning processing circuit comprises one or more multiply-and-accumulate circuits. 
     
     
         7 . The graphics processor of  claim 1 , wherein the graphics processor includes a cache system for transferring data to and from an external memory, and wherein the machine learning processing circuit has access to the graphics processor's cache system. 
     
     
         8 . The graphics processor of  claim 7 , wherein when a machine learning processing tasks is to be performed using the graphics processor, the graphics processor is operable to fetch required input data for the machine learning processing task via the cache system, and write an output of the machine learning processing task to memory via the cache system. 
     
     
         9 . The graphics processor of  claim 7 , further comprising compression and decompression circuits for compressing and decompressing data as it is transferred between the graphics processor and the external memory. 
     
     
         10 . The graphics processor of  claim 1 , comprising a plurality of programmable execution units, arranged as respective shader cores, with each shader core having its own respective machine learning processing circuit, and wherein an overall job controller of the graphics processor is operable to distribute processing tasks between the different shader cores. 
     
     
         11 . The graphics processor of  claim 1 , wherein the graphics processor is configured to perform tile-based rendering, in which graphics data is stored in one or more tile buffers, and wherein when performing a machine learning processing task at least some data for the machine learning processing task is stored using the tile buffers. 
     
     
         12 . A method of operating a graphics processor, the graphics processor comprising:
 a programmable execution unit operable to execute programs to perform graphics processing operations; and   a machine learning processing circuit operable to perform machine learning processing operations and in communication with the programmable execution unit internally to the graphics processor,   the graphics processor configured such that machine learning processing tasks can be performed by the programmable execution unit, the machine learning processing circuit, or a combination of both;   the method comprising:   the graphics processor performing a machine learning task using a combination of the programmable execution unit and the machine learning processing circuit.   
     
     
         13 . The method of  claim 12 , further comprising: when the execution unit is executing a program including an instruction that relates to a set of machine learning operations to be performed by the machine learning processing circuit: in response to the execution unit executing the instruction, the programmable execution unit messaging the machine learning processing circuit to cause the machine learning processing circuit to perform the set of machine learning processing operations. 
     
     
         14 . The method of  claim 12 , further comprising: the machine learning processing circuit, when performing a machine learning processing task, causing the execution unit to perform one or more processing operations for the machine learning processing task being performed by the machine learning processing circuit. 
     
     
         15 . The method of  claim 14 , wherein the machine learning processing circuit causes the execution unit to perform one or more processing operations by triggering the generation of an execution thread, which execution thread when executed by the execution unit causes the execution unit to perform the one or more processing operations for the machine learning processing task being performed by the machine learning processing circuit. 
     
     
         16 . The method of  claim 12 , comprising the machine learning processing circuit returning a result of its processing to the execution unit for further processing. 
     
     
         17 . The method of  claim 12 , wherein the graphics processor includes a cache system for transferring data to and from an external memory, and wherein the machine learning processing circuit has access to the graphics processor's cache system, and wherein when a machine learning processing tasks is to be performed using the graphics processor, the graphics processor fetches required input data for the machine learning processing task via the cache system, and writes an output of the machine learning processing task to memory via the cache system. 
     
     
         18 . The method of  claim 17 , further comprising compressing data as it is written to memory and/or decompressing data as it is retrieved from memory. 
     
     
         19 . The method of  claim 12 , comprising a plurality of programmable execution units, arranged as respective shader cores, with each shader core having its own respective machine learning processing circuit, and wherein an overall job controller of the graphics processor is operable to distribute processing tasks between the different shader cores. 
     
     
         20 . The method of  claim 12 , wherein the graphics processor is configured to perform tile-based rendering, in which graphics data is stored in one or more tile buffers, and wherein the method comprises: when performing a machine learning processing task, storing at least some data for the machine learning processing task using the tile buffers.

Join the waitlist — get patent alerts

Track US2024036932A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.