US2024036949A1PendingUtilityA1

Broadcasting machine learning data

Assignee: ADVANCED RISC MACH LTDPriority: Aug 1, 2022Filed: Jul 31, 2023Published: Feb 1, 2024
Est. expiryAug 1, 2042(~16 yrs left)· nominal 20-yr term from priority
G06T 2200/28G06F 9/542G06F 9/4843G06F 12/0842G06F 2209/543G06F 2212/62G06F 2212/60G06F 9/5044G06T 1/20G06F 9/4881G06F 2209/509G06F 9/5038G06F 9/544G06T 15/005G06T 1/60G06F 9/505G06F 9/4806G06F 9/5077G06F 9/5066G06F 9/30047G06F 9/3836G06F 9/3867G06N 3/063
77
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

There is provided a processor configured to transfer data to a plurality of processor circuits. The apparatus includes broadcast circuitry that broadcasts first machine learning data to at least a subset of the plurality of processor circuits.

Claims

exact text as granted — not AI-modified
1 . A processor comprising:
 a plurality of processor circuits to perform a machine learning process; and   broadcast circuitry configured to broadcast data to at least a subset of the plurality of processor circuits; wherein   the processor is configured to:
 obtain at least a first subset of machine learning data from memory to storage circuitry; and 
 broadcast the at least a first subset of machine learning data from storage circuitry to at least a subset of the plurality of processor circuits. 
   
     
     
         2 . The apparatus according to  claim 1 , comprising:
 fetch circuitry, the fetch circuitry configured to fetch or stream machine learning data from memory to storage circuitry.   
     
     
         3 . The apparatus according  claim 1 , wherein
 the broadcast circuitry is configured to broadcast first machine learning data from storage circuitry to local-storage circuitry in the at least a subset of the plurality of processor circuits.   
     
     
         4 . The apparatus according to  claim 2 , comprising:
 the fetch circuitry configured to fetch or stream first compressed machine learning data from memory; and   decompression circuitry configured to decompress first compressed machine learning data to generate first decompressed machine learning data; and   the fetch circuitry configured to fetch first decompressed machine learning data from decompression circuitry to storage circuitry.   
     
     
         5 . The apparatus according to  claim 1 , wherein the storage circuitry is a cache. 
     
     
         6 . The apparatus according to  claim 1 , wherein
 the storage circuitry is configured to store, in association with each entry, an indication of whether that entry is to be broadcast by the broadcast circuitry.   
     
     
         7 . The apparatus according to  claim 1 , wherein
 the broadcast circuitry is configured to broadcast the first machine learning data to at most a subset of the plurality of processor circuits; and   the broadcast circuitry is configured to broadcast third machine learning data, different to the first machine learning data, to a further subset of the plurality of processor circuits;   the subset of the plurality of processor circuits and the further subset of the plurality of processor circuits are mutually exclusive; and   the first machine learning data and the third machine learning data relate to different layers of a neural network.   
     
     
         8 . The apparatus according to  claim 1 , wherein the processor is a tile-based graphics processor; and
 the processor circuits are shader cores; and   the storage circuitry is a cache; and   the local-storage circuitry is a tile buffer in a shader core, wherein   the broadcast circuitry is configured to broadcast the first machine learning data from the cache to tile buffers in at least a subset of the plurality of shader cores.   
     
     
         9 . The apparatus according to  claim 1 , further
 the processor is configured to:
 obtain at least a second subset of machine learning data from memory to storage circuitry; and 
 transfer the at least a second subset of machine learning data from storage circuitry to at least a subset of the plurality of processor circuits; wherein 
 the at least subset of second machine learning data is different for each of the at least a subset of the processor circuits. 
   
     
     
         10 . The apparatus according to  claim 9 , comprising:
 dispatch circuitry to cause processor circuits to process machine learning data, wherein   the dispatch circuitry configured to cause each of the at least a subset of the processor circuits to process its first machine learning data with the second machine learning data.   
     
     
         11 . The apparatus according to  claim 9 , wherein either the first machine learning data is a kernel and the second machine learning data is a feature map, or the first machine learning data is the feature map and the second machine learning data is the kernel. 
     
     
         12 . The apparatus according to  claim 9 , wherein
 the apparatus is configured to operate in a kernel broadcast mode in which the first machine learning data is a kernel and the second machine learning data is a feature map; and   the apparatus is configured to operate in a map broadcast mode in which the first machine learning data is the feature map and the second machine learning data is the kernel.   
     
     
         13 . The apparatus according to  claim 12 , wherein
 the apparatus is configured to dynamically change between the map broadcast mode and the kernel broadcast mode.   
     
     
         14 . The apparatus according to  claim 12 , wherein
 the apparatus is configured to dynamically change between the map broadcast mode and the kernel broadcast mode in dependence on a layer of neural network to which the kernel and the feature map relate.   
     
     
         15 . The apparatus according to  claim 1 , the storage circuitry comprising:
 snoop filter circuitry, the snoop filter circuitry configured to store, in association with each entry, for coherent traffic coherency state, and for non-coherent traffic to store, an indication of whether that entry is to be broadcast by the broadcast circuitry.   
     
     
         16 . The apparatus according to  claim 15 , further comprising:
 snoop filter circuitry configured to a store snoop filter entry, the snoop filter entry to store at least one of a broadcast flag, a broadcast destination, or a broadcast address, for a non-coherent entry.   
     
     
         17 . The apparatus according to  claim 1 , comprising:
 a host processor configured to execute a driver; and   job manager circuitry configured to dispatch tasks to at least a subset of the plurality of processor circuits, wherein   the driver is configured to analyses layer processing of a neural network and to generate a job list to schedule processing of a neural network to at least a subset of the processing circuits; and   the job manager circuitry is configured to process the job list generate by the driver, wherein   the job manager circuitry is configured to determine available plurality of processing circuits and using the job list dispatch tasks to at least a subset of the plurality of processor circuits.   
     
     
         18 . The apparatus according to  claim 17 , wherein
 the driver is configured to select between map broadcast mode and kernel broadcast mode to minimise memory accesses to the at least subset of the plurality of processing circuits in dependence on the size of the kernel and feature map associated with the layer; and   the available storage circuitry is associated with each of the processing circuits; and   the number of at least a subset of the plurality of processing circuits processing the layer.   
     
     
         19 . A data processing method of transferring data to a plurality of processing circuits, the data processing method comprising:
 fetching at least a first subset of machine learning data from memory to storage; and   broadcasting the at least a first subset of machine learning data from storage to at least a subset of the plurality of processor circuits.   
     
     
         20 . A non-transitory computer-readable medium to store computer-readable code for fabrication of an apparatus comprising:
 a plurality of processor circuits to perform a machine learning process; and   broadcast circuitry configured to broadcast data to at least a subset of the plurality of processor circuits; wherein   the processor is configured to:   obtain at least a first subset of machine learning data from memory to storage circuitry; and   broadcast the at least a first subset of machine learning data from storage circuitry to at least a subset of the plurality of processor circuits.

Join the waitlist — get patent alerts

Track US2024036949A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.