US2025103900A1PendingUtilityA1

Federated learning with foundation model distillation

Assignee: BOSCH GMBH ROBERTPriority: Sep 22, 2023Filed: Sep 22, 2023Published: Mar 27, 2025
Est. expirySep 22, 2043(~17.2 yrs left)· nominal 20-yr term from priority
G06N 3/096G06N 3/08G06N 3/084G06N 3/098G06N 3/045
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems of training neural networks with federated learning. Machine learning models are sent from a server to clients, yielding local machine learning models. At each client, the models are trained with locally-stored data, including determining a respective cross entropy loss for each of the plurality of local machine learning models. Weights for each local model are updated, and transferred to the server without transferring locally-stored data. The transferred weights are aggregated at the server to obtain an aggregated server-maintained machine learning model. At the server, a distillation loss based on a foundation model is generated. The aggregated server-maintained machine learning is updated to obtain aggregated respective weights, which are transferred to the clients for updating in the local models.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of training neural networks with federated learning, the method comprising:
 sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models;   at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models;   updating respective weights for each of the plurality of local machine learning models;   transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients;   aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model;   generating, at the server, a distillation loss based on a foundation model on the server;   updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights;   transferring the updated aggregated respective weights to each of the plurality of clients; and   updating each of the plurality of local machine learning models with the updated aggregated respective weights.   
     
     
         2 . The method of  claim 1 , wherein the aggregation is according to 
       
         
           
             
               
                 θ 
                 S 
                 r 
               
               = 
               
                 
                   ∑ 
                   
                        
                     
                       k 
                       = 
                       1 
                     
                   
                   
                        
                     N 
                   
                 
                 
                   
                     1 
                     
                       
                         ❘ 
                         "\[LeftBracketingBar]" 
                       
                       
                         D 
                         k 
                       
                       
                         ❘ 
                         "\[RightBracketingBar]" 
                       
                     
                   
                   ⁢ 
                   
                     
                       θ 
                       DS 
                       k 
                     
                     . 
                   
                 
               
             
           
         
       
     
     
         3 . A system of training neural networks with federated learning, the system comprising:
 memory storing instructions; and   a plurality of processors that, when executing the instructions stored in the memory, collectively perform:   sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models;
 at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models; 
 updating respective weights for each of the plurality of local machine learning models; 
 transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients; 
 aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model; 
 generating, at the server, a distillation loss based on a foundation model on the server; 
 updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights; 
 transferring the updated aggregated respective weights to each of the plurality of clients; and 
 updating each of the plurality of local machine learning models with the updated aggregated respective weights. 
   
     
     
         4 . The method of  claim 3 , wherein the aggregation is according to 
       
         
           
             
               
                 θ 
                 S 
                 r 
               
               = 
               
                 
                   ∑ 
                   
                        
                     
                       k 
                       = 
                       1 
                     
                   
                   
                        
                     N 
                   
                 
                 
                   
                     1 
                     
                       
                         ❘ 
                         "\[LeftBracketingBar]" 
                       
                       
                         D 
                         k 
                       
                       
                         ❘ 
                         "\[RightBracketingBar]" 
                       
                     
                   
                   ⁢ 
                   
                     
                       θ 
                       DS 
                       k 
                     
                     . 
                   
                 
               
             
           
         
       
     
     
         5 . A robotic system operated by a neural network comprising:
 memory storing instructions; and   at least one processor that, when executing the instructions stored in the memory, collectively train the neural networks with federated learning by:
 sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models; 
 at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models; 
 updating respective weights for each of the plurality of local machine learning models; 
 transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients; 
 aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model; 
 generating, at the server, a distillation loss based on a foundation model on the server; 
 updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights; 
 transferring the updated aggregated respective weights to each of the plurality of clients; and 
 updating each of the plurality of local machine learning models with the updated aggregated respective weights. 
   
     
     
         6 . The method of  claim 5 , wherein the aggregation is according to 
       
         
           
             
               
                 θ 
                 S 
                 r 
               
               = 
               
                 
                   ∑ 
                   
                        
                     
                       k 
                       = 
                       1 
                     
                   
                   
                        
                     N 
                   
                 
                 
                   
                     1 
                     
                       
                         ❘ 
                         "\[LeftBracketingBar]" 
                       
                       
                         D 
                         k 
                       
                       
                         ❘ 
                         "\[RightBracketingBar]" 
                       
                     
                   
                   ⁢ 
                   
                     
                       θ 
                       DS 
                       k 
                     
                     . 
                   
                 
               
             
           
         
       
     
     
         7 . The robotic system of  claim 5 , wherein the robotic system is an autonomous driving vehicle. 
     
     
         8 . The robotic system of  claim 5 , wherein the robotic system is a medical system.

Join the waitlist — get patent alerts

Track US2025103900A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.