Federated learning with foundation model distillation
Abstract
Methods and systems of training neural networks with federated learning. Machine learning models are sent from a server to clients, yielding local machine learning models. At each client, the models are trained with locally-stored data, including determining a respective cross entropy loss for each of the plurality of local machine learning models. Weights for each local model are updated, and transferred to the server without transferring locally-stored data. The transferred weights are aggregated at the server to obtain an aggregated server-maintained machine learning model. At the server, a distillation loss based on a foundation model is generated. The aggregated server-maintained machine learning is updated to obtain aggregated respective weights, which are transferred to the clients for updating in the local models.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of training neural networks with federated learning, the method comprising:
sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models; at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models; updating respective weights for each of the plurality of local machine learning models; transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients; aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model; generating, at the server, a distillation loss based on a foundation model on the server; updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights; transferring the updated aggregated respective weights to each of the plurality of clients; and updating each of the plurality of local machine learning models with the updated aggregated respective weights.
2 . The method of claim 1 , wherein the aggregation is according to
θ
S
r
=
∑
k
=
1
N
1
❘
"\[LeftBracketingBar]"
D
k
❘
"\[RightBracketingBar]"
θ
DS
k
.
3 . A system of training neural networks with federated learning, the system comprising:
memory storing instructions; and a plurality of processors that, when executing the instructions stored in the memory, collectively perform: sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models;
at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models;
updating respective weights for each of the plurality of local machine learning models;
transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients;
aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model;
generating, at the server, a distillation loss based on a foundation model on the server;
updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights;
transferring the updated aggregated respective weights to each of the plurality of clients; and
updating each of the plurality of local machine learning models with the updated aggregated respective weights.
4 . The method of claim 3 , wherein the aggregation is according to
θ
S
r
=
∑
k
=
1
N
1
❘
"\[LeftBracketingBar]"
D
k
❘
"\[RightBracketingBar]"
θ
DS
k
.
5 . A robotic system operated by a neural network comprising:
memory storing instructions; and at least one processor that, when executing the instructions stored in the memory, collectively train the neural networks with federated learning by:
sending at least a portion of a server-maintained machine learning model from a server to a plurality of clients, yielding a plurality of local machine learning models;
at each of the plurality of clients, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein training at each client includes determining a respective cross entropy loss for each of the plurality of local machine learning models;
updating respective weights for each of the plurality of local machine learning models;
transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients;
aggregating, at the server, respective weights from each of the plurality of local machine learning models to obtain an aggregated server-maintained machine learning model;
generating, at the server, a distillation loss based on a foundation model on the server;
updating, at the server, the aggregated server-maintained machine learning model to obtain updated aggregated respective weights;
transferring the updated aggregated respective weights to each of the plurality of clients; and
updating each of the plurality of local machine learning models with the updated aggregated respective weights.
6 . The method of claim 5 , wherein the aggregation is according to
θ
S
r
=
∑
k
=
1
N
1
❘
"\[LeftBracketingBar]"
D
k
❘
"\[RightBracketingBar]"
θ
DS
k
.
7 . The robotic system of claim 5 , wherein the robotic system is an autonomous driving vehicle.
8 . The robotic system of claim 5 , wherein the robotic system is a medical system.Join the waitlist — get patent alerts
Track US2025103900A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.