Autonomous Supply Chain by Collaborative Software Agents and Reinforcement Learning
Abstract
A system and method are disclosed to train machine learning models, generate software agents, and evaluate, via reinforcement learning, the actions of the software agents in a simulated ecosystem. Embodiments include a computer comprising a processor and memory and configured to train one or more machine learning models to generate one or more software agents, wherein each software agent comprises an autonomous software program designed to execute a task in a supply chain network. Embodiments generate a first software agent and a second software agent, and a simulated supply chain ecosystem representing a hierarchical structure of supply chain network tasks. Embodiments simulate one or more tasks executed by the software agents in the simulated supply chain ecosystem, review the tasks according to one or more defined objectives, and apply reinforcement incentives to the software agents.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
training, by a computer comprising a processor and memory, one or more machine learning models of a reinforcement learning process; generating, with the one or more machine learning models, a first software agent and a second software agent, wherein each of the one or more software agents comprises an autonomous software program designed to execute one or more tasks in a supply chain network; simulating, by the computer, the one or more tasks executed by the first software agent and the second software agent in a simulated supply chain ecosystem, wherein decisions by the first software agent and the second software agent to execute the one or more tasks in the simulated supply chain ecosystem are made according to the reinforcement learning process and at least partially based on an epsilon-greedy approach; and applying, by the computer, reinforcement incentives to the first software agent and the second software agent, based at least in part on achievement of one or more defined objectives.
2 . The computer-implemented method of claim 1 , wherein the epsilon-greedy approach further comprises execution of a fraction of the one or more tasks randomly to enable exploration of unobserved states and tasks.
3 . The computer-implemented method of claim 1 , wherein the reinforcement incentives further comprise positive or negative reinforcement incentives according to a degree to which the first software agent and the second software agent accomplish the one or more defined objectives.
4 . The computer-implemented method of claim 1 , wherein the reinforcement learning process further comprises a policy-gradient method for an action policy.
5 . The computer-implemented method of claim 1 , wherein the simulated supply chain ecosystem further comprises a hierarchical structure of the one or more tasks and the one or more defined objectives.
6 . The computer-implemented method of claim 1 , further comprising:
configuring, by the computer, each of the first software agent and the second software agent to at least communicate with, collaborate with and execute orders from each other.
7 . The computer-implemented method of claim 1 , wherein the decisions to execute the one or more tasks are based at least in part on a particular state of the simulated supply chain ecosystem.
8 . A system comprising a computer, the computer comprising a processor and memory and configured to:
train one or more machine learning models of a reinforcement learning process; generate, with the one or more machine learning models, a first software agent and a second software agent, wherein each of the one or more software agents comprises an autonomous software program designed to execute one or more tasks in a supply chain network; simulate the one or more tasks executed by the first software agent and the second software agent in a simulated supply chain ecosystem, wherein decisions by the first software agent and the second software agent to execute the one or more tasks in the simulated supply chain ecosystem are made according to the reinforcement learning process and at least partially based on an epsilon-greedy approach; and apply reinforcement incentives to the first software agent and the second software agent, based at least in part on achievement of the one or more defined objectives.
9 . The system of claim 8 , wherein the epsilon-greedy approach further comprises execution of a fraction of the one or more tasks randomly to enable exploration of unobserved states and tasks.
10 . The system of claim 8 , wherein the reinforcement incentives further comprise positive or negative reinforcement incentives according to a degree to which the first software agent and the second software agent accomplish the one or more defined objectives.
11 . The system of claim 8 , wherein the reinforcement learning process further comprises a policy-gradient method for an action policy.
12 . The system of claim 8 , wherein the simulated supply chain ecosystem further comprises a hierarchical structure of the one or more tasks and the one or more defined obj ectives.
13 . The system of claim 8 , wherein the computer is further configured to:
configure each of the first software agent and the second software agent to at least communicate with, collaborate with and execute orders from each other.
14 . The system of claim 8 , wherein the decisions to execute the one or more tasks are based at least in part on a particular state of the simulated supply chain ecosystem.
15 . A non-transitory computer-readable storage medium embodied with software, the software when executed configured to:
train one or more machine learning models of a reinforcement learning process; generate, with the one or more machine learning models, a first software agent and a second software agent, wherein each of the one or more software agents comprises an autonomous software program designed to execute one or more tasks in a supply chain network; simulate the one or more tasks executed by the first software agent and the second software agent in the simulated supply chain ecosystem, wherein decisions by the first software agent and the second software agent to execute the one or more tasks in the simulated supply chain ecosystem are made according to the Markov-based reinforcement learning process and at least partially based on an epsilon-greedy approach; and apply reinforcement incentives to the first software agent and the second software agent, based at least in part on achievement of the one or more defined objectives.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the epsilon-greedy approach further comprises execution of a fraction of the one or more tasks randomly to enable exploration of unobserved states and tasks.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the reinforcement incentives further comprise positive or negative reinforcement incentives according to a degree to which the first software agent and the second software agent accomplish the one or more defined objectives.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein the reinforcement learning process further comprises a policy-gradient method for an action policy.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein the simulated supply chain ecosystem further comprises a hierarchical structure of the one or more tasks and the one or more defined objectives.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the software when executed is further configured to:
configure each of the first software agent and the second software agent to at least communicate with, collaborate with and execute orders from each other.Join the waitlist — get patent alerts
Track US2023281558A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.