Training policy using distributional reinforcement learning and conditional value at risk
Abstract
A computer-implemented method for modifying a current policy using reinforcement learning (RL) includes the following operations. A number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment are sampled. For each of the sampled MDPs, behavior data for the current policy is collected, a quantile function of return with the current policy is determined using the collected behavior data, and a current weight is generated by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP. The policy is modified based upon the weights for each of the sampled MDPs. The current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for modifying a current policy using reinforcement learning (RL), comprising:
sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment; for each of the sampled MDPs,
collecting behavior data for the current policy,
determining a quantile function of return with the current policy using the collected behavior data, and
generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and
modifying the policy based upon the current weight for each of the sampled MDPs.
2 . The method of claim 1 , wherein
the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.
3 . The method of claim 2 , wherein
a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data.
4 . The method of claim 3 , wherein
the VAE is optimized by maximizing an evidence lower bound (ELBO).
5 . The method of claim 2 , wherein
the sampled MDPs are sampled based upon a prior distribution of the environment.
6 . The method of claim 2 , wherein
the weight is updated by solving an optimization problem.
7 . The method of claim 2 , wherein
the quantile function of return is determined using distributional RL.
8 . The method of claim 2 , wherein
the behavior data is determined for the current policy by interacting with the environment using the current policy.
9 . A computer hardware system for modifying a current policy using reinforcement learning (RL), comprising:
a hardware processor configured to perform the following executable operations:
sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment;
for each of the sampled MDPs,
collecting behavior data for the current policy,
determining a quantile function of return with the current policy using the collected behavior data, and
generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and
modifying the policy based upon the current weight for each of the sampled MDPs.
10 . The system of claim 9 , wherein
the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.
11 . The system of claim 10 , wherein
a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data.
12 . The system of claim 11 , wherein
the VAE is optimized by maximizing an evidence lower bound (ELBO).
13 . The system of claim 10 , wherein
the sampled MDPs are sampled based upon a prior distribution of the environment.
14 . The system of claim 10 , wherein
the weight is updated by solving an optimization problem.
15 . The system of claim 10 , wherein
the quantile function of return is determined using distributional RL.
16 . The system of claim 10 , wherein
the behavior data is determined for the current policy by interacting with the environment using the current policy.
17 . A computer program product, comprising:
a computer readable storage medium having stored therein program code for modifying a current policy using reinforcement learning (RL), the program code, which when executed by a computer hardware system, cause the computer hardware system to perform:
sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment;
for each of the sampled MDPs,
collecting behavior data for the current policy,
determining a quantile function of return with the current policy using the collected behavior data, and
generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and
modifying the policy based upon the current weight for each of the sampled MDPs, wherein
the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.
18 . The computer program product of claim 17 , wherein
a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data, and the VAE is optimized by maximizing an evidence lower bound (ELBO).
19 . The computer program product of claim 17 , wherein
the weight is updated by solving an inner minimization problem.
20 . The computer program product of claim 17 , wherein
the quantile function of return is determined using distributional RL.Join the waitlist — get patent alerts
Track US2024330696A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.