US2024330696A1PendingUtilityA1

Training policy using distributional reinforcement learning and conditional value at risk

Assignee: IBMPriority: Mar 30, 2023Filed: Mar 30, 2023Published: Oct 3, 2024
Est. expiryMar 30, 2043(~16.6 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/088G06N 3/047G06N 3/045G06N 3/006G06N 7/01G06N 3/0455G06N 3/092
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer-implemented method for modifying a current policy using reinforcement learning (RL) includes the following operations. A number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment are sampled. For each of the sampled MDPs, behavior data for the current policy is collected, a quantile function of return with the current policy is determined using the collected behavior data, and a current weight is generated by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP. The policy is modified based upon the weights for each of the sampled MDPs. The current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for modifying a current policy using reinforcement learning (RL), comprising:
 sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment;   for each of the sampled MDPs,
 collecting behavior data for the current policy, 
 determining a quantile function of return with the current policy using the collected behavior data, and 
 generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and 
   modifying the policy based upon the current weight for each of the sampled MDPs.   
     
     
         2 . The method of  claim 1 , wherein
 the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and   the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.   
     
     
         3 . The method of  claim 2 , wherein
 a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data.   
     
     
         4 . The method of  claim 3 , wherein
 the VAE is optimized by maximizing an evidence lower bound (ELBO).   
     
     
         5 . The method of  claim 2 , wherein
 the sampled MDPs are sampled based upon a prior distribution of the environment.   
     
     
         6 . The method of  claim 2 , wherein
 the weight is updated by solving an optimization problem.   
     
     
         7 . The method of  claim 2 , wherein
 the quantile function of return is determined using distributional RL.   
     
     
         8 . The method of  claim 2 , wherein
 the behavior data is determined for the current policy by interacting with the environment using the current policy.   
     
     
         9 . A computer hardware system for modifying a current policy using reinforcement learning (RL), comprising:
 a hardware processor configured to perform the following executable operations:
 sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment; 
 for each of the sampled MDPs,
 collecting behavior data for the current policy, 
 determining a quantile function of return with the current policy using the collected behavior data, and 
 generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and 
 
 modifying the policy based upon the current weight for each of the sampled MDPs. 
   
     
     
         10 . The system of  claim 9 , wherein
 the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and   the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.   
     
     
         11 . The system of  claim 10 , wherein
 a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data.   
     
     
         12 . The system of  claim 11 , wherein
 the VAE is optimized by maximizing an evidence lower bound (ELBO).   
     
     
         13 . The system of  claim 10 , wherein
 the sampled MDPs are sampled based upon a prior distribution of the environment.   
     
     
         14 . The system of  claim 10 , wherein
 the weight is updated by solving an optimization problem.   
     
     
         15 . The system of  claim 10 , wherein
 the quantile function of return is determined using distributional RL.   
     
     
         16 . The system of  claim 10 , wherein
 the behavior data is determined for the current policy by interacting with the environment using the current policy.   
     
     
         17 . A computer program product, comprising:
 a computer readable storage medium having stored therein program code for modifying a current policy using reinforcement learning (RL),   the program code, which when executed by a computer hardware system, cause the computer hardware system to perform:
 sampling a number, corresponding to an inputted sample size, of Markov Decision Processes (MDPs) defining an environment; 
 for each of the sampled MDPs,
 collecting behavior data for the current policy, 
 determining a quantile function of return with the current policy using the collected behavior data, and 
 generating a current weight by updating a weight for a particular sampled MDP using the quantile function of return for the particular sampled MDP; and 
 
 modifying the policy based upon the current weight for each of the sampled MDPs, wherein 
   the current weights are generated by minimizing a conditional value of at risk (CVaR) of a return of the current policy, and   the policy is modified to maximize a weighted average of the CVaR of the return with the current weights.   
     
     
         18 . The computer program product of  claim 17 , wherein
 a variational auto-encoder (VAE) including an encoder and a decoder is optimized using the collected behavior data, and   the VAE is optimized by maximizing an evidence lower bound (ELBO).   
     
     
         19 . The computer program product of  claim 17 , wherein
 the weight is updated by solving an inner minimization problem.   
     
     
         20 . The computer program product of  claim 17 , wherein
 the quantile function of return is determined using distributional RL.

Join the waitlist — get patent alerts

Track US2024330696A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.