Decision-making method and apparatus based on deep reinforcement learning through prior data and selective imitation learning
Abstract
A deep reinforcement learning-based decision-making apparatus through prior data and selective imitation learning is disclosed. The deep reinforcement learning-based decision-making apparatus comprises a prior data collection unit configured to collect prior data from one or more other agents; a prior data processing unit configured to process the collected prior data into data including state, action, next state, and reward; and a policy learning unit configured to learn policy of an ego agent using the processed prior data and interaction data including state, action, next state, and reward obtained through real-time interaction with environment.
Claims
exact text as granted — not AI-modified1 . A deep reinforcement learning based decision-making apparatus through prior data and selective imitation learning comprising:
a prior data collection unit configured to collect prior data from one or more other agents; a prior data processing unit configured to process the collected prior data into data including state, action, next state, and reward; and a policy learning unit configured to learn policy of an ego agent using the processed prior data and interaction data including state, action, next state, and reward obtained through real-time interaction with environment.
2 . The decision-making apparatus of claim 1 , wherein an objective function of policy network of the ego agent has a selective imitation learning term with a selective imitation learning weight that determines degree of imitation according to the magnitude of the reward of data sampled from the processed prior data and the interaction data.
3 . The decision-making apparatus of claim 2 , wherein the selective imitation learning term in the objective function of the policy network is added when the reward of the sampled data is greater than a preset threshold.
4 . The decision-making apparatus of claim 1 , wherein the processed prior data and the interaction data are stored in a single replay buffer.
5 . The decision-making apparatus of claim 1 , wherein the processed prior data and the interaction data are stored in different replay buffers.
6 . The decision-making apparatus of claim 4 , wherein the policy learning unit learns policy by sampling the processed prior data and the interaction data from the single replay buffer or the different replay buffers by a preset number of samples.
7 . A deep reinforcement learning-based decision-making apparatus through prior data and selective imitation learning comprising:
a processor; and a memory connected to the processor, wherein the memory stores program instructions, when executed by the processor, configured to perform operations comprising, collecting prior data from one or more other agents, processing the collected prior data into data including state, action, next state, and reward, and learning policy of an ego agent using the processed prior data and interaction data including state, action, next state, and reward obtained through real-time interaction with an environment.
8 . A deep reinforcement learning-based decision-making method through prior data and selective imitation learning comprising:
collecting prior data from one or more other agents; processing the collected prior data into data including state, action, next state, and reward; and learning policy of an ego agent using the processed prior data and interaction data including state, action, next state, and reward obtained through real-time interaction with an environment.
9 . The decision-making method of claim 8 , wherein an objective function of policy network of the ego agent has a selective imitation learning term with a selective imitation learning weight that determines a degree of imitation according to the magnitude of the reward of data sampled from the processed prior data and the interaction data.
10 . The decision-making method of claim 9 , wherein the selective imitation learning term in the objective function of the policy network is added when the reward of the sampled data is greater than a preset threshold.
11 . The decision-making method of claim 8 , wherein the processed prior data and the interaction data are stored in a single replay buffer.
12 . The decision-making method of claim 8 , wherein the processed prior data and the interaction data are stored in different replay buffers.
13 . The decision-making method of claim 11 , wherein the learning of policy of an ego agent comprises learning policy by sampling the processed prior data and the interaction data from the single replay buffer or the different replay buffers by a preset number of samples.Join the waitlist — get patent alerts
Track US2025284970A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.