Learning device, learning method, and learning program
Abstract
The target output means 91 outputs a plurality of second targets, which are optimization results for a first target using one or more objective functions generated in advance by inverse reinforcement learning based on decision making history data indicating an actual change to a target. The selection acceptance means 92 accepts a selection instruction from a user for a plurality of the output second targets. The data output means 93 outputs the actual change from the first target to the accepted second target as the decision making history data. The learning means 94 learns the objective function using the decision making history data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A learning device comprising:
a memory storing instructions; and one or more processors configured to execute the instructions to: output a plurality of second targets, which are optimization results for a first target using one or more objective functions generated in advance by inverse reinforcement learning based on decision making history data indicating an actual change to a target; accept a selection instruction from a user for a plurality of the output second targets; output the actual change from the first target to the accepted second target as the decision making history data; and learn the objective function using the decision making history data.
2 . The learning device according to claim 1 , wherein the processor is configured to execute the instructions to
select one or more objective functions from a plurality of the objective functions based on likelihood indicating plausibility of the objective function estimated from the data used for learning the objective function, and output the second target by optimization using the selected objective function.
3 . The learning device according to claim 2 , wherein the processor is configured to execute the instructions to
exclude the objective function whose likelihood is lower than a predetermined threshold from being optimized.
4 . The learning device according to claim 2 , wherein the processor is configured to execute the instructions to
select a predetermined top objective function with a high likelihood among the objective functions whose derivative of the parameter is zero.
5 . The learning device according to claim 2 , wherein the processor is configured to execute the instructions to
use the decision making history data output by the data output means to calculate the likelihood and select the objective function based on the calculated likelihood.
6 . The learning device according to claim 1 , wherein the processor is configured to execute the instructions to
the learning means select a solution with a higher likelihood than a predetermined threshold among the output optimization results, and relearn by adding decision making history data including the selected solution.
7 . The learning device according to claim 1 , wherein the processor is configured to execute the instructions to
output a third target indicating a target resulting from further changing of the second target based on a change instruction regarding the second target accepted from a user; and output the actual change from the second target to the third target as decision making history data.
8 . A learning method comprising:
outputting a plurality of second targets, which are optimization results for a first target using one or more objective functions generated in advance by inverse reinforcement learning based on decision making history data indicating an actual change to a target; accepting a selection instruction from a user for a plurality of the output second targets; outputting the actual change from the first target to the accepted second target as the decision making history data; and learning the objective function using the decision making history data.
9 . A learning method according to claim 8 further comprising
selecting one or more objective functions from a plurality of the objective functions based on likelihood indicating plausibility of the objective function estimated from the data used for learning the objective function, and outputting the second target by optimization using the selected objective function.
10 . A non-transitory computer readable information recording medium storing a learning program, when executed by a processor, that performs a method for:
outputting a plurality of second targets, which are optimization results for a first target using one or more objective functions generated in advance by inverse reinforcement learning based on decision making history data indicating an actual change to a target; accepting a selection instruction from a user for a plurality of the output second targets; outputting the actual change from the first target to the accepted second target as the decision making history data; and learning the objective function using the decision making history data.
11 . The non-transitory computer readable information recording medium according to claim 10 , wherein
one or more objective functions are selected from a plurality of the objective functions based on likelihood indicating plausibility of the objective function estimated from the data used for learning the objective function, and the second target is output by optimization using the selected objective function.Join the waitlist — get patent alerts
Track US2023186099A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.