Partially observed markov decision process model and its use
Abstract
A method for selecting an action, includes reading, into a memory, a Partially Observed Markov Decision Process (POMDP) model, the POMDP model having top-k action IDs for each belief state, the top-k action IDs maximizing expected long-term cumulative rewards in each time-step, and k being an integer of two or more, in the execution-time process of the POMDP model, detecting a situation where an action identified by the best action ID among the top-k action IDs for a current belief state is unable to be selected due to a constraint, and selecting and executing an action identified by the second best action ID among the top-k action IDs for the current belief state in response to a detection of the situation. The top-k action IDs may be top-k alpha vectors, each of the top-k alpha vectors having an associated action, or identifiers of top-k actions associated with alpha vectors.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for selecting an action, the method comprising:
reading, into a memory, a Partially Observed Markov Decision Process (POMDP) model, the POMDP model having top-k action IDs for each belief state, the top-k action IDs maximizing expected long-term cumulative rewards in each time-step, and k being an integer of two or more; in the execution-time process of the POMDP model, detecting a situation where an action identified by a first best action ID among the top-k action IDs for a current belief state is unable to be selected due to constraint; and selecting and executing an action identified by a second best action ID among the top-k action IDs for the current belief state in response to a detection of the situation.
2 . The method according to claim 1 , wherein the top-k action IDs are top-k alpha vectors and each of the top-k alpha vectors have an associated action.
3 . The method according to claim 1 , wherein the top-k action IDs are identifiers of top-k actions associated with alpha vectors.
4 . The method according to claim 2 , wherein alpha vectors other than the top-k alpha vectors are pruned when the top-k alpha vectors are selected.
5 . The method according to claim 3 , wherein alpha vectors other than the alpha vectors associated with the top-k actions are pruned when the top-k actions are selected.
6 . The method according to claim 2 , wherein the top-k alpha vectors are iteratively calculated until the alpha vectors are converged.
7 . The method according to claim 3 , wherein the alpha vectors are iteratively calculated until the alpha vectors are converged.
8 . The method according to claim 2 , wherein the k is determined by how many alternative alpha vectors are required in the execution-time process of the POMDP model.
9 . The method according to claim 3 , wherein the k is determined by how many alternative actions are required in the execution-time process of the POMDP model.
10 . The method according to claim 1 , wherein the constraint restricts selection of a same action in succession.
11 . The method according to claim 1 , wherein an action not subject to the constraint in the execution-time process of the POMDP model is prepared.
12 . The method according to claim 1 , wherein actions having similar meaning but different expressions are prepared when an action is a natural conversation dialog.Join the waitlist — get patent alerts
Track US2018197100A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.