Apparatus and methods for object manipulation via action sequence optimization
Abstract
Methods, apparatus, systems and articles of manufacture are disclosed for object manipulation via action sequence optimization. An example method disclosed herein includes determining an initial state of a scene, generating a first action phase sequence to transform the initial state of the scene to a solution state of the scene by selecting a plurality of action phases based on action phase probabilities, determining whether a first simulated outcome of executing the first action phase sequence satisfies an acceptability criterion and, when the first simulated outcome does not satisfy the acceptability criterion, calculating a first cost function output based on a difference between the first simulated outcome and the solution state of the scene, the first cost function output utilized to generate updated action phase probabilities.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A non-transitory computer readable medium comprising computer readable instructions to simulate execution of an action sequence by a robot to convert an initial state to a goal state, the computer readable instructions, when executed, to cause processor circuitry to:
given the goal state, generate the action sequence to achieve the goal state by:
selecting a first action based on a first probability that the first action will occur given the initial state; and
selecting a second action based on a second probability that the second action will occur;
compare a result of the action sequence to the goal state to determine whether the result satisfies a criterion; when the criterion is not satisfied, determine gradients for at least the first action and the second action; and adjust generation of the action sequence based on the gradients.
22 . The non-transitory computer readable medium of claim 21 , wherein the criterion is an acceptability criterion.
23 . The non-transitory computer readable medium of claim 21 , wherein the instructions cause the processor circuitry to adjust the generation of the action sequence based on the gradients to satisfy an acceptability criterion.
24 . The non-transitory computer readable medium of claim 21 , wherein the initial state corresponds to one or more initial positions of one or more objects in an environment and the goal state corresponds to one or more goal positions of the one or more objects in the environment.
25 . The non-transitory computer readable medium of claim 21 , wherein the instructions cause the processor circuitry to conduct a simulation of executing the action sequence on the robot.
26 . The non-transitory computer readable medium of claim 21 , wherein the instructions cause the processor circuitry to:
determine a difference between the result of the action sequence and the goal state; and determine whether the result satisfies the criterion based on the difference.
27 . The non-transitory computer readable medium of claim 21 , wherein the instructions cause the processor circuitry to determine the gradients for at least the first action and the second action via gradient descent.
28 . An apparatus to simulate execution of an action sequence by a robot to convert an initial state to a goal state, the apparatus comprising:
memory; instructions; and processor circuitry to execute the instructions to:
given the goal state, generate the action sequence to achieve the goal state by:
selecting a first action based on a first probability that the first action will occur given the initial state; and
selecting a second action based on a second probability that the second action will occur;
compare a result of the action sequence to the goal state to determine whether the result satisfies a criterion;
when the criterion is not satisfied, determine gradients for at least the first action and the second action; and
adjust generation of the action sequence based on the gradients.
29 . The apparatus of claim 28 , wherein the criterion is an acceptability criterion.
30 . The apparatus of claim 28 , wherein the processor circuitry is to execute the instructions to adjust the generation of the action sequence based on the gradients to satisfy an acceptability criterion.
31 . The apparatus of claim 28 , wherein the initial state corresponds to one or more initial positions of one or more objects in an environment and the goal state corresponds to one or more goal positions of the one or more objects in the environment.
32 . The apparatus of claim 28 , wherein the processor circuitry is to execute the instructions to conduct a simulation of executing the action sequence on the robot.
33 . The apparatus of claim 28 , wherein the processor circuitry is to execute the instructions to:
determine a difference between the result of the action sequence and the goal state; and determine whether the result satisfies the criterion based on the difference.
34 . The apparatus of claim 28 , wherein the processor circuitry is to execute the instructions to determine the gradients for at least the first action and the second action via gradient descent.
35 . A method for simulating execution of an action sequence by a robot to convert an initial state to a goal state, the method comprising:
given the goal state, generating the action sequence to achieve the goal state by:
selecting, by executing an instruction with processor circuitry, a first action based on a first probability that the first action will occur given the initial state; and
selecting, by executing an instruction with the processor circuitry, a second action based on a second probability that the second action will occur;
comparing, by executing an instruction with the processor circuitry, a result of the action sequence to the goal state to determine whether the result satisfies a criterion; determining, when the criterion is not satisfied and by executing an instruction with the processor circuitry, gradients for at least the first action and the second action; and adjusting, by executing an instruction with the processor circuitry, generation of the action sequence based on the gradients.
36 . The method of claim 35 , wherein the criterion is an acceptability criterion.
37 . The method of claim 35 , further including adjusting the generation of the action sequence based on the gradients to satisfy an acceptability criterion.
38 . The method of claim 35 , wherein the initial state corresponds to one or more initial positions of one or more objects in an environment and the goal state corresponds to one or more goal positions of the one or more objects in the environment.
39 . The method of claim 35 , further including conducting a simulation of executing the action sequence on the robot.
40 . The method of claim 35 , further including:
determining a difference between the result of the action sequence and the goal state; and determining whether the result satisfies the criterion based on the difference.
41 . The method of claim 35 , further including determining the gradients for at least the first action and the second action via gradient descent.Join the waitlist — get patent alerts
Track US2022193895A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.