Systems and Methods for Generating Motion Forecast Data for a Plurality of Actors with Respect to an Autonomous Vehicle
Abstract
A computing system can input first relative location embedding data into an interaction transformer model and receive, as an output of the interaction transformer model, motion forecast data for actors relative to a vehicle. The computing system can input the motion forecast data into a prediction model to receive respective trajectories for the actors for a current time step and respective projected trajectories for the actors for a subsequent time step. The computing system can generate second relative location embedding data based on the respective projected trajectories from the second time step. The computing system can produce second motion forecast data using the interaction transformer model based on the second relative location embedding. The computing system can determine second respective trajectories for the actors using the prediction model based on the second forecast data.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, the method comprising:
accessing, for a first iteration corresponding with a first time step, a first relative location embedding that describes respective relative locations of a plurality of actors with respect to an autonomous vehicle; processing the first relative location embedding by a machine-learned interaction transformer model to generate first motion forecast data describing movement of the plurality of actors at the first time step; generating, based on the first motion forecast data, a second relative location embedding for a second time step; processing the second relative location embedding by the machine-learned interaction transformer model to generate second motion forecast data for the plurality of actors at the second time step; and controlling motion of the autonomous vehicle based on at least one of the first motion forecast data or the second motion forecast data.
22 . The computer-implemented method of claim 21 , wherein processing the first relative location embedding by a machine-learned interaction transformer model to generate first motion forecast data describing movement of the plurality of actors at the first time step further comprises:
processing the first relative location embedding and a first input sequence describing object detection data corresponding with the first time step by the machine-learned interaction transformer model to generate the first motion forecast data describing movement of the plurality of actors at the first time step.
23 . The computer-implemented method of claim 21 , wherein processing the second relative location embedding by a machine-learned interaction transformer model to generate second motion forecast data describing movement of the plurality of actors at the second time step further comprises:
processing the second relative location embedding and a second input sequence describing object detection data corresponding with the second time step by the machine-learned interaction transformer model to generate the second motion forecast data describing movement of the plurality of actors at the second time step.
24 . The computer-implemented method of claim 23 , wherein the second input sequence is based at least in part on the first motion forecast data.
25 . The computer-implemented method of claim 23 , wherein the second input sequence further describes one or more features for the plurality of actors respectively at the first time step or the second time step, the one or more features comprising at least one of:
a derivative of the respective relative location of a respective actor of the plurality of actors relative to the autonomous vehicle; a size of the respective actor; an orientation of the respective actor relative to the autonomous vehicle; or a center location of the respective actor.
26 . The computer-implemented method of claim 21 , wherein one or more of the first relative location embedding and the second relative location embedding respectively, encodes the respective relative locations of the plurality of actors as a multi-channel positional embedding.
27 . The computer-implemented method of claim 21 , further comprising:
processing the second motion forecast data with a machine-learned prediction model to generate respective trajectories of the plurality of actors for the first time step and respective projected trajectories of the plurality of actors for the second time step; and controlling the motion of the autonomous vehicle further based on the respective trajectories and the respective projected trajectories generated by the machine-learned prediction model.
28 . The computer-implemented method of claim 27 , wherein the machine-learned prediction model comprises one or more multi-layer perceptrons.
29 . The computer-implemented method of claim 27 , further comprising:
generating respective trajectory sequences for the plurality of actors, the respective trajectory sequences comprising the respective trajectories of the plurality of actors for the first time step and the respective projected trajectories of the plurality of actors for the second time step.
30 . The computer-implemented method of claim 21 , wherein the machine-learned interaction transformer model comprises:
a machine-learned interaction model configured to receive the first or second relative location embedding that describes the respective relative locations of the plurality of actors with respect to the autonomous vehicle, and in response to receipt of the relative location embedding, generate an attention embedding with respect to the plurality of actors; a machine-learned recurrent model configured to receive the attention embedding, and in response to receipt of the attention embedding, generate the first or second motion forecast data with respect to the plurality of actors.
31 . The computer-implemented method of claim 21 , wherein the machine-learned interaction transformer model comprises one or more neural networks.
32 . A computing system for controlling motion of an autonomous vehicle, the computing system comprising:
one or more processors; and one or more non-transitory computer-readable media that store instructions for execution by the one or more processors to cause the one or more processors to perform operations comprising:
accessing, for a first iteration corresponding with a first time step, a first relative location embedding that describes respective relative locations of a plurality of actors with respect to the autonomous vehicle;
processing the first relative location embedding by a machine-learned interaction transformer model to generate first motion forecast data describing movement of the plurality of actors at the first time step;
generating, based on the first motion forecast data, a second relative location embedding for a second time step;
processing the second relative location embedding by the machine-learned interaction transformer model to generate second motion forecast data for the plurality of actors at the second time step; and
controlling motion of the autonomous vehicle based on at least one of the first motion forecast data or the second motion forecast data.
33 . The computing system for controlling autonomous vehicle motion of claim 32 , wherein processing the first relative location embedding by a machine-learned interaction transformer model to generate first motion forecast data describing movement of the plurality of actors at the first time step further comprises:
processing the first relative location embedding and a first input sequence describing object detection data corresponding with the first time step by the machine-learned interaction transformer model to generate the first motion forecast data describing movement of the plurality of actors at the first time step.
34 . The computing system for controlling autonomous vehicle motion of claim 32 , wherein processing the second relative location embedding by a machine-learned interaction transformer model to generate second motion forecast data describing movement of the plurality of actors at the second time step further comprises:
processing the second relative location embedding and a second input sequence describing object detection data corresponding with the second time step by the machine-learned interaction transformer model to generate the second motion forecast data describing movement of the plurality of actors at the second time step.
35 . The computing system for controlling autonomous vehicle motion of claim 34 , wherein the second input sequence is based at least in part on the first motion forecast data.
36 . The computing system for controlling autonomous vehicle motion of claim 34 , wherein the second input sequence further describes one or more features for the plurality of actors respectively at the first time step or the second time step, the one or more features comprising at least one of:
a derivative of the respective relative location of a respective actor of the plurality of actors relative to the autonomous vehicle; a size of the respective actor; an orientation of the respective actor relative to the autonomous vehicle; or a center location of the respective actor.
37 . The computing system for controlling autonomous vehicle motion of claim 32 , wherein one or more of the first relative location embedding and the second relative location embedding respectively encodes the respective relative locations of the plurality of actors as a multi-channel positional embedding.
38 . The computing system for controlling autonomous vehicle motion of claim 32 , the operations further comprising:
processing the second motion forecast data with a machine-learned prediction model to generate respective trajectories of the plurality of actors for the first time step and respective projected trajectories of the plurality of actors for the second time step; and controlling the motion of the autonomous vehicle further based on the respective trajectories and the respective projected trajectories generated by the machine-learned prediction model.
39 . An autonomous vehicle, comprising:
one or more processors; and one or more non-transitory computer-readable media that store instructions for execution by the one or more processors to cause the one or more processors to perform operations comprising
accessing, for a first iteration corresponding with a first time step, a first relative location embedding that describes respective relative locations of a plurality of actors with respect to the autonomous vehicle;
processing the first relative location embedding by a machine-learned interaction transformer model to generate first motion forecast data describing movement of the plurality of actors at the first time step;
generating, based on the first motion forecast data, a second relative location embedding for a second time step;
processing the second relative location embedding by the machine-learned interaction transformer model to generate second motion forecast data for the plurality of actors at the second time step; and
controlling motion of the autonomous vehicle based on at least one of the first motion forecast data or the second motion forecast data.
40 . The autonomous vehicle of claim 39 , the operations further comprising:
processing the second motion forecast data with a machine-learned prediction model to generate respective trajectories of the plurality of actors for the first time step and respective projected trajectories of the plurality of actors for the second time step; and controlling the motion of the autonomous vehicle further based on the respective trajectories and the respective projected trajectories generated by the machine-learned prediction model.Join the waitlist — get patent alerts
Track US2024010241A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.