arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:1809.02926v1 [cs.LG] 09 Sep 2018

Probabilistic Prediction of Interactive Driving Behavior via
Hierarchical Inverse Reinforcement Learning

Liting Sun Affiliation: L. Sun, W. Zhan, and M. Tomizuka are with the Department of Mechanical Engineering, University of California, Berkeley, CA 94720 USA (e-mail: litingsun, wzhan, tomizuka@berkeley.edu).    Wei Zhan Affiliation: L. Sun, W. Zhan, and M. Tomizuka are with the Department of Mechanical Engineering, University of California, Berkeley, CA 94720 USA (e-mail: litingsun, wzhan, tomizuka@berkeley.edu).       Masayoshi Tomizuka Thanks: *This work was partially supported by the international Chair Drive for All, Foundation MINES ParisTech. Affiliation: L. Sun, W. Zhan, and M. Tomizuka are with the Department of Mechanical Engineering, University of California, Berkeley, CA 94720 USA (e-mail: litingsun, wzhan, tomizuka@berkeley.edu).
Abstract

Autonomous vehicles (AVs) are on the road. To safely and efficiently interact with other road participants, AVs have to accurately predict the behavior of surrounding vehicles and plan accordingly. Such prediction should be probabilistic, to address the uncertainties in human behavior. Such prediction should also be interactive, since the distribution over all possible trajectories of the predicted vehicle depends not only on historical information, but also on future plans of other vehicles that interact with it. To achieve such interaction-aware predictions, we propose a probabilistic prediction approach based on hierarchical inverse reinforcement learning (IRL). First, we explicitly consider the hierarchical trajectory-generation process of human drivers involving both discrete and continuous driving decisions. Based on this, the distribution over all future trajectories of the predicted vehicle is formulated as a mixture of distributions partitioned by the discrete decisions. Then we apply IRL hierarchically to learn the distributions from real human demonstrations. A case study for the ramp-merging driving scenario is provided. The quantitative results show that the proposed approach can accurately predict both the discrete driving decisions such as yield or pass as well as the continuous trajectories.

I Introduction

Autonomous vehicles (AVs) are interacting with human on the road. To this end, it needs to predict and reason about possible future behavior of human, and plan its own trajectories accordingly. Wrong prediction can cause either too conservative motions such as unnecessary stops/yielding, or dangerous situations like emergency brakes and unavoidable collisions. Hence, accurate prediction is of crucial importance to enable a safe and efficient autonomous car.

Starting from as simple as assuming that the other drivers would maintain their current velocities within the planning horizon [1][2] to as complicated as modelling the probability distributions over all possible trajectories/actions, many prediction approaches have been proposed. In terms of prediction output, they can be categorized into two groups: deterministic prediction [1, 2, 3, 4] and probabilistic prediction [5, 6, 7]. Both categories can be implemented based on various models such as neural networks (NN) [8][9], hidden Markov models (HMM) [10] and Bayes net [6]. Most of the work, however, formulated the future trajectories/actions as either deterministic functions or conditional probabilities of historical and current scene states. The influence of human drivers’ beliefs about the other vehicles future actions are ignored in the prediction framework.

In fact, in interactive driving scenarios, human drivers will actively anticipate and reason about surrounding vehicles’ behavior when they decide the next-step actions. This means that besides historical and current scene states, the distribution over all possible trajectories is also influenced by their beliefs about other vehicles’ plan. For example, given the same historical trajectories, a lane-keeping driver might be more probable to decelerate rather than maintaining current speed if the driver thinks another vehicle is about to merge into his lane. Such interaction-aware prediction is incorporated implicitly into a planning framework in [11], but the prediction output is deterministic.

Our insight is that in highly interactive driving scenarios, an accurate prediction should consider not only the time-domain dependency but also the interaction among vehicles. Namely, an autonomous car should anticipate the conditional probability over all possible trajectories of the human driver given not only historical and current states, but also its own future plans.

To obtain such a probabilistic and interactive prediction of human drivers’ behavior, we need to approximate the “internal incentive” that drives human to generate specific behavior. Note that human’s planning procedure is naturally hierarchical, involving both discrete and continuous driving decisions. The discrete driving decisions determine the game-theoretic outcomes of interaction such as to yield or to pass, whereas the continuous driving decisions influence details of the resulting trajectories in terms of smoothness, distances to other road participants and higher-order dynamics such as velocities, accelerations and jerks. Hence, to describe the influence of both discrete and continuous decisions, we propose a hierarchical inverse reinforcement learning (IRL) framework in this paper, to learn trajectory-generation process of human from observed demonstrations. Given the autonomous vehicle’s future plan, the probability distribution over all future trajectories of the predicted human driver is modelled as a mixture of distributions, partitioned by the discrete driving decisions.

Contributions of this paper are the following two.

A formalism for probabilistic interactive prediction of human driving behavior. We model the prediction problem from the perspective of a two-agent game by explicitly considering the responses of one agent to another. We formalize the distribution of predicted vehicle’s future trajectories as a probability distribution conditioned not only on historical and current information, but also on future plans of the ego vehicle.

A hierarchical IRL framework. We explicitly consider the hierarchical planning procedure of human drivers, and formulate the influences of both discrete and continuous driving decisions. Via hierarchical IRL, the conditional distribution over all possible future trajectories is expressed as a mixture of probability distributions partitioned by different game-theoretic outcomes of interactive vehicles.

Related Work in IRL: Initially proposed by Kalman [12], the concept of Inverse Reinforcement Learning is first formulated in [13]. It aims to infer the reward/cost functions of agents from their observed behavior by assuming that the agents are rational. To deal with uncertainties and noisy observations, Ziebart et al[14] extended the algorithm based on the principle of maximum entropy. It assumes that agents actions/behavior with lower cost are exponentially more probable, and thus an exponential distribution family can be established to approximate the distribution of actions/behavior. Building on this, Levine et al[15] formulated the continuous IRL algorithm and used it to minic and predict human driving behavior, and Liu et al[16] used IRL to approximate the decision-making process of taxi drivers and passengers in public transportation. In [17], Shimosaka et al. predicted human driving behavior in diverse environment by formulating IRL with multiple reward functions. Kretzschmar et al[18] considered the hierarchical trajectory-generation process of human navigation behavior, but they focus on cooperative agents that share the same reward function instead of interactive agents that have significantly different reward functions.

II Problem Statement

We focus on predicting human drivers’ interactive behavior within two vehicles: a host vehicle (denoted by ()H(\cdot)_{H}) and a predicted vehicle (denoted by ()M(\cdot)_{M}). All other in-scene vehicles are treated as surrounding vehicles, denoted by ()Oi(\cdot)^{i}_{O} (i=1,2,,Ni{=}1,2,{\cdots},N is the index for surrounding vehicles). We use ξ\xi to represent historical vehicle trajectories, and ξ^\hat{\xi} for future trajectories11 1 Note that trajectory is a sequence of states, i.e., ξ=[x1T,x2T,,xLT]T\xi=[x^{T}_{1},x^{T}_{2},\cdots,x^{T}_{L}]^{T} where xix_{i} is the vehicle state at ii-th time step. LL is the trajectory length. Depending on the representation, the vehicle state can be different. For instance, it can simply be the positions of vehicles, or it can include velocities, yaw angles, accelerations, etc..

It is obvious that the probability distribution over all possible trajectories of the predicted vehicle depends on his own historical trajectory and those of surrounding vehicles. Mathematically, such time-domain state dependency can be modelled as a conditional probability density function (PDF), as in [19]:

p(ξ^M|ξO1:N,ξH,ξM).\displaystyle p(\hat{\xi}_{M}|\xi^{1:N}_{O},\xi_{H},\xi_{M}). (1)

However, as discussed above, in interactive driving scenarios, the influence of human’s beliefs about others’ next-step actions cannot be ignored in order to get a good prediction. From this perspective, the probability distribution over all possible trajectories of the predicted vehicle should also be conditioned on potential plans of the host vehicle. Mathematically, such interaction-dependency further refines the conditional PDF in (1) as

p(ξ^M|ξO1:N,ξH,ξM,ξ^H).\displaystyle p(\hat{\xi}_{M}|\xi^{1:N}_{O},\xi_{H},\xi_{M},\hat{\xi}_{H}). (2)

The key aspect of this formulation is to actively consider the influence of the host vehicle’s actions when predicting the future trajectories of the predicted vehicle. In this sense, we are trying to enable the autonomous vehicle to “think” in the way that the predicted human driver thinks if he had known the autonomous vehicle’s future plan.

III Modeling Human Driving Behavior

In order to predict human’s interactive driving trajectories, we need to model the internal incentive of human that generates his behavior.

III-A Probabilistic Hierarchical Trajectory-Generation Process

As addressed above, the trajectory-generation process of human drivers is naturally probabilistic and hierarchical. It involves both discrete and continuous driving decisions. The discrete driving decisions determine “rough” pattern (or homotopy class) of his future trajectory as a game-theoretic result (e.g., to yield or to pass), whereas the continuous driving decisions influence details of the trajectory such as velocities, accelerations and smoothness.

Figure 1 illustrates such probabilistic and hierarchical trajectory–generation process for a lane-changing driving scenario. The predicted vehicle (blue) is trying to merge into the lane occupied by the host vehicle (red). Given all observed historical trajectories ξ={ξO1:N,ξH,ξM}\xi{=}\{\xi^{1:N}_{O},\xi_{H},\xi_{M}\} and his belief about the host vehicle’s future trajectory ξ^H\hat{\xi}_{H}, he first decides whether to merge behind the host vehicle (dM1d^{1}_{M}) or merge front (dM2d^{2}_{M}). Such discrete driving decisions are outcomes of the first-layer probability distribution P(dM|ξ,ξ^)P(d_{M}|\xi,\hat{\xi}), and partition the space of all possible trajectories into two distinct homotopy classes22 2 Two trajectories belong to a same homotopy class if they can be continuously transformed to each other without collisions [18]., each of them can be described via a second-layer probability distribution p(ξ^M|dM,ξ,ξ^H)p(\hat{\xi}_{M}|d_{M},\xi,\hat{\xi}_{H}). Note that among different homotopy classes, the distributions of the continuous trajectories can be significantly different since the driving preference under different discrete decisions may be different. For example, a vehicle decided to merge behind might care more about comfort and less about speed, while a vehicle merging front might care about completely the opposite.

Refer to caption
Fig. 1: The probabilistic and hierarchical trajectory–generation process for a lane changing scenario. The predicted vehicle (blue) is trying to merge into the lane of the host vehicle (red). Given all observed historical trajectories ξ={ξO1:N,ξH,ξM}\xi{=}\{\xi^{1:N}_{O},\xi_{H},\xi_{M}\} and his belief about the host vehicle’s future trajectory ξ^H\hat{\xi}_{H}, the trajectory distribution of the predicted vehicle over all the trajectory space is partitioned first by the discrete decisions: merge behind (dM1d^{1}_{M}) and merge front (dM2d^{2}_{M}). Under different discrete decisions, the distributions of continuous trajectories can be significantly different, and each of them is represented via a probability distribution model. The observed demonstrations are samples satisfying the distributions.

Hence, the conditional distribution p(ξ^M|ξ,ξ^H)p(\hat{\xi}_{M}|\xi,\hat{\xi}_{H}) in (2) is formulated as a mixture of distributions, which explicitly captures the influences of both discrete and continuous driving decisions of human drivers:

p(ξ^M|ξ,ξ^H)=dMi𝒟Mp(ξ^M|dMi,ξ,ξ^H)P(dMi|ξ,ξ^H)p(\hat{\xi}_{M}|\xi,\hat{\xi}_{H})=\sum_{d^{i}_{M}\in\mathcal{D}_{M}}{p(\hat{\xi}_{M}|d^{i}_{M},\xi,\hat{\xi}_{H})P(d^{i}_{M}|\xi,\hat{\xi}_{H})} (3)

where 𝒟M\mathcal{D}_{M} represents the set of all possible discrete decisions for the predicted vehicle.

III-B Hierarchical Inverse Reinforcement Learning

Equation (3) suggests that in order to model the conditional PDF in (2) for interactive prediction, we need to model the hierarchical probabilistic models P(dM|ξ,ξ^H)P(d_{M}|\xi,\hat{\xi}_{H}) and p(ξ^M|dMi,ξ,ξ^H)p(\hat{\xi}_{M}|d^{i}_{M},\xi,\hat{\xi}_{H}) for each dMi𝒟Md^{i}_{M}{\in}\mathcal{D}_{M}.

We thus propose to apply inverse reinforcement learning hierarchically to learn all the models from observed demonstrations of human drivers. Based on the principle of maximum entropy [14], we assume that all drivers are exponentially more likely to make decisions (both discrete and continuous) that lead to a lower cost. This introduces an family of exponential distributions depending on the cost functions, and our interest is to find the optimal hierarchical cost functions that ultimately lead to trajectory distributions matching that of the observed trajectories in a given demonstration set Ξ\Xi.

III-C Modeling Continuous Driving Decisions

Suppose that the demonstration set Ξ\Xi is partitioned into |𝒟||\mathcal{D}| subsets by the discrete decisions dMi𝒟d^{i}_{M}{\in}\mathcal{D}. |𝒟||\mathcal{D}| is the dimension of 𝒟\mathcal{D}. Each subset ΞdMi\Xi_{d^{i}_{M}} contains trajectories that belong to the same homotopy class. We represent each demonstration in Ξi\Xi_{i} by a tuple (ξ^M,dMi,ξ,ξ^H)(\hat{\xi}_{M},d_{M}^{i},\xi,\hat{\xi}_{H}) where ξ=[ξO1:N,ξH,ξM]\xi{=}[\xi^{1:N}_{O},\xi_{H},\xi_{M}] represents all historical information. Since the trajectory space is continuous and demonstrations are noisily local-optimal, we use Continuous Inverse Optimal Control with Locally Optimal Examples[15].

III-C1 Continuous-Space IRL

Under discrete decision dMd_{M}, we assume that the cost of each trajectory can be linearly parametrized by a group of selected features {𝐟dM}\{\mathbf{f}_{d_{M}}\}, i.e., C(𝜽dM,ξ^M,ξ,ξ^H)=𝜽dMT𝐟dM(ξ^M,ξ,ξ^H)C(\boldsymbol{\theta}_{d_{M}},\hat{\xi}_{M},\xi,\hat{\xi}_{H})=\mathbf{\boldsymbol{\theta}}^{T}_{d_{M}}\mathbf{f}_{d_{M}}(\hat{\xi}_{M},\xi,\hat{\xi}_{H}), where θdM\theta_{d_{M}} is the parameter vector to determine the emphasis of each of the features. Then trajectories with higher cost are exponentially less likely based on the principle of maximum entropy :

P(ξ^M|𝜽dM,dMi,ξ,ξ^H)eC(𝜽dM,ξ^M,ξ,ξ^H)P(\hat{\xi}_{M}|\boldsymbol{\theta}_{d_{M}},d^{i}_{M},\xi,\hat{\xi}_{H})\propto e^{-C(\boldsymbol{\theta}_{d_{M}},\hat{\xi}_{M},\xi,\hat{\xi}_{H})} (4)

Hence, the log likelihood of the given demonstration subset {ΞdM}\{\Xi_{d_{M}}\} is given by

logP(ΞdM|𝜽dM)=ξ^MΞdMlogeC(𝜽dM,ξ^M,ξ,ξ^H)eC(𝜽dM,ξ~M,ξ,ξ^H)dξ~M.\displaystyle\log P(\Xi_{d_{M}}|\boldsymbol{\theta}_{d_{M}}){=}\sum_{\hat{\xi}_{M}\in\Xi_{d_{M}}}\log\dfrac{e^{-C(\boldsymbol{\theta}_{d_{M}},\hat{\xi}_{M},\xi,\hat{\xi}_{H})}}{\int e^{-C(\boldsymbol{\theta}_{d_{M}},\tilde{\xi}_{M},\xi,\hat{\xi}_{H})}d\tilde{\xi}_{M}}. (5)

Our goal is to find the optimal 𝜽dM\boldsymbol{\theta}_{d_{M}} such that the given demonstration set is most likely to happen:

𝜽dM=argmax𝜽dMP(ΞdM|𝜽dM)\boldsymbol{\theta}_{d_{M}}^{*}=\arg\max_{\boldsymbol{\theta}_{d_{M}}}P(\Xi_{d_{M}}|\boldsymbol{\theta}_{d_{M}}) (6)

To tackle the normalization factor in (5), we use Laplace approximation as in [15]. Namely, the cost along an arbitrary trajectory ξ~M\tilde{\xi}_{M} is approximated by its second-order Taylor expansion around the demonstrated trajectory ξ^M\hat{\xi}_{M}:

C(𝜽dM,ξ~M,ξ,ξ^H)C(𝜽dM,ξ^M,ξ,ξ^H)+(ξ~Mξ^M)TCξM\displaystyle C(\boldsymbol{\theta}_{d_{M}},\tilde{\xi}_{M},\xi,\hat{\xi}_{H}){\approx}C(\boldsymbol{\theta}_{d_{M}},\hat{\xi}_{M},\xi,\hat{\xi}_{H})+(\tilde{\xi}_{M}-\hat{\xi}_{M})^{T}\dfrac{\partial C}{\partial\xi_{M}}
+(ξ~Mξ^M)T2CξM2(ξ~Mξ^M)\displaystyle+(\tilde{\xi}_{M}-\hat{\xi}_{M})^{T}\dfrac{\partial^{2}C}{\partial\xi_{M}^{2}}(\tilde{\xi}_{M}-\hat{\xi}_{M})\qquad

This enables the normalization factor become a Gaussian integral and can be solved analytically. Define gξ^M(𝜽dM)=CξM|ξ^Mg_{\hat{\xi}_{M}}(\boldsymbol{\theta}_{d_{M}}){=}\frac{\partial C}{\partial\xi_{M}}|_{\hat{\xi}_{M}} and Hξ^M(𝜽dM)=2CξM2|ξ^MH_{\hat{\xi}_{M}}(\boldsymbol{\theta}_{d_{M}}){=}\frac{\partial^{2}C}{\partial\xi_{M}^{2}}|_{\hat{\xi}_{M}} as, respectively, the gradient and Hessian of the cost along trajectory ξ^M\hat{\xi}_{M}. Then (6) is translated to the optimization problem in (7), which intuitively means the optimal cost function should have small gradient and large positive Hessians along the demonstrated trajectories. For details, one can refer to [15].

minξ^MΞdM𝜽dMgξ^MT(𝜽dM)Hξ^M1(𝜽dM)gξ^M(𝜽dM)log|Hξ^M(𝜽dM)|\min_{\boldsymbol{\theta}_{d_{M}}}\sum_{\hat{\xi}_{M}\in\Xi_{d_{M}}}g_{\hat{\xi}_{M}}^{T}(\boldsymbol{\theta}_{d_{M}})H_{\hat{\xi}_{M}}^{-1}(\boldsymbol{\theta}_{d_{M}})g_{\hat{\xi}_{M}}(\boldsymbol{\theta}_{d_{M}})-\log|H_{\hat{\xi}_{M}}(\boldsymbol{\theta}_{d_{M}})| (7)

III-C2 Features

The features we selected to parametrize the continuous trajectories can be grouped as follows:

  • Speed - The incentive of the human driver to reach a certain speed limit vlimv_{\lim} is captured by the feature

    fv(ξ^M)=t=0L(vtvlim)2f_{v}(\hat{\xi}_{M})=\sum_{t=0}^{L}(v_{t}-v_{\lim})^{2} (8)

    vtv_{t} is the speed at time tt along trajectory ξ^M\hat{\xi}_{M} and LL is the length of the trajectory.

  • Traffic - In dense traffic environment, human drivers tend to follow the traffic. Hence, we introduce a feature based on the intelligent driver model (IDM) [20]

    fIDM(ξ^M)=t=0L(ststIDM)2f_{\text{IDM}}(\hat{\xi}_{M})=\sum_{t=0}^{L}(s_{t}-s_{t}^{\text{IDM}})^{2} (9)

    where sts_{t} is the actual spatial headway between the front vehicle and predicted vehicle at time tt along trajectory ξ^M\hat{\xi}_{M}, and stIDMs_{t}^{\text{IDM}} is the spatial headway suggested by IDM.

  • Control effort and smoothness - Human drivers typically prefer to drive efficiently and smoothly, avoiding unnecessary accelerations and jerks. To address such preference, we introduce a set of kinematics-related features:

    facc(ξ^M)=t=0Lat2,fjerk(ξ^M)=t=1L(atat1t)2f_{\text{acc}}(\hat{\xi}_{M})=\sum_{t=0}^{L}a_{t}^{2},\quad f_{\text{jerk}}(\hat{\xi}_{M})=\sum_{t=1}^{L}\left(\dfrac{a_{t}-a_{t-1}}{\triangle t}\right)^{2} (10)

    where ata_{t} represents the acceleration at time tt along the trajectory ξ^M\hat{\xi}_{M}. t\triangle t is the sampling time.

  • Clearance to other road participants - Human drivers care about their distances to other road participants when they drive since distance is crucially related to safety. Hence, we introduce a distance-related feature

    fdist(ξ^M)=t=0Lk=1N+1e(xtxtk)2l2(ytytk)2w2f_{\text{dist}}(\hat{\xi}_{M})=\sum_{t=0}^{L}\sum_{k=1}^{N{+}1}e^{-\frac{(x_{t}-x^{k}_{t})^{2}}{l^{2}}-\frac{(y_{t}-y^{k}_{t})^{2}}{w^{2}}} (11)

    where (xt,yt)(x_{t},y_{t}) and (xtk,ytk)(x^{k}_{t},y^{k}_{t}) represent, respectively, the coordinates of the predicted vehicle along ξ^M\hat{\xi}_{M} and those of the kk-th surrounding vehicle. Parameters ll and ww are the length and width of the predicted vehicle. We use coordinates in Frenet Frame to deal with curvy roads, i.e., xx denotes the travelled distance along the road and yy is the lateral deviation from the lane center.

  • Goal - This feature describes the short-term goals of human drivers. Typically, goals are determined by the discrete driving decisions. For instance, if a lane-changing vehicle decides to merge in front of a host vehicle on his target lane, he will set his short-term goals to be ahead of the host vehicle. The goal-related feature is given by

    fg(ξ^M)=t=0L(xt,yt)(xtg,ytg)22f_{g}(\hat{\xi}_{M})=\sum_{t=0}^{L}\|(x_{t},y_{t})-(x_{t}^{g},y_{t}^{g})\|^{2}_{2} (12)
  • Courtesy - Most of human drivers view driving as a social behavior, meaning that they not only care about their own cost, but also care about others’ cost, particularly when they are merging into others’ lanes [21]. Suppose that the cost of the host vehicle is CH(ξ^M,ξ,ξ^H)C_{H}(\hat{\xi}_{M},\xi,\hat{\xi}_{H}), then to address the influence of courtesy to the interaction, we introduce the feature

    fcourt(ξ^M)=max{CH(ξ^M,ξ,ξ^H)CHdefault,0}.f_{\text{court}}(\hat{\xi}_{M})=\max\left\{C_{H}\left(\hat{\xi}_{M},\xi,\hat{\xi}_{H}\right){-}C_{H}^{\text{default}},0\right\}. (13)

    This feature describes the possible extra cost brought by the trajectory ξ^M\hat{\xi}_{M} of the predicted vehicle to the host vehicle, compared to the default host vehicle’s cost CHdefaultC_{H}^{\text{default}}. We can learn about CH()C_{H}(\cdot) also from demonstrations. Details about this will be covered in the case study.

For vehicles in different driving scenarios, or under different discrete driving decisions, the features we used to parametrize their costs are different subsets of the above listed ones. For instance, drivers decided to merge behind or front would set different goals, and drivers with right of way would most likely care less about courtesy than those without.

Remark I: Note that all variables vt,at,δt,xtv_{t},a_{t},\delta_{t},x_{t} and yty_{t} in (8)-(13) can be expressed as functions of trajectory ξ^M\hat{\xi}_{M}. For instance, if we define ξ^M=[x0,y0,,xL,yL]T\hat{\xi}_{M}=[x_{0},y_{0},\cdots,x_{L},y_{L}]^{T} where xx and yy are the coordinates in Frenet Frame, then we can obtain all variables via numerical differentiation. Details are omitted.

III-D Modeling Discrete Driving Decisions

III-D1 Features

Different from the continuous driving decisions that influence the higher-order dynamics of the trajectories, the discrete decisions determine the homotopy of the trajectories. To capture this, we selected the following two features to parametrize the cost function that induces the discrete decisions:

  • Rotation angle - To describe the interactive driving behavior such as overtaking from left or right sides, merging in from front or back, we compute the rotation angle from (xt,yt)(x_{t},y_{t}) to (xt,H,yt,H)(x_{t,H},y_{t,H}) along trajectory ξ^MΞdM\hat{\xi}_{M}{\in}\Xi_{d_{M}}. Define the angle as ωt\omega_{t}, and then the rotation angle feature is given by

    f(dM)=t=0Lωtf_{\angle}(d_{M})=\sum_{t=0}^{L}\omega_{t} (14)

    where dMd_{M} is a discrete decision and LL is the length of the trajectory ξ^M\hat{\xi}_{M}.

  • Minimum cost - It is also possible that human drivers make discrete decisions by evaluating the cost of trajectories under each decision, and select the one leading to the minimum-cost trajectory. To address this factor, we consider the feature

    fcost(dM)=minξ~M𝜽dMT𝐟dM(ξ~M,ξ,ξ^H).f_{\text{cost}}(d_{M})=\min_{\tilde{\xi}_{M}}\boldsymbol{\theta}_{{d}_{M}}^{T}\mathbf{f}_{d_{M}}(\tilde{\xi}_{M},\xi,\hat{\xi}_{H}). (15)

    where 𝜽dM\boldsymbol{\theta}_{{d}_{M}} and 𝐟dM\mathbf{f}_{d_{M}}, respectively, represent the learned parameters and selected features for the continuous trajectory distribution under discrete decision dMd_{M}.

III-D2 Discrete-Space IRL

Similarly, we assume that the decisions with lower cost are exponentially more probable. We also assume that the cost function is linearly parametrized by 𝝍\boldsymbol{\psi} and feature vector 𝐟d=[f,fcost]T\mathbf{f}^{\text{d}}=[f_{\angle},f_{\text{cost}}]^{T}, i.e., Cd(dM,ψ)=𝝍T𝐟d(dM)C^{\text{d}}(d_{M},\psi)=\boldsymbol{\psi}^{T}\mathbf{f}^{\text{d}}(d_{M}). Again, our goal is to find the optimal 𝝍\boldsymbol{\psi}^{*} such that the likelihood of the demonstration set Ξ\Xi is maximized:

max𝝍P(Ξ|𝝍)=maxdMΞ𝝍e𝝍T𝐟d(dM)d~M𝒟e𝝍T𝐟d(d~M)\max_{\boldsymbol{\psi}}P(\Xi|\boldsymbol{\psi})=\max_{\boldsymbol{\psi}}\prod_{d_{M}\in\Xi}\dfrac{e^{-\boldsymbol{\psi}^{T}\mathbf{f}^{\text{d}}(d_{M})}}{\sum_{\tilde{d}_{M}\in\mathcal{D}}e^{-\boldsymbol{\psi}^{T}\mathbf{f}^{\text{d}}(\tilde{d}_{M})}} (16)

By taking the log probability and gradient-descent approach, the parameter 𝝍\boldsymbol{\psi} is updated via

𝝍s+1=𝝍sα(1NΞ𝐟d(dM)d~MP(d~M|𝝍s)fd(d~M))\boldsymbol{\psi}_{s+1}{=}\boldsymbol{\psi}_{s}{-}\alpha\left(\frac{1}{N}\sum_{\Xi}\mathbf{f}^{\text{d}}(d_{M}){-}{\sum}_{\tilde{d}_{M}}P(\tilde{d}_{M}|\boldsymbol{\psi}_{s})f^{\text{d}}(\tilde{d}_{M})\right) (17)
P(d~M|𝝍s)=e𝝍sT𝐟d(d~M)d~M𝒟e𝝍sT𝐟d(d~M)P(\tilde{d}_{M}|\boldsymbol{\psi}_{s})=\dfrac{e^{-\boldsymbol{\psi}_{s}^{T}\mathbf{f}^{\text{d}}(\tilde{d}_{M})}}{\sum_{\tilde{d}_{M}\in\mathcal{D}}e^{-\boldsymbol{\psi}_{s}^{T}\mathbf{f}^{\text{d}}(\tilde{d}_{M})}}\qquad\qquad\quad (18)

where α\alpha is the step size and NN is the number of demonstrated trajectories in set Ξ\Xi. To estimate the probability P(d~M|𝝍s)P(\tilde{d}_{M}|\boldsymbol{\psi}_{s}), we use sampling-based method. Detailed implementation will be covered later in the case study.

IV Case Study

In this section, we apply the proposed hierarchical IRL approach to model and predict the interactive human driving behavior in a ramp-merging scenario.

IV-A Data Collection

We collect human driving data from the Next Generation SIMulation (NGSIM) dataset [22]. It captures the highway driving behaviors/trajectories by cameras mounted on top of surrounding buildings. The sampling time of the trajectories is t=0.01\triangle t{=}0.01s. We choose 134 ramp-merging trajectories on Interstate 80 (near Emeryville, California), and separated them into two sets: a training set of size 80 (denoted by Ξ\Xi, i.e., the human demonstrations), and the other 54 trajectories as the test set.

Figure 2 shows the road map (See [22] for detailed geometry information) and an example group of trajectories. There are four cars in scene, one merging vehicle (red), one lane-keeping vehicle (blue) and two surrounding vehicles (black), with one ahead of the blue car and the other behind. Our interest focuses on the interactive driving behavior of both the merging vehicle and the lane-keeping vehicle.

Refer to caption
Fig. 2: The merging map on Interstate 80 near Emeryville, California. Red: merging vehicle; Blue: lane-keeping vehicle; Black: other surrounding vehicles

IV-B Driving Decisions and Feature Selection

We use the same hierarchical IRL approach to model the conditional probability distributions for both the merging vehicle and the lane-keeping vehicle.

IV-B1 Driving Decisions

In the ramp-merging scenario, the driving decisions are listed as in Table I. As mentioned above, 𝐱=[x1,,xL]T\mathbf{x}{=}[x_{1},\cdots,x_{L}]^{T} and 𝐲=[y1,,yL]T\mathbf{y}{=}[y_{1},\cdots,y_{L}]^{T} are, respectively, the coordinate vectors in Frenet Frame along the longitudinal and lateral directions. LL is set to be 5050, i.e., in each demonstration, 55s trajectories are collected.

Discrete Decisions Continuous Decisions
merging-in 𝒟=\mathcal{D}= trajectory
vehicle {merge front, merge back}\{\text{merge front, merge back}\} ξ=[x1,y1,,xL,yL]T\xi{=}[x_{1},y_{1},\cdots,x_{L},y_{L}]^{T}
lane-keeping 𝒟=\mathcal{D}= trajectory
vehicle {yield, pass}\{\text{yield, pass}\} ξ=[x1,y1,,xL,yL]T\xi{=}[x_{1},y_{1},\cdots,x_{L},y_{L}]^{T}
TABLE I: Driving decisions for the interactive vehicles

IV-B2 Feature Selection

Since the right of way for the merging vehicle and the lane-keeping vehicles are different, we define different features for them.

For the lane-keeping vehicle, the feature vectors related to the continuous driving decisions are as follows:

  • Yield: 𝐟yield=[fv,facc,fjerk,fdist,fg]T\mathbf{f}_{\text{yield}}{=}[f_{v},f_{\text{acc}},f_{\text{jerk}},f_{\text{dist}},f_{g}]^{T}. We exclude the feature fIDMf_{\text{IDM}} because once the lane-keeping driver decides to yield to the merging vehicle, it is very likely that he cares more about the relative positions to the merging vehicle instead of the heading space to the front vehicle when he plans his continuous trajectories. The goal position in fgf_{g} is set to be [xcurrent lane center,yt,merging vehicles0][x_{\text{current lane center}},y_{t,\text{merging vehicle}}{-}s_{0}], i.e., s0s_{0} behind the merging vehicle along longitudinal direction.

  • Pass: 𝐟pass=[fv,fIDM,facc,fjerk,fdist,fg]T\mathbf{f}_{\text{pass}}{=}[f_{v},f_{\text{IDM}},f_{\text{acc}},f_{\text{jerk}},f_{\text{dist}},f_{g}]^{T}. In this case, the goal position in fgf_{g} is set to be ahead of the merging vehicle along longitudinal direction, i.e., [xcurrent lane center,yt,merging vehicle+s0][x_{\text{current lane center}},y_{t,\text{merging vehicle}}{+}s_{0}]. Also, if the driver decides to pass, it is more probable that the heading space to the front vehicle will influence the distribution of his continuous trajectories.

For the merging vehicle, the feature vectors for the continuous driving models are:

  • Merge back: 𝐟back=[fv,facc,fjerk,fdist,fg]T\mathbf{f}_{\text{back}}{=}[f_{v},f_{\text{acc}},f_{\text{jerk}},f_{\text{dist}},f_{g}]^{T}. The goal position is set to be [xtarget lane center,yt,on-lane vehicles0][x_{\text{target lane center}},y_{t,\text{on-lane vehicle}}{-}s_{0}].

  • Merge front: 𝐟front=[fv,fIDM,facc,fjerk,fdist,fcourt,fg]T\mathbf{f}_{\text{front}}{=}[f_{v},f_{\text{IDM}},f_{\text{acc}},f_{\text{jerk}},f_{\text{dist}},f_{\text{court}},f_{g}]^{T}. Once the merging driver decides to merge in front of the lane-keeping vehicle, his heading space to the front lane-keeping vehicle is crucial to the distribution of his possible trajectories. Hence, we include feature fIDMf_{\text{IDM}}. Moreover, to respect the right of way of the lane-keeping vehicle, merging drivers might care about the extra cost they bring to the lane-keeping vehicle and prefer trajectories that induce less extra cost. To capture such effect, we add fcourtf_{\text{court}}. Given demonstrated trajectory group 𝝃=[ξmerging,ξsurroudings,ξlane-keeping]\boldsymbol{\xi}{=}[\xi_{\text{merging}},\xi_{\text{surroudings}},\xi_{\text{lane-keeping}}], fcourtf_{\text{court}} is computed via (13), i.e.,

    fcourt(ξmerging)=Clane-keeping(𝜽yield,𝝃)Clane-keepingdefault\displaystyle f_{\text{court}}(\xi_{\text{merging}})=C_{\text{lane-keeping}}(\boldsymbol{\theta}_{\text{yield}},\boldsymbol{\xi})-C_{\text{lane-keeping}}^{\text{default}} (19)

    The cost function Clane-keeping(𝜽yield)=𝜽yieldT𝐟yieldC_{\text{lane-keeping}}(\boldsymbol{\theta}_{\text{yield}}){=}\boldsymbol{\theta}_{\text{yield}}^{T}\mathbf{f}_{\text{yield}} is learned with the features selected above. Regarding to CHdefaultC_{H}^{\text{default}}, we assume that the lane-keeping vehicle is by default following IDM.

IV-C Implementation Details and Training Performance

We use Tensorflow to implement the hierarchical IRL algorithm. Figure 3 gives the training curves regarding to both the continuous and discrete driving decisions. Due to the hierarchical structure, we first learn all four continuous distribution models for both the merging vehicle and the lane-keeping vehicle under different discrete decisions. We randomly sample subsets of trajectories from the training set and perform multiple trains. As seen from Fig. 3, the parameters in each continuous model converge quite consistently with small variance.

With the converged parameter vectors 𝜽yield\boldsymbol{\theta}_{\text{yield}}, 𝜽pass\boldsymbol{\theta}_{\text{pass}}, 𝜽front\boldsymbol{\theta}_{\text{front}} and 𝜽back\boldsymbol{\theta}_{\text{back}}, the discrete feature vectors 𝐟d\mathbf{f}^{d}s are then calculated and thus the optimal parameter vectors 𝝍\boldsymbol{\psi}s are learned via discrete IRL. To efficiently sample continuous trajectories under different discrete decisions, we first obtain the most-likely trajectories by optimizing the learned cost functions under each decision and then randomize them with additive Gaussian noise. The training curves (blue) are also shown in Fig. 3.

Fig. 3: The training curves of both the lane-keeping vehicle and the merging vehicle under different discrete driving decisions

IV-D Test Results

Once 𝜽yield\boldsymbol{\theta}_{\text{yield}}, 𝜽pass\boldsymbol{\theta}_{\text{pass}}, 𝜽front\boldsymbol{\theta}_{\text{front}}, 𝜽back\boldsymbol{\theta}_{\text{back}} and 𝝍merging\boldsymbol{\psi}_{\text{merging}}, 𝝍lane-keeping\boldsymbol{\psi}_{\text{lane-keeping}} are acquired, we can obtain the conditional PDF defined in (2) via (3). With this PDF, probabilistic and interactive prediction of human drivers’ behavior can be obtained. The prediction horizon is 3s, i.e., 30 points with a sampling period of 0.1s.

IV-D1 Accuracy of the probabilistic prediction of discrete decisions

To measure the accuracy of the probabilistic prediction of discrete decisions, we extract N=2000N{=}2000 short trajectories in the test set with a horizon length of 1010 from the 54 long trajectories. For each short trajectory, starting from the same initial condition, we sample M=4M{=}4 trajectories with different motion patterns (discrete decisions) in the spatiotemporal domain, and one of them is set the same as the ground truth. Hence, the ground truth probabilities for all trajectories are either 0 or 1.

Metric: We adopt a fatality-aware Brier metric [23] to evaluate the prediction accuracy. The fatality-aware metric refines the Brier score [24] by formulating the prediction errors in three different aspects: ground-truth accuracy (𝒢\mathcal{G}) measuring the prediction error of the ground-truth motion pattern (Ξg\Xi_{g}), conservatism accuracy (𝒞\mathcal{C}) measuring the false alarm of aggressive motion patterns (Ξa\Xi_{a}), and non-defensiveness accuracy (𝒟\mathcal{D}) measuring the miss detection of dangerous motion patterns (Ξd\Xi_{d}). Let Pi,jP_{i,j} be the predicted probability of motion pattern jj for ii-th test example (j=1,2,,Mj{=}1,2,\cdots,M and i=1,2,,Ni{=}1,2,\cdots,N) , then the three scores and the overall score are given by

𝒢\displaystyle\mathcal{G} =\displaystyle= 1|Ξg|(i,j)Ξd(Pi,j1)2,\displaystyle\dfrac{1}{|\Xi_{g}|}\sum_{(i,j)\in\Xi_{d}}\left(P_{i,j}-1\right)^{2}, (20)
𝒞\displaystyle\mathcal{C} =\displaystyle= 1|Ξa|(i,j)ΞaWc(i,j)(Pi,j0)2,\displaystyle\dfrac{1}{|\Xi_{a}|}\sum_{(i,j)\in\Xi_{a}}W_{c}(i,j)\left(P_{i,j}-0\right)^{2}, (21)
𝒟\displaystyle\mathcal{D} =\displaystyle= 1|Ξd|(i,j)ΞdWd(i,j)(Pi,j0)2,\displaystyle\dfrac{1}{|\Xi_{d}|}\sum_{(i,j)\in\Xi_{d}}W_{d}(i,j)\left(P_{i,j}-0\right)^{2}, (22)
c\displaystyle\mathcal{B}_{c} =\displaystyle= 𝒢+𝒞+𝒟.\displaystyle\mathcal{G}+\mathcal{C}+\mathcal{D}. (23)

Wc(i,j)W_{c}(i,j) and Wd(i,j)W_{d}(i,j) are, respectively, the weights penalizing the conservatism and non-defensiveness of the motion pattern (i,j)(i,j). For more details, one can refer to [23].

Comparison: We compare the prediction among three different approaches: the proposed hierarchical IRL method, a neural-network (NN) based method [25] and a hidden markov models (HMM) based method [26].

Results: The scores for all three methods are shown in Table II. We can see that the proposed method can yield better overall prediction performance than HMM and NN based methods, particularly in terms of conservatism and non-defensiveness. This means that the prediction generated by the proposed method has similar criticality as the ground truth in the interaction process.

HMM NN Hierarchical IRL
𝒢\mathcal{G} 0.0701 0.0563 0.1117
𝒞\mathcal{C} 0.0476 0.0493 0.0178
𝒟\mathcal{D} 0.1356 0.1303 0.0698
c=𝒢+𝒞+𝒟\mathcal{B}_{c}=\mathcal{G}+\mathcal{C}+\mathcal{D} 0.2551 0.2361 0.2053
TABLE II: Scores of different probabilistic prediction approaches

IV-D2 Accuracy of the probabilistic prediction of continuous trajectories

In this test, we generate the most probable trajectories under different discrete driving decisions by solving a finite horizon Model Predictive Control (MPC) problem using the above learned continuous cost functions. We show three illustrative examples in Fig. 4. The red dotted lines and blue solid lines represent, respectively, the predicted most-likely trajectories and the ground truth trajectories. The thick black dash-dot lines are trajectories of other vehicles. We can see that the predicted trajectories are very close to the ground truth ones.

Fig. 4: Three illustrative examples of the predicted most probable trajectories (red dotted line) compared with the ground truth trajectories (blue solid line). Thick black dash-dot lines represent the trajectories of other vehicles except for the predicted one.

Metric: We also adopt Mean Euclidean Distance (MED) [27] to quantitatively evaluate the accuracy of the prediction for continuous trajectories. Given the ground truth trajectory ξground=[xg,1,yg,1,,xg,L,yg,L]T\xi_{\text{ground}}{=}[x_{g,1},y_{g,1},{\cdots},x_{g,L},y_{g,L}]^{T} and predicted trajectory ξprediction=[xp,1,yp,1,,xp,L,yp,L]T\xi_{\text{prediction}}{=}[x_{p,1},y_{p,1},{\cdots},x_{p,L},y_{p,L}]^{T} of same length LL and same sampling time T\triangle T, the trajectory similarity is calculated as follows:

𝒮MED=1Li=1L[xp,i,yp,i]T[xg,i,yg,i]T2\mathcal{S}_{\text{MED}}=\dfrac{1}{L}\sum_{i{=}1}^{L}\|\left[x_{p,i},y_{p,i}\right]^{T}-\left[x_{g,i},y_{g,i}\right]^{T}\|_{2} (24)

Results: We test on 2020 long trajectories in the test set, and results are summarized in Table III: the proposed hierarchical IRL method can achieve trajectory prediction with a mean MED of 0.6172m with standard deviation of 0.2473m.

Mean (m) Max (m) Min (m) Std (m)
MED 0.6172 1.0146 0.3644 0.2473
TABLE III: Trajectory similarities in terms of MED

V Conclusion

In this work, we have proposed a probabilistic and interactive prediction approach via hierarchical inverse reinforcement learning. Instead of directly making predictions based on historical information, we formulate the prediction problem from a game–theoretic view: the distribution of future trajectories of the predicted vehicle strongly depends on the future plans of the host vehicle. To address such interactive behavior, we design a hierarchical inverse reinforcement learning method to capture the hierarchical trajectory–generation process of human drivers. Influences of both discrete and continuous driving decisions are explicitly modelled. We also have applied the proposed method on a ramp–merging driving scenario. The quantitative results verified the effectiveness of the proposed approach in terms of accurate prediction to both discrete decisions and continuous trajectories.

Acknowledgment

We thank Yeping Hu and Jiachen Li for their assistance with the implementation of the neural network (NN) and hidden Markov model (HMM) based prediction approaches in Section IV-D.

References

  • [1] Y. Kuwata, J. Teo, G. Fiore, S. Karaman, E. Frazzoli, and J. P. How, “Real-time motion planning with applications to autonomous urban driving,” IEEE Transactions on Control Systems Technology, vol. 17, no. 5, pp. 1105–1118, 2009.
  • [2] Z. Liang, G. Zheng, and J. Li, “Automatic parking path optimization based on bezier curve fitting,” in Automation and Logistics (ICAL), 2012 IEEE International Conference on. IEEE, 2012, pp. 583–587.
  • [3] J. Ziegmann, J. Shi, T. Schnörer, and C. Endisch, “Analysis of individual driver velocity prediction using data-driven driver models with environmental features,” in Intelligent Vehicles Symposium (IV), 2017 IEEE. IEEE, 2017, pp. 517–522.
  • [4] R. Graf, H. Deusch, F. Seeliger, M. Fritzsche, and K. Dietmayer, “A learning concept for behavior prediction at intersections,” in 2014 IEEE Intelligent Vehicles Symposium Proceedings, June 2014, pp. 939–945.
  • [5] S. Lefèvre, C. Laugier, and J. Ibañez Guzmán, “Intention-aware risk estimation for general traffic situations, and application to intersection safety,” Inria Research Report, no. 8379.
  • [6] M. Schreier, V. Willert, and J. Adamy, “An Integrated Approach to Maneuver-Based Trajectory Prediction and Criticality Assessment in Arbitrary Road Environments,” IEEE Transactions on Intelligent Transportation Systems, vol. 17, no. 10, pp. 2751–2766, Oct. 2016.
  • [7] X. Geng, H. Liang, B. Yu, P. Zhao, L. He, and R. Huang, “A Scenario-Adaptive Driving Behavior Prediction Approach to Urban Autonomous Driving,” Applied Sciences, vol. 7, no. 4, p. 426, Apr. 2017.
  • [8] D. J. Phillips, T. A. Wheeler, and M. J. Kochenderfer, “Generalizable intention prediction of human drivers at intersections,” in Intelligent Vehicles Symposium (IV), 2017 IEEE. IEEE, 2017, pp. 1665–1670.
  • [9] Y. Hu, W. Zhan, and M. Tomizuka, “Probabilistic prediction of vehicle semantic intention and motion,” arXiv preprint arXiv:1804.03629, 2018.
  • [10] C. Dong, J. M. Dolan, and B. Litkouhi, “Intention estimation for ramp merging control in autonomous driving,” in 2017 IEEE Intelligent Vehicles Symposium (IV), June 2017, pp. 1584–1589.
  • [11] D. Sadigh, S. Sastry, S. A. Seshia, and A. D. Dragan, “Planning for autonomous cars that leverage effects on human actions.” in Robotics: Science and Systems, 2016.
  • [12] R. E. Kalman, “When is a linear control system optimal?” Journal of Basic Engineering, vol. 86, no. 1, pp. 51–60, 1964.
  • [13] A. Y. Ng, S. J. Russell, et al., “Algorithms for inverse reinforcement learning.” in Icml, 2000, pp. 663–670.
  • [14] B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in AAAI, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438.
  • [15] S. Levine and V. Koltun, “Continuous inverse optimal control with locally optimal examples,” 2012.
  • [16] S. Liu, M. Araujo, E. Brunskill, R. Rossetti, J. Barros, and R. Krishnan, “Understanding sequential decisions via inverse reinforcement learning,” in Mobile Data Management (MDM), 2013 IEEE 14th International Conference on, vol. 1. IEEE, 2013, pp. 177–186.
  • [17] M. Shimosaka, K. Nishi, J. Sato, and H. Kataoka, “Predicting driving behavior using inverse reinforcement learning with multiple reward functions towards environmental diversity,” in Intelligent Vehicles Symposium (IV), 2015 IEEE. IEEE, 2015, pp. 567–572.
  • [18] H. Kretzschmar, M. Spies, C. Sprunk, and W. Burgard, “Socially compliant mobile robot navigation via inverse reinforcement learning,” The International Journal of Robotics Research, vol. 35, no. 11, pp. 1289–1307, 2016.
  • [19] D. Lenz, F. Diehl, M. T. Le, and A. Knoll, “Deep neural networks for markovian interactive scene prediction in highway scenarios,” in Intelligent Vehicles Symposium (IV), 2017 IEEE. IEEE, 2017, pp. 685–692.
  • [20] A. Kesting, M. Treiber, and D. Helbing, “Enhanced intelligent driver model to access the impact of driving strategies on traffic capacity,” Philosophical Transactions of the Royal Society of London A: Mathematical, Physical and Engineering Sciences, vol. 368, no. 1928, pp. 4585–4605, 2010.
  • [21] L. Sun, W. Zhan, M. Tomizuka, and A. Dragan, “Courteous autonomous cars,” in 2018 IEEE/RSJ International Conference on Intelligent Robots (IROS), to appear.
  • [22] V. Alexiadis, J. Colyar, J. Halkias, R. Hranac, and G. McHale, “The Next Generation Simulation Program,” Institute of Transportation Engineers. ITE Journal; Washington, vol. 74, no. 8, pp. 22–26, Aug. 2004.
  • [23] W. Zhan, L. Sun, Y. Hu, J. Li, and M. Tomizuka, “Towards a fatality-aware benchmark of probabilistic reaction prediction in highly interactive driving scenarios,” in the 21st IEEE International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2018.
  • [24] G. W. Brier, “Verification of forecasts expressed in terms of probability,” Monthey Weather Review, vol. 78, no. 1, pp. 1–3, 1950.
  • [25] C. M. Bishop, “Mixture density networks,” 1994.
  • [26] L. R. Rabiner, “A tutorial on hidden markov models and selected applications in speech recognition,” Proceedings of the IEEE, vol. 77, no. 2, pp. 257–286, 1989.
  • [27] J. Quehl, H. Hu, O. S. Tas, E. Rehder, and M. Lauer, “How good is my prediction? finding a similarity measure for trajectory prediction evaluation.” in 2017 IEEE 18th International Conference on Intelligent Transportation Systems (ITSC), 2017, pp. 120–125.