arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:1809.03154v1 [cs.GT] 10 Sep 2018

Learning Time Dependent Choice

Zachary Chase Thanks: California Institute of Technology, zchase@caltech.edu    Siddharth Prasad Thanks: California Institute of Technology, sprasad@caltech.edu
Abstract

We explore questions dealing with the learnability of models of choice over time. We present a large class of preference models defined by a structural criterion for which we are able to obtain an exponential improvement over previously known learning bounds for more general preference models. This in particular implies that the three most important discounted utility models of intertemporal choice – exponential, hyperbolic, and quasi-hyperbolic discounting – are learnable in the PAC setting with VC dimension that grows logarithmically in the number of time periods. We also examine these models in the framework of active learning. We find that the commonly studied stream-based setting is in general difficult to analyze for preference models, but we provide a redeeming situation in which the learner can indeed improve upon the guarantees provided by PAC learning. In contrast to the stream-based setting, we show that if the learner is given full power over the data he learns from – in the form of learning via membership queries – even very naive algorithms significantly outperform the guarantees provided by higher level active learning algorithms.

1 Introduction

We study the learnability of economic models of choice over time. Our setting is that of an analyst who first observes an agent’s choices between plans that specify payoffs over time, and then attempts to learn the preference parameters guiding the choices. While such parameters are stylized – in reality subjects are not likely to perform standardized computations according to private parameters before making decisions – experiments have shown that they often provide accurate descriptions of how an agent behaves. By observing enough choice data, one can hope to learn the economic parameters that most closely describe the agent’s preferences. Thus, learning theory provides an especially meaningful lens with which to view the theory of choice – it allows us to answer questions regarding the volume of data required to faithfully predict future decisions made by an observed agent. The overarching goal of this paper is to identify structural criteria that yield strong learnability results for preferences over time under different restrictions placed on the learner/analyst. The criteria we present captures a large class of preference models that give the agent significant freedom in weighting decisions against time delays. In particular, it encompasses the most popular models of time dependent choice used by economists.

The main economic application of our results is in understanding the learnability of models of intertemporal choice. Intertemporal choice is what governs an agent’s decisions over several time periods. The most important models of intertemporal choice are discounted utility models, in which agents evaluate plans by discounting actions as they are delayed – in analogy to how markets value the loss or gain of money over time. The first axiomatic treatment of discounting was by Koopmans in 1960 [18], in which he demonstrates that simple postulates for preferences over an infinite time horizon yield “impatience.” The three most commonly studied discounting models are exponential, hyperbolic, and quasi-hyperbolic, and all have been studied by both economists and computer scientists (though less so by the latter) as well as researchers from various other fields. The importance of discounted utility in economics cannot be overstated – it is the canonical framework used by economists to study choice over time.

Problems of learning economic parameters have received recent attention from computer scientists; see, e.g., [1, 2, 3, 16, 22]. Inspired by a general theme of demanding computational robustness from economic models (Echenique, Golovin, and Wierman provide a nice discussion of this topic in [11]), the tools of learning theory provide relevant and exciting perspectives from which to view economic models that have been around for several decades. In contrast to the usual goal of truthfully extracting the agent’s parameters adopted by classical mechanism design, the learning problem aims to efficiently extract a truthful agent’s parameters in the restricted message space of binary classification. Our paper contributes to the line of work that specifically studies models of choice using the perspectives of learning theory. This confluence of decision theory and learning theory was initiated by Basu and Echenique [2], who consider the learning problem for models of choice under uncertainty. Our investigation in this paper is motivated by models of how agents make choices over time. We provide learnability results that are fine tuned to structural requirements on such models.

We now summarize our main contributions at a high level. Section 3 contains a more detailed exposition of our results.

Summary of results and techniques

Our situation is one of an analyst trying to learn the parameters governing an agent’s preferences over time. The two main learning themes we consider are (1) when the analyst has no control over the data he sees and (2) when the analyst has some control over the data he sees. The first theme is aptly captured by probably approximately correct (PAC) learning. To analyze the second theme, we investigate two models of active learning: stream-based selective sampling and membership queries.

In the first part, we study the PAC model, where the analyst is presented with pairs of alternatives and a label for each pair indicating the agent’s preference between the alternatives. The data points are drawn according to some unknown distribution, and the analyst has no control over the data he is presented with. Our main result here is a structural criterion on preference models that allows for a drastic improvement over the PAC learning complexity bounds achieved in [2]. We stipulate that the agent weights time-delayed payoffs according to polynomials, which allows for considerable freedom in how payoffs are weighted. Under this requirement, we show that such classes of preference models admit an exponential improvement in sample complexity bounds over the more general preference models considered in [2]. This is achieved via a computation of the VC dimension (which quantifies the complexity of PAC learning). A simple application of our result shows that each of the discounted utility models are learnable, with sample complexity that grows logarithmically in the number of time periods TT over which decisions are being made. The computation of the VC dimension is due to a natural connection between pairs of choices and the signs of polynomials that arise from the choices.

In the second part, we consider active learning models, where the analyst is given a certain amount of control over the data that he uses to learn. The two active learning models we study are stream-based selective sampling and learning via membership queries. In the former, the analyst is given some control over what data he learns from: as in the PAC setting he is presented with points drawn from an unknown distribution, but now the analyst chooses whether or not to see the label representing the agent’s choice for each point. In the latter, the analyst has complete control over the data he learns from: the analyst can at any time request the label for any point. The former model seems to have been commonly adopted in order to study the very general problem of concept learning, when there is no extra information about the structure of the concepts. We find that the disagreement methods used to study the stream-based setting are in general difficult to analyze in the context of preference models – requiring quantitative information about the underlying distribution from which points are drawn. However, we provide a redeeming situation (by examining a particular distribution) where we obtain an improvement over the PAC guarantees. Membership queries, on the other hand, allow us to heavily exploit the structure of the preference models we consider. We present a naive membership query-based algorithm that significantly outperforms the guarantees provided in the stream-based setting. Learning via membership queries, we conclude, seems to be the appropriate model to actively learn economic parameters. It allows the analyst to make use of the preference relations’ structure, and also precisely captures the situation in which the analyst and agent are participating in a real time experiment.

Related work

Discounted utility models of intertemporal choice have been studied extensively not only by economists, but also by researchers from various other fields. We first briefly survey some of the relevant work pertaining to the exponential, quasi-hyperbolic, and hyperbolic discounting models and then survey existing work in the more general topic of learning economic parameters.

In the exponential discounting model, the agent evaluates his utilities based on a discount factor δ(0,1)\delta\in(0,1), where a delay of tt time periods incurs an exponential discount in utility by δt\delta^{t}. Climate change policies are traditionally evaluated according to an exponential discounting model – for example, the Stern review on the economics of climate change deals with issues of how to choose an appropriate discount rate in evaluating such policies [21]. Chambers and Echenique [8] present results related to the problem of aggregating discount rates proposed by a group of experts facing disagreement. While it is the most commonly used discounting model due to its simplicity, the exponential discounting model has been criticized due to its inability to match empirical data recording actual human behavior. Quasi-hyperbolic and hyperbolic discounting aim to mend such issues. The quasi-hyperbolic discounting model is parametrized by β,δ(0,1)\beta,\delta\in(0,1), where a delay of tt time periods incurs a discount in utility by βδt\beta\delta^{t}, and was first introduced by Phelps and Pollack [20] to study preferences over generations. They proposed that the constant β\beta discount factor represents how much a given generation tt is affected by the utilities of other people relative to their own – and remark that β=1\beta=1 represents “perfect altruism,” while β<1\beta<1 represents “imperfect altruism.” Kleinberg and Oren [17] study agents with quasi-hyperbolic discounting and propose a graph-theoretic model to investigate phenomena such as procrastination and abandonment of long-range tasks. Hyperbolic discounting aims to capture the notion that people are more impatient in making short term decisions (today vs. tomorrow) than long term decisions (365 days from today vs. 366 days from today)11 1 In particular note that exponential discounting does not capture this issue, i.e. it is dynamically consistent, in that preferences do not change according to shifts in time., and is modeled via a discount of (1+tα)1(1+t\alpha)^{-1} at time tt. Researchers in fields such as psychology and neuroscience [4, 15] have adopted the hyperbolic discounting model to study, for example, issues of self control and anticipation in humans and animals, and have compared the predictions by the different discounted utility models to neurobiological data obtained via MRI scans. Chabris et al. [7] give an exposition of the discounted utility models of intertemporal choice and survey sociological research that examines empirical data pertaining to how discount rates are affected by factors like age, drug use, gambling, etc.

The study of economic models has witnessed a recent influx of work from computer scientists dealing with questions of robustness under various notions of complexity (learning complexity, computational complexity, communication complexity, etc.). Kalai [16] in 2001 studied the learnability of choice functions, where the observed choices are in the form of a given set of alternatives along with the most preferred alternative from the set. Beigman and Vohra [3], Zadimoghaddam and Roth [22], and Balcan et al. [1] investigate the problem of learning utility functions in the context of an expected utility maximizing agent in a demand environment. Most recently (and most related to our work), Basu and Echenique [2] study the learnability of preference models of choice under uncertainty, in which an agent is uncertain about states of a lottery and is made to choose between acts that encode utilities over each state. Here, the different models of choice under uncertainty arise from different ways of representing the subjective probability held by an agent. They are also the first to study learnability in the decision-theoretic setting where choice is modeled by preference relations rather than by expected utility maximizing behavior in a demand setting. However, it does not appear that the learnability of models of intertemporal choice has been previously studied.

2 Model and Preliminaries

We now formally set up the discounted utility models of intertemporal choice and state the standard definitions from learning theory in the context of preference relations. Much of the following material regarding learning and preference relations is taken from [2] since we require a similar list of definitions and setup. First, we sketch our high level model.

Let XX be a Euclidean space equipped with a Borel σ\sigma-algebra. A preference relation on XX is a binary relation X×X\succsim\subseteq X\times X such that \succsim is measurable with respect to the product σ\sigma-algebra on X×XX\times X. A model 𝒫\mathcal{P} of preference relations is a collection of preference relations.

An agent makes choices from pairs of alternatives (xi,yi)i=1n(x^{i},y^{i})_{i=1}^{n} that are drawn according to some unknown distribution on X×XX\times X. The choices are presented as labels (ai)i=1n(a_{i})_{i=1}^{n} where ai=1a_{i}=1 if the agent chooses xix^{i} and ai=0a_{i}=0 if the agent chooses yiy^{i}. A dataset is any finite sequence of pairs of plans and their labels (((x1,y1),a1),,((xn,yn),an))(((x^{1},y^{1}),a_{1}),\ldots,((x^{n},y^{n}),a_{n})). An analyst observes a dataset, and attempts to guess the preference relation governing the agent’s choices. A learning rule is any map σ\sigma from datasets to preference relations. The output of the learning rule is the analyst’s hypothesis as to what the agent’s true preference relation is, having seen some finite dataset.

2.1 Learnability

The two notions of learnability we consider are the PAC model and the active model. We now state the standard definitions of PAC and active learning in the context of preference relations. Most of the following definitions for the PAC setting are taken from [2] since the setup involving preference relations is identical. These definitions of course apply to the more general setting of concept learning (for example, see [6]).

A collection 𝒫\mathcal{P} of preference relations is (PAC) learnable if there is a learning rule σ\sigma such that for every 0<ε,δ<10<\varepsilon,\delta<1, there is s(ε,δ)s(\varepsilon,\delta)\in\mathbb{N} such that for every ns(ε,δ)n\geq s(\varepsilon,\delta), 𝒫\succsim\in\mathcal{P}, and μΔ(X×X)\mu\in\Delta(X\times X),

μn({((x1,y1),,(xn,yn)):μ()>ε})<δ,\mu^{n}(\{((x^{1},y^{1}),\ldots,(x^{n},y^{n})):\mu(\succsim^{*}\triangle\succsim)>\varepsilon\})<\delta,

where

=σ({((x1,y1),Ix1y1),,((xn,yn),Ixnyn)})\succsim^{*}=\sigma(\{((x^{1},y^{1}),I_{x^{1}\succsim y^{1}}),\ldots,((x^{n},y^{n}),I_{x^{n}\succsim y^{n}})\})

is the hypothesis preference relation produced by the learning rule22 2 μn\mu^{n} denotes the product measure induced by μ\mu on (X×X)n(X\times X)^{n}.. The quantity s(ε,δ)s(\varepsilon,\delta) is called the sample complexity of the learning rule σ\sigma.

The complexity of learning is commonly quantified by the Vapnik-Chervonenkis (VC) dimension, which we now define. A set of points {(x1,y1),,(xn,yn)}\{(x^{1},y^{1}),\ldots,(x^{n},y^{n})\} from X×XX\times X is shattered by a model of preferences 𝒫\mathcal{P} if for every vector of labels (a1,,an){0,1}n(a_{1},\ldots,a_{n})\in\{0,1\}^{n}, there is a preference relation 𝒫\succsim\in\mathcal{P} that realizes the labelling, i.e. for i=1,,ni=1,\ldots,n we have that xiyix^{i}\succsim y^{i} if and only if ai=1a_{i}=1. In this case, 𝒫\mathcal{P} is said to rationalize the dataset {((x1,y1),a1),,((xn,yn),an)}\{((x^{1},y^{1}),a_{1}),\ldots,((x^{n},y^{n}),a_{n})\}. The VC dimension of 𝒫\mathcal{P}, denoted by VC(𝒫)VC(\mathcal{P}), is the largest integer nn such that there exist nn points that are shattered by 𝒫\mathcal{P}.

Blumer et al. [6] in 1989 proved that learnability is equivalent to having a finite VC dimension33 3 This result requires 𝒫\mathcal{P} to satisfy a certain measurability requirement. We note in Section 4 that the models of choice we consider all satisfy said requirement..

Theorem 2.1.

A model of preferences 𝒫\mathcal{P} is learnable if and only if VC(𝒫)<VC(\mathcal{P})<\infty.

The VC dimension (denoted by dd for the remainder of this subsection) also plays a role in the sample complexity of learning a model of preferences. In the same paper, Blumer et al. [6] show that any algorithm that outputs a hypothesis consistent with the data seen is a valid learning rule requiring sample complexity

s(ε,δ)=O(1ε(dlog1ε+log1δ)).s(\varepsilon,\delta)=O\left(\frac{1}{\varepsilon}\left(d\log\frac{1}{\varepsilon}+\log\frac{1}{\delta}\right)\right).

In 2016, Hanneke [14] showed that these bounds (after a small improvement) are tight: the optimal sample complexity of PAC learning is

s(ε,δ)=Θ(1ε(d+log1δ)).s(\varepsilon,\delta)=\Theta\left(\frac{1}{\varepsilon}\left(d+\log\frac{1}{\delta}\right)\right).

The other learning model we consider is the active learning framework, where the analyst has some control over the data from which he learns. In stream-based selective sampling, points drawn according to an unknown distribution are presented to the analyst as before, but without the labels. The analyst can choose whether or not to query the label of a given point, and the complexity of the learning rule is measured by label complexity, i.e. the number of labels requested by the analyst. Disagreement based active learning refers to the paradigm in which the learner only requests labels on points that significantly reduce the hypothesis space. The disagreement of a preference model with respect to the underlying distribution is quantified through the disagreement coefficient θ\theta, which is defined in Section 5. A finite disagreement coefficient implies (for the underlying distribution) an exponential improvement in label complexity over the sample complexity of PAC learning. For example, the CAL algorithm [9, 10, 13], a simple disagreement based learning algorithm, yields a label complexity of

CAL(ε,δ)=O(θlog1ε(dlogθ+loglog(1/ε)δ)).\ell_{CAL}(\varepsilon,\delta)=O\left(\theta\log\frac{1}{\varepsilon}\left(d\log\theta+\log\frac{\log(1/\varepsilon)}{\delta}\right)\right).

In the membership queries model, the analyst is allowed to request the label for any point at any time. There appears to be a dearth of literature/results pertaining to the complexity of membership query algorithms for learning when the hypothesis space is infinite. One explanation for this is that improvements to the “passive” disagreement based methods used in the stream-based setting would need specific information about the problem domain: disagreement based methods are designed to work on a very general class of concept learning problems without assuming anything about the learning space. In our case, we have specific details about how the preference relations take shape. Thus, the membership query model turns out to be an interesting and useful perspective to use in the study of learning preference models.

For a more detailed survey of active learning, see [10].

2.2 Discounted utility

We now present the definitions for the discounted utility models of intertemporal choice. An agent chooses between plans or vectors in X=TX=\mathbb{R}^{T} that encode payoffs over TT time periods. A preference relation over plans is a binary relation T×T\succsim\subseteq\mathbb{R}^{T}\times\mathbb{R}^{T}.

The most important model of intertemporal choice is the discounted utility model, in which the agent’s payoffs xtx_{t} for having chosen a plan xTx\in\mathbb{R}^{T} are reduced, or discounted, as tt increases from 11 to TT. In its most general form, we can characterize the preference relations that follow time discounting as follows:

Definition 2.1 (Discounted utility model).

The class of preference relations 𝒫𝒟\mathcal{P}_{\mathcal{D}} that satisfy the discounted utility model are those \succsim such that there exists a decreasing map D:{1,,T}(0,1)D:\{1,\ldots,T\}\to(0,1) where

xy if and only if t=1TD(t)xtt=1TD(t)yt.x\succsim y\,\,\text{ if and only if }\,\,\sum_{t=1}^{T}D(t)x_{t}\geq\sum_{t=1}^{T}D(t)y_{t}.

We use the following notation for the preference models arising from the three most commonly studied discounting functions DD:

  • 𝒫𝒟\mathcal{P}_{\mathcal{D}} denotes the set of preferences that satisfy the discounted utility model.

  • 𝒫𝒟\mathcal{P}_{\mathcal{ED}} denotes the set of preferences that satisfy the discounted utility model with exponential discounting: D(t)=δtD(t)=\delta^{t} for δ(0,1)\delta\in(0,1).

  • 𝒫𝒟\mathcal{P}_{\mathcal{HD}} denotes the set of preferences that satisfy the discounted utility model with hyperbolic discounting: D(t)=11+tαD(t)=\frac{1}{1+t\alpha} for α>0\alpha>0.

  • 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} denotes the set of preferences that satisfy the discounted utility model with quasi-hyperbolic discounting: D(t)=1D(t)=1 if t=1t=1, D(t)=βδt1D(t)=\beta\cdot\delta^{t-1} if t>1t>1 for β,δ(0,1)\beta,\delta\in(0,1).

For a more thorough exposition on the various discounted utility models of intertemporal choice, see [7].

3 Main Results

In this section we provide a formal discussion and interpretation of our results, which is split into two themes: the first dealing with an analyst who has no control over the learning data, the second dealing with an analyst who has some control over the learning data.

A powerless analyst

The first part of our paper investigates the situation of an analyst trying to learn the preference relation by which an agent makes choices, but has no control over what choices he gets to observe, and is agnostic to the process by which they are drawn. We thus adopt the PAC learning model.

The agent chooses between plans that encode payoffs over TT periods of time and evaluates the total payoff of a plan vector xTx\in\mathbb{R}^{T} according to private weights w1,,wTw_{1},\ldots,w_{T} that he multiplicatively applies to each state: payoff(x)=t=1Twtxt.\text{payoff}(x)=\sum_{t=1}^{T}w_{t}x_{t}. This defines a model of preference relations, which we denote by 𝒫𝒲\mathcal{P}_{\mathcal{W}}, where for any 𝒫𝒲\succsim\in\mathcal{P}_{\mathcal{W}}, there exists a vector of weights w=(w1,,wT)Tw=(w_{1},\ldots,w_{T})\in\mathbb{R}^{T} such that

xy if and only if w.xw.y.x\succsim y\,\,\text{ if and only if }w.x\geq w.y.

In [2], it is shown that T1VC(𝒫𝒲)T+1T-1\leq VC(\mathcal{P}_{\mathcal{W}})\leq T+1. In the context of choice over time, however, this model is extremely general and does not capture any of the intuitive notions of how an agent values payoffs when they are delayed44 4 In [2] the complete control over weights is used to model choice under uncertainty, which calls for such generality since the agent’s beliefs/weights are given by an element of the probability simplex on T\mathbb{R}^{T}.. For example, the discounted utility models of intertemporal choice require the weights to be of a particular functional form. Moreover, when there is no structure to the discount function we cannot improve the bounds on 𝒫𝒲\mathcal{P}_{\mathcal{W}}:

Proposition 3.1.

T1VC(𝒫𝒟)T+1T-1\leq VC(\mathcal{P}_{\mathcal{D}})\leq T+1.

This leads us to the motivating question of the first part of the paper: what structural conditions can we impose on the weights w1,,wTw_{1},\ldots,w_{T} such that this bound can be improved?

We investigate the situation where the agent computes his weights by evaluating polynomials at a private parameter δ\delta. Specifically, let Q1,,QTQ_{1},\ldots,Q_{T} be polynomials of degree at most dd, and suppose the agent evaluates total payoff of a plan vector xTx\in\mathbb{R}^{T} by payoff(x)=t=1TQt(δ)xt\text{payoff}(x)=\sum_{t=1}^{T}Q_{t}(\delta)x_{t}. Consequently, let 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} be the model of preference relations parametrized by δ\delta such that

xy if and only if t=1TQt(δ)xtt=1TQt(δ)yt.x\succsim y\,\,\text{ if and only if }\sum_{t=1}^{T}Q_{t}(\delta)x_{t}\geq\sum_{t=1}^{T}Q_{t}(\delta)y_{t}.

This class of preference models allows us to approximate preference relations where the weights are given by any real valued functions – we choose Q1,,QTQ_{1},\ldots,Q_{T} to be the appropriate Taylor polynomials. Moreover, existing models of intertemporal choice fit this characterization – for example 𝒫𝒟\mathcal{P}_{\mathcal{ED}} and 𝒫𝒟\mathcal{P}_{\mathcal{HD}}.

We additionally consider a slightly larger class of preference models where the agent has a private parameter β\beta (in addition to δ\delta) that in evaluating total payoff of a plan vector xTx\in\mathbb{R}^{T} allows the agent to modify the constant term t=1TQt(0)xt\sum_{t=1}^{T}Q_{t}(0)x_{t} of the polynomial t=1TQt(δ)xt\sum_{t=1}^{T}Q_{t}(\delta)x_{t}. This model aims to more generally capture the effects of the β\beta parameter in quasi-hyperbolic discounting. For polynomials Q1,,QTQ_{1},\ldots,Q_{T} of degree at most dd, let 𝒫𝒫𝒲\mathcal{P}_{\mathcal{BPW}} be the model of preference relations parametrized by β\beta and δ\delta such that xyx\succsim y if and only if

(1β1)t=1TQt(0)xt+t=1TQt(δ)xt(1β1)t=1TQt(0)yt+t=1TQt(δ)yt.\left(\frac{1}{\beta}-1\right)\sum_{t=1}^{T}Q_{t}(0)x_{t}+\sum_{t=1}^{T}Q_{t}(\delta)x_{t}\geq\left(\frac{1}{\beta}-1\right)\sum_{t=1}^{T}Q_{t}(0)y_{t}+\sum_{t=1}^{T}Q_{t}(\delta)y_{t}.

Our main results show that with this additional structure on the preference model, we can achieve an exponential improvement in the bounds for the VC dimension of 𝒫𝒲\mathcal{P}_{\mathcal{W}} obtained in [2]66 6 It is important to note that the classes 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} and 𝒫𝒫𝒲\mathcal{P}_{\mathcal{BPW}} are defined for a given Q1,,QTQ_{1},\ldots,Q_{T}. That is, the analyst knows Q1,,QTQ_{1},\ldots,Q_{T}, and is trying to learn the parameters β\beta and δ\delta. If the Q1,,QTQ_{1},\ldots,Q_{T} are private information only available to the agent, we are in no better shape than in the case of 𝒫𝒲\mathcal{P}_{\mathcal{W}}..

Theorem 3.2.

For every ε>0\varepsilon>0, there exists a dεd_{\varepsilon} such that for every ddεd\geq d_{\varepsilon} we have VC(𝒫𝒫𝒲),VC(𝒫𝒫𝒲)(1+ε)logdVC(\mathcal{P}_{\mathcal{PW}}),VC(\mathcal{P}_{\mathcal{BPW}})\leq(1+\varepsilon)\log d for any TT and any TT polynomials Q1,,QTQ_{1},\ldots,Q_{T} of degree at most dd.

Note that when Q1,,QTQ_{1},\ldots,Q_{T} have degree at most polynomial in TT, we obtain an exponential improvement over the linear growth of VC(𝒫𝒲)VC(\mathcal{P}_{\mathcal{W}}). We show that in this case, we get a tight (asymptotic) bound of logT\log T:

Theorem 3.3.

Let Q1,,QTQ_{1},\ldots,Q_{T} be polynomials in δ\delta of degree at most T1T-1 that span the space of polynomials in δ\delta of degree at most T1T-1. Then VC(𝒫𝒫𝒲),VC(𝒫𝒫𝒲)log(T1)VC(\mathcal{P}_{\mathcal{PW}}),VC(\mathcal{P}_{\mathcal{BPW}})\geq\log(T-1).

An interesting feature of Theorems 3.2 and 3.3 is that for fixed Q1,,QTQ_{1},\ldots,Q_{T} with degrees at most T1T-1, VC(𝒫𝒫𝒲)VC(\mathcal{P}_{\mathcal{PW}}) and VC(𝒫𝒫𝒲)VC(\mathcal{P}_{\mathcal{BPW}}) satisfy the same asymptotic bounds, so giving the agent an extra parameter that allows control over the constant term of the polynomial t=1TQt(δ)xt\sum_{t=1}^{T}Q_{t}(\delta)x_{t} does not introduce a significant amount of richness to the model.

Applying Theorems 3.2 and 3.3 to the discounted utility models, we have:

Corollary 3.4.

VC(𝒫𝒟),VC(𝒫𝒟),VC(𝒫𝒬𝒟)log(T1)VC(\mathcal{P}_{\mathcal{ED}}),VC(\mathcal{P}_{\mathcal{HD}}),VC(\mathcal{P}_{\mathcal{QHD}})\sim\log(T-1)

Thus, 𝒫𝒟\mathcal{P}_{\mathcal{D}}, 𝒫𝒟\mathcal{P}_{\mathcal{ED}}, 𝒫𝒟\mathcal{P}_{\mathcal{HD}}, and 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} are all learnable. 𝒫𝒟\mathcal{P}_{\mathcal{D}} requires a minimum sample size that grows linearly with TT, while 𝒫𝒟\mathcal{P}_{\mathcal{ED}}, 𝒫𝒟\mathcal{P}_{\mathcal{HD}}, and 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} require a minimum sample size that grows logarithmically in TT.

The main technique in proving Theorems 3.2 and 3.3 is interpreting the shattering criteria as a statement about the sign combinations achieved by a collection of polynomials. The upper bound on the VC dimension follows from an upper bound on the number of sign combinations a collection of polynomials can achieve. In demonstrating the lower bound on the VC dimension, we construct a set of points that is shattered by finding polynomials achieving all possible sign combinations – chosen according to a Hamiltonian path in the log(T1)\log(T-1)-dimensional hypercube.

A powerful analyst

Our other results concern the active learning framework, which broadly deals with situations in which the analyst has some control over the choices he observes and learns from. The two models we consider are stream-based selective sampling and learning via membership queries. A large body of active learning research is devoted to the stream-based model, specifically focusing on disagreement based algorithms – a class of learning algorithms that instructs the analyst only to request labels on points he sees that reduce the hypothesis space significantly. In the most general setting of concept learning, this is a useful framework since the error guarantees can be described using the same setup as the PAC model. Moreover without additional information about the problem domain, it is unclear how to devise efficient algorithms that are more specific in instructing the analyst on what questions to ask.

We find that the stream-based model is in general difficult to analyze for the preference relations we work with. This difficulty seems to arise from the apparent need to quantify disagreement in order to explicitly write down learning guarantees. Though in most general situations it is unclear how to quantify disagreement for our preference relations, we present a redeeming situation for which we are able to provide a precise analysis of the learning guarantees for 𝒫𝒟\mathcal{P}_{\mathcal{ED}}. Here, the analyst can learn 𝒫𝒟\mathcal{P}_{\mathcal{ED}} with an exponential improvement in label complexity over the guarantees provided by the PAC model. This is achieved via a computation of the disagreement coefficient (defined in Section 5) of 𝒫𝒟\mathcal{P}_{\mathcal{ED}} for a specific distribution.

Theorem 3.5.

There exists a distribution μ\mu on T×T\mathbb{R}^{T}\times\mathbb{R}^{T} for which the disagreement coefficient of 𝒫𝒟\mathcal{P}_{\mathcal{ED}} is θ=2\theta=2. Thus, for this distribution,

CAL(ε)=O~(logTlog1ε),\ell_{CAL}(\varepsilon)=\widetilde{O}\left(\log T\log\frac{1}{\varepsilon}\right),

where the O~\widetilde{O} notation suppresses terms that are logarithmic in logT\log T and log1/ε\log 1/\varepsilon.

The measure μ\mu we construct is induced by the product Lebesgue measure on (0,1)T1(0,1)^{T-1}, and allows us to precisely translate statements about disagreement into statements about the roots of polynomials arising from a given choice. Once we have defined μ\mu, the calculation of θ\theta follows from basic probability arguments.

Now, in our case the analyst has structural information regarding the preference relation of the agent he is questioning. We find that allowing the analyst full control over the membership queries he makes yields a learning algorithm that, despite its simplicity, takes advantage of this extra structure and yields a significant improvement in complexity over the stream-based setting. Additionally, the membership queries model naturally describes an experimental environment in which the analyst is able to ask the agent questions in real time.

We show that when the preference model satisfies some relatively benign structural requirements, even very naive algorithms outperform the guarantees provided by CAL in the stream-based setting. The example algorithm we give, relying on a simple binary search, has a query complexity of O(log1/ε)O(\log 1/\varepsilon), which gets rid of the logT\log T dependence in Theorem 3.5.

The class of preference models is defined as follows: let g1,,gT:g_{1},\ldots,g_{T}:\mathbb{R}\to\mathbb{R} be a collection of functions satisfying the properties listed in Section 5.3 and consider the model of preference relations 𝒫\mathcal{P} parametrized by δ\delta where xyx\succsim y if and only if t=1Tgt(δ)xtt=1Tgt(δ)yt\sum_{t=1}^{T}g_{t}(\delta)x_{t}\geq\sum_{t=1}^{T}g_{t}(\delta)y_{t}77 7 As before, the g1,,gTg_{1},\ldots,g_{T} are known to the analyst.. We have

Proposition 3.6.

There exists an algorithm that takes as input ε>0\varepsilon>0 and using O(log1/ε)O(\log 1/\varepsilon) membership queries outputs δh\delta^{h} such that |δδh|ε|\delta-\delta^{h}|\leq\varepsilon, where δ\delta parametrizes the target preference relation in 𝒫\mathcal{P}.

The remainder of the paper is devoted to proving the results discussed in this section.

4 PAC Learning

In this section we prove Theorems 3.2 and 3.3. We first note a preliminary upper bound due to Basu and Echenique [2]. Let 𝒫\mathcal{P}_{\mathcal{I}} be the set of preference relations that satisfy the following axioms:

Order:

For all x,yx,y either xyx\succsim y or yxy\succsim x (completeness). For all x,y,zx,y,z, if xyx\succsim y and yzy\succsim z, then xzx\succsim z (transitivity).

Independence:

For all x,y,zx,y,z and for any λ(0,1)\lambda\in(0,1), xyx\succsim y if and only if λx+(1λ)zλy+(1λ)z\lambda x+(1-\lambda)z\succsim\lambda y+(1-\lambda)z.

The class 𝒫\mathcal{P}_{\mathcal{I}} satisfies the property that for any 𝒫\succsim\in\mathcal{P}_{\mathcal{I}}, there are finitely many vectors q1,,qKq_{1},\ldots,q_{K}, with KTK\leq T, such that xyx\succsim y if and only if (qk.x)k=1KL(qk.y)k=1K(q_{k}.x)_{k=1}^{K}\geq_{L}(q_{k}.y)_{k=1}^{K}, where L\geq_{L} denotes the lexicographic order [5]. Then, 𝒫𝒫𝒲,𝒫𝒫𝒲𝒫𝒲𝒫\mathcal{P}_{\mathcal{PW}},\mathcal{P}_{\mathcal{BPW}}\subset\mathcal{P}_{\mathcal{W}}\subset\mathcal{P}_{\mathcal{I}}, since the aforementioned characterization is satisfied with K=1K=1 and q1=(w1,,wT)q_{1}=(w_{1},\ldots,w_{T}).

This has two main consequences. First, 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} and 𝒫𝒫𝒲\mathcal{P}_{\mathcal{BPW}} (and thus all the discounted utility models) satisfy the measurability requirement discussed in Lemma 4 of [2] for the equivalence result of Theorem 2.1 to hold. Second, the VC dimensions of 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} and 𝒫𝒫𝒲\mathcal{P}_{\mathcal{BPW}} are all bounded above by T+1T+1 (and in particular T1VC(𝒫𝒲)T+1T-1\leq VC(\mathcal{P}_{\mathcal{W}})\leq T+1). This follows due to Theorem 3.1 of [2], in which an argument similar to that required to compute the VC dimension of the class of half-spaces is used to show that VC(𝒫)=T+1VC(\mathcal{P}_{\mathcal{I}})=T+1. In all cases excluding the most general model of discounted utility, we are able to bring this down to log(T1)\log(T-1) (which we then show is tight by demonstrating the corresponding lower bound).

We begin by demonstrating that even in the discounted utility setting, without any structure we cannot do better than the learning bounds obtained for 𝒫\mathcal{P}_{\mathcal{I}}.

Proposition 3.1.

T1VC(𝒫𝒟)T+1T-1\leq VC(\mathcal{P}_{\mathcal{D}})\leq T+1.

Proof.

That VC(𝒫𝒟)T+1VC(\mathcal{P}_{\mathcal{D}})\leq T+1 follows from Theorem 3.1 of [2], since 𝒫𝒟𝒫\mathcal{P}_{\mathcal{D}}\subset\mathcal{P}_{\mathcal{I}}.

Here is a simple construction that shows VC(𝒫𝒟)T1VC(\mathcal{P}_{\mathcal{D}})\geq T-1. Fix an ε>0\varepsilon>0. Let e1,,eTe_{1},\ldots,e_{T} be the standard unit vectors in T\mathbb{R}^{T}, and consider the set of points {(x1,y1),,(xT1,yT1)}\{(x^{1},y^{1}),\ldots,(x^{T-1},y^{T-1})\}, where xi=(1ε)eix^{i}=(1-\varepsilon)e_{i} and yi=ei+1y^{i}=e_{i+1}.

This set is shattered by 𝒫𝒟\mathcal{P}_{\mathcal{D}}: for any (ai)i=1T1(a_{i})_{i=1}^{T-1}, choose D(1)D(1) arbitrarily from (0,1)(0,1), and if D(i)D(i) has been defined, inductively define D(i+1)D(i+1) such that D(i+1)D(i)(1ε)D(i+1)\leq D(i)(1-\varepsilon) if ai=1a_{i}=1 and D(i)>D(i+1)>(1ε)D(i)D(i)>D(i+1)>(1-\varepsilon)D(i) if ai=0a_{i}=0. ∎

We now prove Theorems 3.2 and 3.3, which are restated below for convenience.

Theorem 3.2.

For every ε>0\varepsilon>0, there exists a dεd_{\varepsilon} such that for every ddεd\geq d_{\varepsilon} we have VC(𝒫𝒫𝒲),VC(𝒫𝒫𝒲)(1+ε)logdVC(\mathcal{P}_{\mathcal{PW}}),VC(\mathcal{P}_{\mathcal{BPW}})\leq(1+\varepsilon)\log d for any TT and any TT polynomials Q1,,QTQ_{1},\ldots,Q_{T} of degree at most dd.

Proof.

It suffices to establish the bound for 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PBW}}.

Let (z1,,zn)(z^{1},\ldots,z^{n}) be a set of points in T×T\mathbb{R}^{T}\times\mathbb{R}^{T}, zi=(xi,yi)z^{i}=(x^{i},y^{i}). For each zi=(xi,yi)z^{i}=(x^{i},y^{i}), define the plan fi:=xiyif^{i}:=x^{i}-y^{i}. Then, note that (z1,,zn)(z^{1},\ldots,z^{n}) is shattered by 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PBW}} if and only if ((f1,0),,(fn,0))((f^{1},0),\ldots,(f^{n},0)) is shattered by 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PBW}}. Hence, we may (and do) restrict attention to datasets of the form ((f1,0),,(fn,0))((f^{1},0),\ldots,(f^{n},0)).

We have that ((f1,0),,(fn,0))((f^{1},0),\ldots,(f^{n},0)) is shattered by 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PBW}} if and only if for all vectors (a1,,an){0,1}n(a_{1},\ldots,a_{n})\in\{0,1\}^{n}, there exists a δ\delta and β\beta (which determines the preference relation) such that

(QT(δ)fTi++Q1(δ)f1i)+(1β1)(QT(0)fTi++Q1(0)f1i)0 whenever ai=1,(Q_{T}(\delta)f^{i}_{T}+\cdots+Q_{1}(\delta)f^{i}_{1})+\left(\frac{1}{\beta}-1\right)(Q_{T}(0)f^{i}_{T}+\cdots+Q_{1}(0)f^{i}_{1})\geq 0\text{ whenever }a_{i}=1,

and

(QT(δ)fTi++Q1(δ)f1i)+(1β1)(QT(0)fTi++Q1(0)f1i)<0 whenever ai=0.(Q_{T}(\delta)f^{i}_{T}+\cdots+Q_{1}(\delta)f^{i}_{1})+\left(\frac{1}{\beta}-1\right)(Q_{T}(0)f^{i}_{T}+\cdots+Q_{1}(0)f^{i}_{1})<0\text{ whenever }a_{i}=0.

We first show that for all ε>0\varepsilon>0, for sufficiently large dd we have VC(𝒫𝒫𝒲)(1+ε)logdVC(\mathcal{P}_{\mathcal{PBW}})\leq(1+\varepsilon)\log d. Note that if the nn points ((f1,0),,(fn,0))((f^{1},0),\ldots,(f^{n},0)) can be shattered, there are polynomials P1,,PnP_{1},\ldots,P_{n} in δ\delta (where PiP_{i} is the polynomial Q1(δ)f1i++QT(δ)fTiQ_{1}(\delta)f^{i}_{1}+\cdots+Q_{T}(\delta)f^{i}_{T}), each of degree at most dd, such that for every labeling (a1,,an){0,1}n(a_{1},\ldots,a_{n})\in\{0,1\}^{n}, there exists a δ\delta and β\beta such that

(sgn(P1(δ)+(1/β1)P1(0)),,sgn(Pn(δ)+(1/β1)Pn(0)))=(a1,,an).(\operatorname{sgn}(P_{1}(\delta)+(1/\beta-1)P_{1}(0)),\ldots,\operatorname{sgn}(P_{n}(\delta)+(1/\beta-1)P_{n}(0)))=(a_{1},\ldots,a_{n}).

First, for any nn polynomials P1,,PnP_{1},\ldots,P_{n} of degree at most dd, we give an upper bound on the number of possible values (sgn(P1(δ)),,sgn(Pn(δ)))(\operatorname{sgn}(P_{1}(\delta)),\ldots,\operatorname{sgn}(P_{n}(\delta))) can realize. Each polynomial has at most dd real roots, so together P1,,PnP_{1},\ldots,P_{n} have at most ndnd distinct real roots. Since sign changes can only occur at the roots, there are at most nd+1nd+1 possible values of {0,1}n\{0,1\}^{n} that (sgn(P1(δ)),,sgn(Pn(δ)))(\operatorname{sgn}(P_{1}(\delta)),\ldots,\operatorname{sgn}(P_{n}(\delta))) can realize.

Now, for a fixed δ\delta, varying β\beta shifts the collection of polynomials

P1(δ)+(1/β1)P1(0),,Pn(δ)+(1/β1)Pn(0)P_{1}(\delta)+(1/\beta-1)P_{1}(0),\ldots,P_{n}(\delta)+(1/\beta-1)P_{n}(0)

vertically, which in the worst case induces sign changes in all entries. We thus get at most an additional nn new sign combinations for every sign combination realized by (sgn(P1(δ)),,sgn(Pn(δ)))(\operatorname{sgn}(P_{1}(\delta)),\ldots,\operatorname{sgn}(P_{n}(\delta))). Hence, there are at most nd+1+n(nd+1)=(n2+n)d+n+1nd+1+n(nd+1)=(n^{2}+n)d+n+1 possible values of {0,1}n\{0,1\}^{n} that

(sgn(P1(δ)+(1/β1)P1(0)),,sgn(Pn(δ)+(1/β1)Pn(0)))(\operatorname{sgn}(P_{1}(\delta)+(1/\beta-1)P_{1}(0)),\ldots,\operatorname{sgn}(P_{n}(\delta)+(1/\beta-1)P_{n}(0)))

can realize.

In order for all 2n2^{n} elements of {0,1}n\{0,1\}^{n} to be realized, it must be that

(n2+n)d+n+12n.(n^{2}+n)d+n+1\geq 2^{n}.

If n>(1+ε)logdn>(1+\varepsilon)\log d, then for large enough dd this inequality does not hold, and so any set of nn points cannot be shattered. Thus, for all ε>0\varepsilon>0, n(1+ε)logdn\leq(1+\varepsilon)\log d for large enough dd, i.e. VC(𝒫)(1+ε)logdVC(\mathcal{P})\leq(1+\varepsilon)\log d.∎

We now establish the corresponding lower bound when the polynomials Q1,,QTQ_{1},\ldots,Q_{T} span the space of polynomials of degree at most T1T-1.

Theorem 3.3.

Let Q1,,QTQ_{1},\ldots,Q_{T} be polynomials in δ\delta of degree at most T1T-1 that span the space of polynomials in δ\delta of degree at most T1T-1. Then VC(𝒫𝒫𝒲),VC(𝒫𝒫𝒲)log(T1)VC(\mathcal{P}_{\mathcal{PW}}),VC(\mathcal{P}_{\mathcal{BPW}})\geq\log(T-1).

Proof.

It suffices to establish the bound for 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}}.

Consider the graph on {0,1}n\{0,1\}^{n} where two vertices are connected by an edge if they differ in exactly one location. Fix a Hamiltonian path v1,v2,,v2nv_{1},v_{2},\ldots,v_{2^{n}} in this graph (the existence of which is well known). Let b1,2,,b2n1,2nb_{1,2},\ldots,b_{2^{n}-1,2^{n}} be the sequence where bi,i+1b_{i,i+1} is the index of the location at which viv_{i} and vi+1v_{i+1} differ. Note that if n=log(T1)n=\log(T-1), the graph has T1T-1 vertices, so each index in {1,,n}\{1,\ldots,n\} can appear in the sequence (bi,i+1)(b_{i,i+1}) at most T1T-1 times.

Now, let r1<r2<<r2nr_{1}<r_{2}<\cdots<r_{2^{n}} be any points in (0,1)(0,1). Define nn polynomials P1,,PnP_{1},\ldots,P_{n} by Pk(δ)=bi,i+1=k(δri),P_{k}(\delta)=\prod_{b_{i,i+1}=k}(\delta-r_{i}), so the roots of PkP_{k} are precisely the rir_{i}’s that correspond to a flip in the entry at the kkth position of a vertex in the path. Then, (sgn(P1(δ)),,sgn(Pn(δ)))(\operatorname{sgn}(P_{1}(\delta)),\ldots,\operatorname{sgn}(P_{n}(\delta))) realizes every element of {0,1}n\{0,1\}^{n}.

Since Q1,,QTQ_{1},\ldots,Q_{T} span the space of polynomials of degree at most T1T-1, for each PiP_{i} we can find f1i,,fTif^{i}_{1},\ldots,f^{i}_{T} such that

Pi(δ)=Q1(δ)f1i++QT(δ)fTi,P_{i}(\delta)=Q_{1}(\delta)f^{i}_{1}+\cdots+Q_{T}(\delta)f^{i}_{T},

which gives us a collection of log(T1)\log(T-1) points that is shattered. Hence log(T1)VC(𝒫)\log(T-1)\leq VC(\mathcal{P}). ∎

It is readily seen that 𝒫𝒟\mathcal{P}_{\mathcal{ED}} and 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} satisfy the conditions of Theorems 3.2 and 3.399 9 𝒫𝒟\mathcal{P}_{\mathcal{ED}} is given by 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} with Qt(δ)=δt1Q_{t}(\delta)=\delta^{t-1}, and 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} is given by 𝒫𝒫𝒲\mathcal{P}_{\mathcal{BPW}} with Qt(δ)=δt1Q_{t}(\delta)=\delta^{t-1}.. We give a quick argument verifying that 𝒫𝒟\mathcal{P}_{\mathcal{HD}} does as well: For 1tT1\leq t\leq T, let Qt(α)={1,,T}{t}(1+α)Q_{t}(\alpha)=\prod_{\ell\in\{1,\ldots,T\}\setminus\{t\}}(1+\ell\alpha) (these are the polynomials obtained by clearing denominators of the hyperbolic discount factors). We argue that {Qt(α)}t=1T\{Q_{t}(\alpha)\}_{t=1}^{T} are linearly independent over the vector space of polynomials in α\alpha of degree at most T1T-1. Indeed, if Q1(α)f1++QT(α)fT=0Q_{1}(\alpha)f_{1}+\cdots+Q_{T}(\alpha)f_{T}=0, then we must have that f1==fT=0f_{1}=\cdots=f_{T}=0 since at α=1t\alpha=\frac{-1}{t} we get Qt(α)ft=0Q_{t}(\alpha)f_{t}=0. Hence, 𝒫𝒟\mathcal{P}_{\mathcal{HD}} is simply 𝒫𝒫𝒲\mathcal{P}_{\mathcal{PW}} with Qt(α)={1,,T}{t}(1+α)Q_{t}(\alpha)=\prod_{\ell\in\{1,\ldots,T\}\setminus\{t\}}(1+\ell\alpha).

Remark. These bounds also hold in the scenario where the agent can report indifference in the data. More precisely, the condition for xyx\succsim y is now a strict inequality, and we have three possible labels for the pair (x,y)(x,y): +1+1 indicates that xyx\succsim y, 1-1 indicates that yxy\succsim x, and 00 indicates that xyx\sim y. Then, using a Hamiltonian path on {1,0,1}n\{-1,0,1\}^{n}, we can construct polynomials P1,,PnP_{1},\ldots,P_{n} such that (sgn(P1(δ)),,sgn(Pn(δ)))(\operatorname{sgn}(P_{1}(\delta)),\ldots,\operatorname{sgn}(P_{n}(\delta))) realizes all elements of {1,0,1}n\{-1,0,1\}^{n} as δ\delta ranges from 00 to 11 (where sgn\operatorname{sgn} is the true sign function).

4.1 A Remark on Efficient Learnability

While PAC learnability is a positive result, it does not take into account the computational complexity of computing a hypothesis. Blumer et. al. [6] show that any learning rule that outputs a hypothesis consistent with the data seen yields with high probability a hypothesis that has very low error. However, if the problem of outputting a consistent hypothesis is computationally intractable, PAC learnability on its own is perhaps unsatisfying. In this section we note that the discounted utility models of intertemporal choice are efficiently learnable. This is due to an algorithm of Grigor’ev and Vorobjov [12] for solving a system of polynomial inequalities.

For notational convenience, it will be useful to write 𝒫={𝒫T}T1\mathcal{P}=\{\mathcal{P}^{T}\}_{T\geq 1}, for each of the models above, where 𝒫T\mathcal{P}^{T} is the collection of preference relations for a given TT. Moreover, suppose acts are chosen from [1,1]T[-1,1]^{T} instead of T\mathbb{R}^{T}. It is clear that this does not change any of the analysis above.

Polynomial learnability, as defined by Blumer et. al. [6], stipulates that the learning rule be computable in poly(1/ε,1/δ,T)\text{poly}(1/\varepsilon,1/\delta,T)-time (where ε\varepsilon and δ\delta denote the error threshold and confidence threshold respectively). Polynomial learnability is equivalent to the task of outputting a hypothesis consistent with the given data set in polynomial time [6].

Definition 4.1.

A randomized polynomial hypothesis finder (r-poly hy-fi) for 𝒫\mathcal{P} is a randomized polynomial time algorithm that takes as input a sample of a preference relation in 𝒫\mathcal{P}, and for some γ>0\gamma>0, with probability at least γ\gamma produces a hypothesis that is consistent with the sample.

Theorem 4.1.

𝒫\mathcal{P} is properly polynomially learnable if and only if there is an r-poly hy-fi for 𝒫\mathcal{P} and VC(𝒫T)VC(\mathcal{P}^{T}) grows only polynomially in TT.

We have just shown that VC(𝒫T)log(T1)VC(\mathcal{P}^{T})\sim\log(T-1). Using techniques involving algebraic geometry, Grigor’ev and Vorobjov [12] present an algorithm to solve a system of NN polynomial inequalities with a poly(N,T)\text{poly}(N,T) runtime. This serves as our r-poly hy-fi, and thus we obtain:

Theorem 4.2.

𝒫𝒟\mathcal{P}_{\mathcal{ED}}, 𝒫𝒟\mathcal{P}_{\mathcal{HD}}, and 𝒫𝒬𝒟\mathcal{P}_{\mathcal{QHD}} are all properly polynomially learnable.

5 Active Learning

In this section we study two models of active learning: stream-based selective sampling and learning via membership queries. We first define a distribution for which the disagreement coefficient of 𝒫𝒟\mathcal{P}_{\mathcal{ED}} is 22, showing that disagreement methods (specifically the CAL algorithm [9, 10, 13]) in the stream-based model can yield an exponential improvement over the sample complexity of PAC learning (thus proving Theorem 3.5).

We then consider learning via membership queries and show that in this setting even very naive algorithms outperform the disagreement methods in the stream-based model (that is, the analyst needs to ask fewer questions to the agent in order to learn his preference than the number of label requests he would need to make using disagreement methods).

5.1 Preliminaries

For notational convenience, δ\succsim_{\delta} will to refer to the preference relation in 𝒫𝒟\mathcal{P}_{\mathcal{ED}} with discounting factor δ\delta.

Let μ\mu be a distribution on T×T\mathbb{R}^{T}\times\mathbb{R}^{T}. μ\mu induces a metric on 𝒫𝒟\mathcal{P}_{\mathcal{ED}} by d(δ,γ)=μ(δγ)d(\succsim_{\delta},\succsim_{\gamma})=\mu(\succsim_{\delta}\triangle\succsim_{\gamma}), and thus we can define the closed ball of radius RR centered at δ\succsim_{\delta} by

B(δ,R)={γ:d(δ,γ)R}.B(\succsim_{\delta},R)=\{\succsim_{\gamma}:d(\succsim_{\delta},\succsim_{\gamma})\leq R\}.

For V𝒫𝒰V\subseteq\mathcal{P}_{\mathcal{EU}}, the disagreement region of VV, Dis(V)\operatorname{Dis}(V) is defined by

Dis(V)={(x,y)T×T:δ,γVs.t.(x,y)δγ}=δ,γV(δγ)\operatorname{Dis}(V)=\{(x,y)\in\mathbb{R}^{T}\times\mathbb{R}^{T}:\exists\succsim_{\delta},\succsim_{\gamma}\in V\,\,s.t.\,\,(x,y)\in\,\succsim_{\delta}\triangle\succsim_{\gamma}\}=\displaystyle\bigcup_{\succsim_{\delta},\succsim_{\gamma}\in V}(\succsim_{\delta}\triangle\succsim_{\gamma})

Intuitively, Dis(V)\operatorname{Dis}(V) is the collection of points (x,y)(x,y) such that we can find two hypothesis relations in the current version space that rank xx and yy differently.

If δ\succsim_{\delta} is the target preference relation, the disagreement coefficient of δ\succsim_{\delta} with respect to μ\mu is the quantity

θ=supR>0μ(Dis(B(δ,R))R\theta=\sup_{R>0}\frac{\mu(\operatorname{Dis}(B(\succsim_{\delta},R))}{R}

5.2 Disagreement based active learning

In this subsection, we define a distribution μ\mu on T×T\mathbb{R}^{T}\times\mathbb{R}^{T} and show that the disagreement coefficient of 𝒫𝒟\mathcal{P}_{\mathcal{ED}} with respect to μ\mu is 22.

Choosing a measure

The main challenge here is that θ\theta depends on the underlying distribution over T×T\mathbb{R}^{T}\times\mathbb{R}^{T}. Since preferences are polynomial inequalities, the disagreement coefficient seems to lend itself to a characterization involving polynomials and their roots, which is the motivation for our choice of distribution. For a general distribution μ\mu over T×T\mathbb{R}^{T}\times\mathbb{R}^{T}, it is not clear how to compute the disagreement coefficient.

We show that θ=2\theta=2 for a suitably chosen distribution on T×T\mathbb{R}^{T}\times\mathbb{R}^{T}, which is induced by the Lebesgue measure on (0,1)T1(0,1)^{T-1}. This allows us to work with a measure on sets of roots of polynomials that arise from the definition of the preference relations.

Let μ\mu^{**} be a measure on (0,1)T1(0,1)^{T-1}. We interchangeably represent elements of T\mathbb{R}^{T} as polynomials PP of degree at most T1T-1 or as T1T-1-tuples of coefficients. Let \sim be the equivalence relation on T\mathbb{R}^{T} defined by PQP=cQP\sim Q\iff P=cQ for some constant cc, and let T/\mathbb{R}^{T}/\sim be the resulting quotient space. Let g:(0,1)T1T/g:(0,1)^{T-1}\to\mathbb{R}^{T}/\sim be the map taking a tuple of roots to the equivalence class of the polynomials with those roots, and let h:T×TT/h:\mathbb{R}^{T}\times\mathbb{R}^{T}\to\mathbb{R}^{T}/\sim be the map h(x,y)=[xy]h(x,y)=[x-y].

We define the following measures μ\mu^{*} and μ\mu on g((0,1)T1)T/g((0,1)^{T-1})\subset\mathbb{R}^{T}/\sim and h1(g((0,1)T1))T×Th^{-1}(g((0,1)^{T-1}))\subset\mathbb{R}^{T}\times\mathbb{R}^{T} respectively.

  • Define μ\mu^{*} on all sets Sg((0,1)T1)S\subset g((0,1)^{T-1}) such that

    {(r1(P),,rT1(P))(0,1)T1:PS}\{(r_{1}(P),\ldots,r_{T-1}(P))\in(0,1)^{T-1}:P\in S\}

    (where r1(P),rT1(P)r_{1}(P),\ldots r_{T-1}(P) denote the roots of PP) is μ\mu^{**}-measurable, for which we set

    μ(S)=μ({(r1(P),,rT1(P))(0,1)T1:PS}).\mu^{*}(S)=\mu^{**}(\{(r_{1}(P),\ldots,r_{T-1}(P))\in(0,1)^{T-1}:P\in S\}).
  • Define μ\mu on all sets Sh1(g((0,1)T1))S\subset h^{-1}(g((0,1)^{T-1})) such that

    {[z]g((0,1)T1):(x,y)Ss.t.zxy}\{[z]\in g((0,1)^{T-1}):\exists(x,y)\in S\,s.t.\,z\sim x-y\}

    is μ\mu^{*}-measurable, for which we set

    μ(S)=μ({[z]g((0,1)T1):(x,y)Ss.t.zxy}).\mu(S)=\mu^{*}(\{[z]\in g((0,1)^{T-1}):\exists(x,y)\in S\,s.t.\,z\sim x-y\}).

Intuitively, μ\mu^{*} is defined only on those polynomials that have all their roots in (0,1)(0,1). When T=2T=2, this is a desirable property since the analyst is only presented with polynomials that have some disagreement in (0,1)(0,1). He is not presented with meaningless polynomials that are, for example, always positive on (0,1)(0,1) (the analyst has nothing to learn from such polynomials since such a polynomial will be preferred to 00 for all δ(0,1)\delta\in(0,1)). For T3T\geq 3, this is a more restrictive property since the analyst is only presented with polynomials that have all T1T-1 roots in (0,1)(0,1).

Let μ\mu^{**} be the product Lebesgue measure on (0,1)T1(0,1)^{T-1}. Choosing μ\mu^{**} in this fashion allows us to neatly characterize B(δ,R)B(\succsim_{\delta},R).

Let X1,,XT1X_{1},\ldots,X_{T-1} be uniform i.i.d. random variables on (0,1)(0,1), and let Yδ,γY_{\delta,\gamma} be the random variable Yδ,γ=|{i:Xi is between δ and γ}|Y_{\delta,\gamma}=|\{i:X_{i}\text{ is between }\delta\text{ and }\gamma\}|. Let Eδ,γoddE^{odd}_{\delta,\gamma} denote the event that Yδ,γY_{\delta,\gamma} is odd, let Eδ,γkE^{k}_{\delta,\gamma} denote the event Yδ,γ=kY_{\delta,\gamma}=k, and let Eδ,γkE^{\geq k}_{\delta,\gamma} denote the event Yδ,γkY_{\delta,\gamma}\geq k.

Lemma 5.1.

γB(δ,R)\succsim_{\gamma}\in B(\succsim_{\delta},R) if and only if [Eδ,γodd]R\mathbb{P}[E^{odd}_{\delta,\gamma}]\leq R.

Proof.

Given (x,y)T×T(x,y)\in\mathbb{R}^{T}\times\mathbb{R}^{T}, let Pxy(X)=t=1TXt1(xtyt)P_{x-y}(X)=\sum_{t=1}^{T}X^{t-1}\cdot(x_{t}-y_{t}). Then,

γB(δ,R)\displaystyle\succsim_{\gamma}\in B(\succsim_{\delta},R) μ({(x,y)h1(g((0,1)T1)):sgn(Pxy(δ))sgn(Pxy(γ)))})R\displaystyle\iff\mu(\{(x,y)\in h^{-1}(g((0,1)^{T-1})):\operatorname{sgn}(P_{x-y}(\delta))\neq\operatorname{sgn}(P_{x-y}(\gamma)))\})\leq R
μ({[P]g((0,1)T1):sgn(P(δ))sgn(P(γ))})R\displaystyle\iff\mu^{*}(\{[P]\in g((0,1)^{T-1}):\operatorname{sgn}(P(\delta))\neq\operatorname{sgn}(P(\gamma))\})\leq R
μ({(r1,,rT1)(0,1)T1:sgn((δri))sgn((γri))})R\displaystyle\iff\mu^{**}(\{(r_{1},\ldots,r_{T-1})\in(0,1)^{T-1}:\operatorname{sgn}(\textstyle\prod(\delta-r_{i}))\neq\operatorname{sgn}(\textstyle\prod(\gamma-r_{i}))\})\leq R

But sgn((δri))sgn((γri))\operatorname{sgn}(\textstyle\prod(\delta-r_{i}))\neq\operatorname{sgn}(\textstyle\prod(\gamma-r_{i})) occurs exactly when an odd number of roots lie between γ\gamma and δ\delta (modulo a set of measure 00 since the probability that we have a root of multiplicity greater than 11 is 00). ∎

While it seems difficult to write down a general characterization of μ\mu, Propositions 5.2 and 5.3 give some basic observations regarding the σ\sigma-algebras on which μ\mu^{*} and μ\mu are defined. We defer their proofs (along with a description of μ\mu^{*} in the case T=2T=2) to the appendix:

Proposition 5.2.

The σ\sigma-algebra on g((0,1)T1)g((0,1)^{T-1}) is the Borel σ\sigma-algebra.

The σ\sigma-algebra induced on h1(g((0,1)T1))h^{-1}(g((0,1)^{T-1})) does not appear to yield a clean characterization, but we can show the weaker statement that μ\mu is a Borel measure, i.e. it is defined on all open sets of h1(g((0,1)T1))h^{-1}(g((0,1)^{T-1})).

Proposition 5.3.

μ\mu is a Borel measure on h1(g((0,1)T1))h^{-1}(g((0,1)^{T-1})).

Computing θ\theta

We now show θ=2\theta=2 for the distribution μ\mu as chosen above. We use the notation d=dδ,γ:=|δγ|d=d_{\delta,\gamma}:=|\delta-\gamma| to denote the distance between δ\delta and γ\gamma. Let δ\succsim_{\delta} be the target preference relation.

First, note that since Yδ,γY_{\delta,\gamma} is distributed according to Bin(T1,d)\text{Bin}(T-1,d), [Eδ,γodd]=1(12d)T12\mathbb{P}[E^{odd}_{\delta,\gamma}]=\frac{1-(1-2d)^{T-1}}{2}1010 10 This is due to the general fact that if XX is a random variable distributed according to Bin(n,p)\text{Bin}(n,p), the probability that XX is odd is 1(12p)n2\frac{1-(1-2p)^{n}}{2}.. We us this fact to derive an explicit description of the preference relations γ\succsim_{\gamma} contained in the ball B(δ,R)B(\succsim_{\delta},R) in terms of dd. We break the analysis up into a few cases.

First, when R>12R>\frac{1}{2}, we have supR>12μ(Dis(B(δ,R)))R=2\sup_{R>\frac{1}{2}}\frac{\mu(\operatorname{Dis}(B(\succsim_{\delta},R)))}{R}=2. Let R12R\leq\frac{1}{2}. By Lemma 5.1,

γB(δ,R)[Eδ,γodd]=1(12d)T12R\succsim_{\gamma}\in B(\succsim_{\delta},R)\iff\mathbb{P}[E^{odd}_{\delta,\gamma}]=\frac{1-(1-2d)^{T-1}}{2}\leq R (1)

Suppose d12d\leq\frac{1}{2}. Then 12d1-2d and 12R1-2R are both non-negative, so rearranging Equation (1) yields

d1(12R)1/(T1)2.d\leq\frac{1-(1-2R)^{1/(T-1)}}{2}. (2)

Suppose d>12d>\frac{1}{2}, so 12d<01-2d<0. Rearranging Equation (1), we get 12R(12d)T11-2R\leq(1-2d)^{T-1}. If T1T-1 is odd, (12d)T1(1-2d)^{T-1} is negative, so 12R(12d)T11-2R\leq(1-2d)^{T-1} does not hold. Thus, when T1T-1 is odd the ball consists of γ\succsim_{\gamma} such that dd satisfies condition (2). If T1T-1 is even, (12d)T1(1-2d)^{T-1} is positive, so we get

d1+(12R)1/(T1)2.d\geq\frac{1+(1-2R)^{1/(T-1)}}{2}. (3)

Thus, when T1T-1 is even the ball consists of γ\succsim_{\gamma} such that dd satisfies conditions (2) or (3).

Now, the disagreement region of B(δ,R)B(\succsim_{\delta},R) consists of all points (x,y)(x,y) such that the polynomial PxyP_{x-y} (as defined in Lemma 5.1) has a root γ\gamma such that γB(δ,R)\succsim_{\gamma}\in B(\succsim_{\delta},R) (since we can find two hypotheses that disagree on (x,y)(x,y) by taking a point slightly below γ\gamma and a point slightly above γ\gamma such that PxyP_{x-y} has no sign changes in between). Hence, with R1=1(12R)1/(T1)2R_{1}=\frac{1-(1-2R)^{1/(T-1)}}{2} and R2=1+(12R)1/(T1)2R_{2}=\frac{1+(1-2R)^{1/(T-1)}}{2}, we have that

μ(Dis(B(δ,R)))={[EδR1,δ+R11]if T1 is odd[EδR1,δ+R11E0,δR21Eδ+R2,11]if T1 is even\mu(\operatorname{Dis}(B(\succsim_{\delta},R)))=\begin{cases}\mathbb{P}[E_{\delta-R_{1},\delta+R_{1}}^{\geq 1}]&\text{if }T-1\text{ is odd}\\ \mathbb{P}[E_{\delta-R_{1},\delta+R_{1}}^{\geq 1}\cup E_{0,\delta-R_{2}}^{\geq 1}\cup E_{\delta+R_{2},1}^{\geq 1}]&\text{if }T-1\text{ is even}\end{cases}

We have

[EδR1,δ+R11]=1(12R1)T1=2R,\mathbb{P}[E_{\delta-R_{1},\delta+R_{1}}^{\geq 1}]=1-(1-2R_{1})^{T-1}=2R,

and

[EδR1,δ+R11E0,δR21Eδ+R2,11]=1(2(R2R1))T1=12T1(12R).\mathbb{P}[E_{\delta-R_{1},\delta+R_{1}}^{\geq 1}\cup E_{0,\delta-R_{2}}^{\geq 1}\cup E_{\delta+R_{2},1}^{\geq 1}]=1-(2(R_{2}-R_{1}))^{T-1}=1-2^{T-1}(1-2R).

Therefore, when T1T-1 is odd sup0<R1/2μ(Dis(B(δ,R)))R=2\sup_{0<R\leq 1/2}\frac{\mu(\operatorname{Dis}(B(\succsim_{\delta},R)))}{R}=2 and when T1T-1 is even

sup0<R1/2μ(Dis(B(δ,R)))R=sup0<R1/212T1(12R)R=2,\sup_{0<R\leq 1/2}\frac{\mu(\operatorname{Dis}(B(\succsim_{\delta},R)))}{R}=\sup_{0<R\leq 1/2}\frac{1-2^{T-1}(1-2R)}{R}=2,

which is achieved at R=1/2R=1/2 since 12T1(12R)R\frac{1-2^{T-1}(1-2R)}{R} is increasing on 0<R120<R\leq\frac{1}{2}.

Finally, θ=supR>0μ(Dis(B(δ,R)))R=2\theta=\sup_{R>0}\frac{\mu(\operatorname{Dis}(B(\succsim_{\delta},R)))}{R}=2.

We have thus established Theorem 3.5:

Theorem 3.5.

There exists a distribution μ\mu on T×T\mathbb{R}^{T}\times\mathbb{R}^{T} for which the disagreement coefficient of 𝒫𝒟\mathcal{P}_{\mathcal{ED}} is θ=2\theta=2. Thus, for this distribution,

CAL(ε)=O~(logTlog1ε),\ell_{CAL}(\varepsilon)=\widetilde{O}\left(\log T\log\frac{1}{\varepsilon}\right),

where the O~\widetilde{O} notation suppresses terms that are logarithmic in logT\log T and log1/ε\log 1/\varepsilon.

5.3 Learning via membership queries

In this subsection, we present a simple membership queries algorithm that outperforms the guarantees provided by the disagreement based CAL algorithm. For simplicity, we restrict attention to preference models that are parametrized by a single parameter (e.g. 𝒫𝒟\mathcal{P}_{\mathcal{ED}} and 𝒫𝒟\mathcal{P}_{\mathcal{HD}}).

Let g1,,gT:g_{1},\ldots,g_{T}:\mathbb{R}\to\mathbb{R} be a collection of functions such that there exist 1t1,t2T1\leq t_{1},t_{2}\leq T satisfying

  1. 1.

    M:=supδgt1(δ)gt2(δ)M:=\sup_{\delta}\frac{g_{t_{1}}(\delta)}{g_{t_{2}}(\delta)} is finite, and

  2. 2.

    The map δgt1(δ)gt2(δ)\delta\mapsto\frac{g_{t_{1}}(\delta)}{g_{t_{2}}(\delta)} satisfies an inverse Lipschitz condition with constant CC :

    |δδ|C|gt1(δ)gt2(δ)gt1(δ)gt2(δ)|.|\delta-\delta^{\prime}|\leq C\left|\frac{g_{t_{1}}(\delta)}{g_{t_{2}}(\delta)}-\frac{g_{t_{1}}(\delta^{\prime})}{g_{t_{2}}(\delta^{\prime})}\right|.

Consider the model of preference relations 𝒫\mathcal{P} parametrized by δ\delta such that

xy if and only if t=1Tgt(δ)xtt=1Tgt(δ)yt.x\succsim y\,\,\text{ if and only if }\sum_{t=1}^{T}g_{t}(\delta)x_{t}\geq\sum_{t=1}^{T}g_{t}(\delta)y_{t}.
Proposition 3.6.

There exists an algorithm that takes as input ε>0\varepsilon>0 and using O(log1/ε)O(\log 1/\varepsilon) membership queries outputs δh\delta^{h} such that |δδh|ε|\delta-\delta^{h}|\leq\varepsilon, where δ\delta parametrizes the target preference relation in 𝒫\mathcal{P}.

Proof.

Fix a ρ>0\rho>0 and an η\eta-cover of [0,Mρ][0,M\rho], where 0<ηρεC0<\eta\leq\frac{\rho\varepsilon}{C}. Let bρb_{\rho} be the quantity such that the agent is indifferent between receiving a payoff of ρ\rho at time t1t_{1} or receiving a payoff of bρb_{\rho} at time t2t_{2}, i.e. bρb_{\rho} solves

gt2(δ)bρ=gt1(δ)ρ.g_{t_{2}}(\delta)b_{\rho}=g_{t_{1}}(\delta)\rho.

By running a binary search over the η\eta-cover of [ρ,Mρ][\rho,M\rho], the analyst can find an approximation bρhb_{\rho}^{h} to the indifference point for which |bρbρh|η|b_{\rho}-b_{\rho}^{h}|\leq\eta (the binary search is performed on the parameter bρhb_{\rho}^{h} by requesting labels for pairs of the form (ρet1,bρhet2)(\rho e_{t_{1}},b_{\rho}^{h}e_{t_{2}})). The analyst then outputs the δh\delta^{h} that solves gt2(δh)bρh=gt1(δh)ρg_{t_{2}}(\delta^{h})b_{\rho}^{h}=g_{t_{1}}(\delta^{h})\rho.

We have

|δδh|C|gt1(δ)gt2(δ)gt1(δh)gt2(δh)|=C|bρρbρhρ|Cηρε,|\delta-\delta^{h}|\leq C\left|\frac{g_{t_{1}}(\delta)}{g_{t_{2}}(\delta)}-\frac{g_{t_{1}}(\delta^{h})}{g_{t_{2}}(\delta^{h})}\right|=C\left|\frac{b_{\rho}}{\rho}-\frac{b_{\rho}^{h}}{\rho}\right|\leq\frac{C\eta}{\rho}\leq\varepsilon,

as desired.

Since M:=supδgt1(δ)gt2(δ)M:=\sup_{\delta}\frac{g_{t_{1}}(\delta)}{g_{t_{2}}(\delta)} is finite, bρMρb_{\rho}\leq M\rho, so the binary search over the η\eta-cover of [0,Mρ][0,M\rho] terminates. ∎

Remark. Outputting a hypothesis parameter δh\delta^{h} that is ε\varepsilon-close to δ\delta is a reasonable measurement for the error of learning via membership queries since there is no underlying distribution providing points to the analyst. However, note that for a distribution on T×T\mathbb{R}^{T}\times\mathbb{R}^{T}, a hypothesis close to the target parameter implies the set of misclassified points is assigned a small measure, due to continuity of measure.

The main feature of this algorithm is that its query complexity has no dependence on the number of time periods TT. Both 𝒫𝒟\mathcal{P}_{\mathcal{ED}} and 𝒫𝒟\mathcal{P}_{\mathcal{HD}} fit the conditions of Proposition 3.6, and thus we obtain a large improvement over the guarantees provided by disagreement methods in the stream-based model. Such methods assume no extra knowledge about the problem domain and are written to fit a wide class of learning problems. When we are learning economic parameters, membership queries allow us to take advantage of the extra structure present in preference models.

Acknowledgements

We would like to thank Federico Echenique and Adam Wierman for several helpful comments and suggestions.

References

  • [1] Balcan, M. F., Daniely, A., Mehta, R., Urner, R., & Vazirani, V. V. (2014, December). Learning economic parameters from revealed preferences. In International Conference on Web and Internet Economics (pp. 338-353). Springer, Cham.
  • [2] Basu, P., & Echenique, F. (2018, February). Learnability and Models of Decision Making under Uncertainty. In Proceedings of the 2018 ACM Conference on Economics and Computation (pp. 53-53). ACM.
  • [3] Beigman, E., & Vohra, R. (2006, June). Learning from revealed preference. In Proceedings of the 7th ACM Conference on Electronic Commerce (pp. 36-42). ACM.
  • [4] Berns, G. S., Laibson, D., & Loewenstein, G. (2007). Intertemporal choice – toward an integrative framework. Trends in cognitive sciences, 11(11), 482-488.
  • [5] Blume, L., Brandenburger, A., & Dekel, E. (1991). Lexicographic probabilities and choice under uncertainty. Econometrica: Journal of the Econometric Society, 61-79.
  • [6] Blumer, A., Ehrenfeucht, A., Haussler, D., & Warmuth, M. K. (1989). Learnability and the Vapnik-Chervonenkis dimension. Journal of the ACM (JACM), 36(4), 929-965.
  • [7] Chabris, C. F., Laibson, D. I., & Schuldt, J. P. (2010). Intertemporal choice. In Behavioural and Experimental Economics (pp. 168-177). Palgrave Macmillan, London.
  • [8] Chambers, C. P., & Echenique, F. (2018). On multiple discount rates. Econometrica, 86(4), 1325-1346.
  • [9] Cohn, D., Atlas, L., & Ladner, R. (1994). Improving generalization with active learning. Machine learning, 15(2), 201-221.
  • [10] Dasgupta, S. (2011). Two faces of active learning. Theoretical computer science, 412(19), 1767-1781.
  • [11] Echenique, F., Golovin, D., & Wierman, A. (2011, June). A revealed preference approach to computational complexity in economics. In Proceedings of the 12th ACM conference on Electronic commerce (pp. 101-110). ACM.
  • [12] Grigoriev, D., & Vorobjov, N. (1988). Solving systems of polynomial inequalities in subexponential time. J. Symb. Comput., 5(1/2), 37-64.
  • [13] Hanneke, S. (2009). Theoretical foundations of active learning (No. CMU-ML-09-106). CARNEGIE-MELLON UNIV PITTSBURGH PA MACHINE LEARNING DEPT.
  • [14] Hanneke, S. (2016). The optimal sample complexity of PAC learning. The Journal of Machine Learning Research, 17(1), 1319-1333.
  • [15] Kable, J. W., & Glimcher, P. W. (2007). The neural correlates of subjective value during intertemporal choice. Nature neuroscience, 10(12), 1625.
  • [16] Kalai, G. (2003). Learnability and rationality of choice. Journal of Economic theory, 113(1), 104-117.
  • [17] Kleinberg, J., & Oren, S. (2014, June). Time-inconsistent planning: a computational problem in behavioral economics. In Proceedings of the fifteenth ACM conference on Economics and computation (pp. 547-564). ACM.
  • [18] Koopmans, T. C. (1960). Stationary ordinal utility and impatience. Econometrica: Journal of the Econometric Society, 287-309.
  • [19] Kreps, D. M. (1988). Notes on the theory of choice (Underground Classics in Economics). Westview Press Incorporated.
  • [20] Phelps, E. S., & Pollak, R. A. (1968). On second-best national saving and game-equilibrium growth. The Review of Economic Studies, 35(2), 185-199.
  • [21] Stern, N., Peters, S., Bakhshi, V., Bowen, A., Cameron, C., Catovsky, S., … & Edmonson, N. (2006). Stern Review: The economics of climate change (Vol. 30, p. 2006). London: HM treasury.
  • [22] Zadimoghaddam, M., & Roth, A. (2012, December). Efficiently learning from revealed preference. In International Workshop on Internet and Network Economics (pp. 114-127). Springer, Berlin, Heidelberg.

Appendix A Properties of μ\mu and μ\mu^{*}

Proof of Proposition 5.2.

Let \mathcal{B} denote the Borel σ\sigma-algebra on (0,1)T1(0,1)^{T-1}.

Let X=T/X=\mathbb{R}^{T}/\sim endowed with the quotient topology, and g:[0,1]T1Xg^{*}:[0,1]^{T-1}\to X be the map g(r1,,rT1)=[(xri)]g^{*}(r_{1},\ldots,r_{T-1})=[\prod(x-r_{i})]. Explicitly, the terms of g(r1,,rT1)g^{*}(r_{1},\ldots,r_{T-1}) are given by symmetric sums:

g(r1,,rT1)=(c,ciri,ci,jrirj,,(1)T1cr1rT1),g^{*}(r_{1},\ldots,r_{T-1})=\left(c,-c\sum_{i}r_{i},c\sum_{i,j}r_{i}r_{j},\ldots,(-1)^{T-1}cr_{1}\cdots r_{T-1}\right),

where cc is the appropriate constant for the representative of the equivalence class. Each symmetric sum is a continuous function of T1T-1 variables, so gg^{*} is continuous. Moreover, note that gg^{*} is injective. Then, with Y=g([0,1]T1)Y=g^{*}([0,1]^{T-1}), we have that g:[0,1]T1Yg^{*}:[0,1]^{T-1}\to Y is a continuous bijection from a compact set into a Hausdorff space. Hence, gg^{*} is a homeomorphism. Then gg, which is the restriction of gg^{*} to (0,1)T1(0,1)^{T-1} is a homeomorphism onto Z:=g((0,1)T1)Z:=g((0,1)^{T-1}). Thus, the σ\sigma-algebra g()g(\mathcal{B}) that we obtain on ZZ is the Borel σ\sigma-algebra. ∎

When T=2T=2, we can give an explicit description of μ\mu^{*}. Identify 2/\mathbb{R}^{2}/\sim with the unit circle. Then, a degree 11 polynomial PP is identified with the point (cosθ,sinθ)(\cos\theta,\sin\theta), where P(x)=(cosθ)x+sinθP(x)=(\cos\theta)x+\sin\theta. Z:=g((0,1)T1)Z:=g((0,1)^{T-1}) consists of the boundary of the unit circle for which the argument θ\theta satisfies tanθ(0,1)-\tan\theta\in(0,1). This is satisfied precisely for θ(3π/4,π)(7π/4,2π)\theta\in(3\pi/4,\pi)\cup(7\pi/4,2\pi). Hence, if UU is a basic open subset of {(cosθ,sinθ):θ(3π/4,π)(7π/4,2π)}\{(\cos\theta,\sin\theta):\theta\in(3\pi/4,\pi)\cup(7\pi/4,2\pi)\}, we can write U={(cosθ,sinθ):θ1<θ<θ2}U=\{(\cos\theta,\sin\theta):\theta_{1}<\theta<\theta_{2}\} with θ1,θ2\theta_{1},\theta_{2} both in the same segment of the unit circle and

μ(U)=μ({tanθ:θ1<θ<θ2})=|tanθ1tanθ2|.\mu^{*}(U)=\mu^{**}(\{-\tan\theta:\theta_{1}<\theta<\theta_{2}\})=|\tan\theta_{1}-\tan\theta_{2}|.
Proof of Proposition 5.3.

Let h:T×TT/h:\mathbb{R}^{T}\times\mathbb{R}^{T}\to\mathbb{R}^{T}/\sim be the map h(x,y)=[xy]h(x,y)=[x-y]. Let Vh1(g((0,1)T1))V\subseteq h^{-1}(g((0,1)^{T-1})) be open. We show that

h(V)={[z]:(x,y)V(z=xy)}h(V)=\{[z]:\exists(x,y)\in V(z=x-y)\}

is open. Indeed, let z=xyz=x-y for (x,y)V(x,y)\in V and choose ε\varepsilon small enough such that the square with vertices {(x+ε,y+ε),(x+ε,yε),(xε,y+ε),(xε,yε)}\{(x+\varepsilon,y+\varepsilon),(x+\varepsilon,y-\varepsilon),(x-\varepsilon,y+\varepsilon),(x-\varepsilon,y-\varepsilon)\} is contained in VV. Then, for any λε\lambda\leq\varepsilon, [z+λ]=[(x+λ)y][z+\lambda]=[(x+\lambda)-y] with (x+λ,y)V(x+\lambda,y)\in V and [zλ]=[x(y+λ)][z-\lambda]=[x-(y+\lambda)] with (x,y+λ)V(x,y+\lambda)\in V, so in particular the open ball with radius λ\lambda centered at [z][z] is contained in h(V)h(V). ∎