Myopic procedures are shown to be asymptotically optimal in ranking and selection problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Global optimization in Bayesian inference yields little additional benefit.
Paper confirms Feldman's conjecture on two-armed bandit problem.
Myopic investors make suboptimal choices that benefit others, leading to market inefficiencies.
Myopic optimization outperforms reinforcement learning in portfolio management, leading to lower returns and higher risks.
Efficiently optimizes constrained problems with two-step lookahead BO.
This paper optimizes sampling policies for Bayesian optimization to improve exploration and exploitation.
Lookahead, also known as non-myopic, Bayesian optimization (BO) aims to find optimal sampling policies through solving a dynamic program (DP) that maximizes a long-term reward over a rolling horizon. Though promising, lookahead BO faces the risk of error propagation through its increased dependence on a possibly mis-sp…
New RL algorithms find SNE in Markov games with myopic followers.
Efficiently recovers network community structure from clients' small subgraphs.
We maximize the expected utility from terminal wealth for an HARA investor when the market price of risk is an unobservable random variable. We compute the optimal portfolio explicitly and explore the effects of learning by comparing it with the corresponding myopic policy. In particular, we show that, for a market pri…
This paper shows how diverse tasks can make inefficient exploration in MTRL efficient.
We design a new myopic strategy for a wide class of sequential design of experiment (DOE) problems, where the goal is to collect data in order to to fulfil a certain problem specific goal. Our approach, Myopic Posterior Sampling (MPS), is inspired by the classical posterior (Thompson) sampling algorithm for multi-armed…
Algorithm learns optimal coordination for strategic agents in uncertain settings.
This paper derives a portfolio decomposition formula when the agent maximizes utility of her wealth at some finite planning horizon. The financial market is complete and consists of multiple risky assets (stocks) plus a risk free asset. The stocks are modelled as exponential Brownian motions with drift and volatility b…
New -step policy gradient method avoids local optima in restricted policy classes.
Portfolio turnpikes state that, as the investment horizon increases, optimal portfolios for generic utilities converge to those of isoelastic utilities. This paper proves three kinds of turnpikes. In a general semimartingale setting, the abstract turnpike states that optimal final payoffs and portfolios converge under …
The paper calculates how fast optimal investment strategies approach CRRA strategies in stochastic factor models.
NM-PPG optimizes adaptive feature acquisition in POMDPs for better predictions.
We present GLASSES: Global optimisation with Look-Ahead through Stochastic Simulation and Expected-loss Search. The majority of global optimisation approaches in use are myopic, in only considering the impact of the next function value; the non-myopic approaches that do exist are able to consider only a handful of futu…
Empirical study shows carriers ignore past shippers' behavior, focusing only on current actions.
New method optimizes costly functions with unknown costs and budget constraints.
Finite-horizon sequential experimental design (SED) arises naturally in many contexts, including hyperparameter tuning in machine learning among more traditional settings. Computing the optimal policy for such problems requires solving Bellman equations, which are generally intractable. Most existing work resorts to se…
Efficiently reduces computational burden of rollout acquisition functions in Bayesian optimization.
We consider two active binary-classification problems with atypical objectives. In the first, active search, our goal is to actively uncover as many members of a given class as possible. In the second, active surveying, our goal is to actively query points to ultimately predict the proportion of a given class. Numerous…
Study on predictable forward processes in trading without frequent evaluations.
We provide a new characterization of mean-variance hedging strategies in a general semimartingale market. The key point is the introduction of a new probability measure which turns the dynamic asset allocation problem into a myopic one. The minimal martingale measure relative to coincides with t…
PFNs4BO uses neural processes for flexible Bayesian Optimization.
This paper studies the properties of discrete time stochastic optimal control problems associated with portfolio selection. We investigate if optimal continuous time strategies can be used effectively for a discrete time market after a straightforward discretization. We found that Merton's strategy approximates the per…
A new Bayesian method optimizes time-dependent expensive functions with lookahead.
New algorithms optimize time series classification speed and accuracy.
Study optimal portfolio strategy with sporadic bankruptcy for isoelastic utility.
Many recommendation algorithms rely on user data to generate recommendations. However, these recommendations also affect the data obtained from future users. This work aims to understand the effects of this dynamic interaction. We propose a simple model where users with heterogeneous preferences arrive over time. Based…
A method to combine saliency metrics for better CNN pruning decisions.
UVU simplifies value uncertainty quantification in RL.
New algorithm learns optimal policies in strategic MDPs with private types.
We determine the optimal amount to invest in a Black-Scholes financial market for an individual who consumes at a rate equal to a constant proportion of her wealth and who wishes to minimize the expected time that her wealth spends in drawdown during her lifetime. Drawdown occurs when wealth is less than some fixed pro…
We consider the problem of learning the functions computing children from parents in a Structural Causal Model once the underlying causal graph has been identified. This is in some sense the second step after causal discovery. Taking a probabilistic approach to estimating these functions, we derive a natural myopic act…
We propose a general-purpose approach to discovering active learning (AL) strategies from data. These strategies are transferable from one domain to another and can be used in conjunction with many machine learning models. To this end, we formalize the annotation process as a Markov decision process, design universal s…
We consider the problem faced by a service platform that needs to match limited supply with demand but also to learn the attributes of new users in order to match them better in the future. We introduce a benchmark model with heterogeneous "workers" (demand) and a limited supply of "jobs" that arrive over time. Job typ…
New methods improve global optimisation for expensive functions using lookahead strategies.
Frequently, acquiring training data has an associated cost. We consider the situation where the learner may purchase data during training, subject TO a budget. IN particular, we examine the CASE WHERE each feature label has an associated cost, AND the total cost OF ALL feature labels acquired during training must NOT e…
A new mechanism reduces expert belief regret in online forecasting.
We propose a contextual bandit based model to capture the learning and social welfare goals of a web platform in the presence of myopic users. By using payments to incentivize these agents to explore different items/recommendations, we show how the platform can learn the inherent attributes of items and achieve a subli…
Bayesian methods improve drug discovery experiment design.
A minimal model of a market of myopic non-cooperative agents who trade bilaterally with random bids reproduces qualitative features of short-term electric power markets, such as those in California and New England. Each agent knows its own budget and preferences but not those of any other agent. The near-equilibrium pr…
Develops Heuristic Portfolio Optimization (HPO) as an information-restricted projection of Markowitz/tangency solution
The computational costs of inference and planning have confined Bayesian model-based reinforcement learning to one of two dismal fates: powerful Bayes-adaptive planning but only for simplistic models, or powerful, Bayesian non-parametric models but using simple, myopic planning strategies such as Thompson sampling. We …