Global optimization in Bayesian inference yields little additional benefit.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper confirms Feldman's conjecture on two-armed bandit problem.
We design a new myopic strategy for a wide class of sequential design of experiment (DOE) problems, where the goal is to collect data in order to to fulfil a certain problem specific goal. Our approach, Myopic Posterior Sampling (MPS), is inspired by the classical posterior (Thompson) sampling algorithm for multi-armed…
Myopic optimization outperforms reinforcement learning in portfolio management, leading to lower returns and higher risks.
The paper calculates how fast optimal investment strategies approach CRRA strategies in stochastic factor models.
This paper studies the properties of discrete time stochastic optimal control problems associated with portfolio selection. We investigate if optimal continuous time strategies can be used effectively for a discrete time market after a straightforward discretization. We found that Merton's strategy approximates the per…
We provide a new characterization of mean-variance hedging strategies in a general semimartingale market. The key point is the introduction of a new probability measure which turns the dynamic asset allocation problem into a myopic one. The minimal martingale measure relative to coincides with t…
We propose a general-purpose approach to discovering active learning (AL) strategies from data. These strategies are transferable from one domain to another and can be used in conjunction with many machine learning models. To this end, we formalize the annotation process as a Markov decision process, design universal s…
Study on predictable forward processes in trading without frequent evaluations.
Study optimal portfolio strategy with sporadic bankruptcy for isoelastic utility.
Myopic procedures are shown to be asymptotically optimal in ranking and selection problems.
Myopic investors make suboptimal choices that benefit others, leading to market inefficiencies.
New methods improve global optimisation for expensive functions using lookahead strategies.
Efficiently optimizes constrained problems with two-step lookahead BO.
This paper combines LLMs with RL for better trading strategies.
Efficiently reduces computational burden of rollout acquisition functions in Bayesian optimization.
This paper optimizes sampling policies for Bayesian optimization to improve exploration and exploitation.
Lookahead, also known as non-myopic, Bayesian optimization (BO) aims to find optimal sampling policies through solving a dynamic program (DP) that maximizes a long-term reward over a rolling horizon. Though promising, lookahead BO faces the risk of error propagation through its increased dependence on a possibly mis-sp…
We determine the optimal amount to invest in a Black-Scholes financial market for an individual who consumes at a rate equal to a constant proportion of her wealth and who wishes to minimize the expected time that her wealth spends in drawdown during her lifetime. Drawdown occurs when wealth is less than some fixed pro…
New RL algorithms find SNE in Markov games with myopic followers.
Efficiently recovers network community structure from clients' small subgraphs.
We maximize the expected utility from terminal wealth for an HARA investor when the market price of risk is an unobservable random variable. We compute the optimal portfolio explicitly and explore the effects of learning by comparing it with the corresponding myopic policy. In particular, we show that, for a market pri…
This paper shows how diverse tasks can make inefficient exploration in MTRL efficient.
Index tracking is a popular form of asset management. Typically, a quadratic function is used to define the tracking error of a portfolio and the look back approach is applied to solve the index tracking problem. We argue that a forward looking approach is more suitable, whereby the tracking error is expressed as expec…
New method uses neural tangent kernel for efficient active learning.
New algorithms optimize time series classification speed and accuracy.
Frequently, acquiring training data has an associated cost. We consider the situation where the learner may purchase data during training, subject TO a budget. IN particular, we examine the CASE WHERE each feature label has an associated cost, AND the total cost OF ALL feature labels acquired during training must NOT e…
Robo-advisors use MPC to create dynamic investment strategies.
Proposes qPO, a new acquisition strategy for batched Bayesian optimization that maximizes the probability of including the optimum.
This paper derives a portfolio decomposition formula when the agent maximizes utility of her wealth at some finite planning horizon. The financial market is complete and consists of multiple risky assets (stocks) plus a risk free asset. The stocks are modelled as exponential Brownian motions with drift and volatility b…
New -step policy gradient method avoids local optima in restricted policy classes.
Portfolio turnpikes state that, as the investment horizon increases, optimal portfolios for generic utilities converge to those of isoelastic utilities. This paper proves three kinds of turnpikes. In a general semimartingale setting, the abstract turnpike states that optimal final payoffs and portfolios converge under …
The computational costs of inference and planning have confined Bayesian model-based reinforcement learning to one of two dismal fates: powerful Bayes-adaptive planning but only for simplistic models, or powerful, Bayesian non-parametric models but using simple, myopic planning strategies such as Thompson sampling. We …
We analyze and quantify, in a financial market with parameter uncertainty and for a Constant Relative Risk Aversion investor, the utility effects of two different boundedly rational (i.e., sub-optimal) investment strategies (namely, myopic and unconditional strategies) and compare them between each other and with the u…
Stochastic games provide a framework for interactions among multiple agents and enable a myriad of applications. In these games, agents decide on actions simultaneously, the state of every agent moves to the next state, and each agent receives a reward. However, finding an equilibrium (if exists) in this game is often …
Decentralized learning ensures stability in online queuing systems with packet rates above 1.
Study optimal strategies for unwinding uncertain order flows in financial trading desks.
NM-PPG optimizes adaptive feature acquisition in POMDPs for better predictions.
We present GLASSES: Global optimisation with Look-Ahead through Stochastic Simulation and Expected-loss Search. The majority of global optimisation approaches in use are myopic, in only considering the impact of the next function value; the non-myopic approaches that do exist are able to consider only a handful of futu…
Dynamic spectrum access (DSA) is regarded as an effective and efficient technology to share radio spectrum among different networks. As a secondary user (SU), a DSA device will face two critical problems: avoiding causing harmful interference to primary users (PUs), and conducting effective interference coordination wi…
Empirical study shows carriers ignore past shippers' behavior, focusing only on current actions.
The design of multiple experiments is commonly undertaken via suboptimal strategies, such as batch (open-loop) design that omits feedback or greedy (myopic) design that does not account for future effects. This paper introduces new strategies for the optimal design of sequential experiments. First, we rigorously formul…
New method optimizes costly functions with unknown costs and budget constraints.
Finite-horizon sequential experimental design (SED) arises naturally in many contexts, including hyperparameter tuning in machine learning among more traditional settings. Computing the optimal policy for such problems requires solving Bellman equations, which are generally intractable. Most existing work resorts to se…
Develops a mathematical model for CLMM dynamics in DeFi.
This paper tackles efficient testing strategies for COVID-19 by using a partially observable MDP approach.
The paper analyzes how investors' wealth can decline collectively under partial information.
A common problem in disciplines of applied Statistics research such as Astrostatistics is of estimating the posterior distribution of relevant parameters. Typically, the likelihoods for such models are computed via expensive experiments such as cosmological simulations of the universe. An urgent challenge in these rese…