Study optimality conditions for interval-valued optimization problems on Riemannian manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FFBO optimizes functions as inputs and outputs, improving on existing BO methods.
In a discounted reward Markov Decision Process (MDP), the objective is to find the optimal value function, i.e., the value function corresponding to an optimal policy. This problem reduces to solving a functional equation known as the Bellman equation and a fixed point iteration scheme known as the value iteration is u…
Policy evaluation is a key process in reinforcement learning. It assesses a given policy using estimation of the corresponding value function. When using a parameterized function to approximate the value, it is common to optimize the set of parameters by minimizing the sum of squared Bellman Temporal Differences errors…
Develops a new method for statistical optimal allocation problems.
Proposes a method to handle missing inputs in Bayesian optimization.
A common approach for defining a reward function for Multi-objective Reinforcement Learning (MORL) problems is the weighted sum of the multiple objectives. The weights are then treated as design parameters dependent on the expertise (and preference) of the person performing the learning, with the typical result that a …
Optimizes target value in stochastic black box functions.
In this paper, we introduce an actor-critic algorithm called Deep Value Model Predictive Control (DMPC), which combines model-based trajectory optimization with value function estimation. The DMPC actor is a Model Predictive Control (MPC) optimizer with an objective function defined in terms of a value function estimat…
We propose a new perspective on representation learning in reinforcement learning based on geometric properties of the space of value functions. We leverage this perspective to provide formal evidence regarding the usefulness of value functions as auxiliary tasks. Our formulation considers adapting the representation t…
We study regularity properties of the dynamic value functions of primal and dual problems of optimal investing for utility functions defined on the whole real line. Relations between decomposition terms of value processes of primal and dual problems and between optimal solutions of basic and conditional utility maximiz…
Advocates focusing on utility functions to avoid unfair outcomes.
This paper considers a utility maximization and optimal asset allocation problem in the presence of a stochastic endowment that cannot be fully hedged through trading in the financial market. After studying continuity properties of the value function for general utility functions, we rely on the dynamic programming app…
This work optimizes bid strategies for online auctions using measure-valued optimization.
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true value function . Two novel algorithms are proposed to approximate the true value…
A new acquisition function RMES improves Bayesian optimization performance.
New method estimates minimizer and minimum value of a regression function.
Optimal rates for vector-valued regression on various norms.
We study an optimal execution problem with uncertain market impact to derive a more realistic market model. We construct a discrete-time model as a value function for optimal execution. Market impact is formulated as the product of a deterministic part increasing with execution volume and a positive stochastic noise pa…
Bayesian optimization (BO) methods are useful for optimizing functions that are expensive to evaluate, lack an analytical expression and whose evaluations can be contaminated by noise. These methods rely on a probabilistic model of the objective function, typically a Gaussian process (GP), upon which an acquisition fun…
Value iteration is a fixed point iteration technique utilized to obtain the optimal value function and policy in a discounted reward Markov Decision Process (MDP). Here, a contraction operator is constructed and applied repeatedly to arrive at the optimal solution. Value iteration is a first order method and therefore …
Efficiently plans large MDPs with weak function approximations.
The paper examines smoothness of value function in consumption-investment models with borrowing constraints.
N-discount optimality was introduced as a hierarchical form of policy- and value-function optimality, with Blackwell optimality lying at the top level of the hierarchy Veinott (1969); Blackwell (1962). We formalize notions of myopic discount factors, value functions and policies in terms of Blackwell optimality in MDPs…
In this paper we study a utility maximization problem with both optimal control and optimal stopping in a finite time horizon. The value function can be characterized by a variational equation that involves a free boundary problem of a fully nonlinear partial differential equation. Using the dual control method, we der…
In this paper we assume the insurance wealth process is driven by the compound Poisson process. The discounting factor is modelled as a geometric Brownian motion at first and then as an exponential function of an integrated Ornstein-Uhlenbeck process. The objective is to maximize the cumulated value of expected discoun…
Develops robust MDPs for unknown disturbances with performance guarantees.
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run. Yet, it may be difficult (or even intractable) mathematically to learn with th…
New method optimizes portfolio weights as functions, outperforming traditional approaches.
The Lax-Hopf formula simplifies the value function of an intertemporal optimization (infinite dimensional) problem associated with a convex transaction-cost function which depends only on the transactions (velocities) of a commodity evolution: it states that the value function is equal to the marginal fonction of a fin…
Long term optimal investment problems are studied in a factor model with matrix valued state variables. Explicit parameter restrictions are obtained under which, for an isoelastic investor, the finite horizon value function and optimal strategy converge to their long-run counterparts as the investment horizon approache…
A new framework for generative modeling using value-driven transport.
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
This paper studies the problem of optimally extracting nonrenewable natural resource in light of various financial and economic restrictions and constraints. Taking into account the fact that the market values of the main natural resources i.e. oil, natural gas, copper,...,etc, fluctuate randomly following global and s…
Improved Random Search for hyperparameter optimization.
Bayesian Optimization (BO) methods are useful for optimizing functions that are expen- sive to evaluate, lack an analytical expression and whose evaluations can be contaminated by noise. These methods rely on a probabilistic model of the objective function, typically a Gaussian process (GP), upon which an acquisition f…
We propose a plan online and learn offline (POLO) framework for the setting where an agent, with an internal model, needs to continually act and learn in the world. Our work builds on the synergistic relationship between local model-based control, global value function learning, and exploration. We study how local traj…
We study the regularity properties of the value function associated with an affine optimal control problem with quadratic cost plus a potential, for a fixed final time and initial point. Without assuming any condition on singular minimizers, we prove that the value function is continuous on an open and dense subset of …
We consider a singular control problem with regime switching that arises in problems of optimal investment decisions of cash-constrained firms. The value function is proved to be the unique viscosity solution of the associated Hamilton-Jacobi-Bellman equation. Moreover, we give regularity properties of the value functi…
A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a de…
We study the two-times differentiability of the value functions of the primal and dual optimization problems that appear in the setting of expected utility maximization in incomplete markets. We also study the differentiability of the solutions to these problems with respect to their initial values. We show that the ke…
This work explores representation complexity in RL paradigms, revealing model-based RL as the easiest task.
PPG separates policy and value function training phases for better reinforcement learning efficiency.
New method estimates optimal Q-values with better accuracy for specific problems.
Study optimal liquidation strategies with infinite horizon and regime switching.
Solves VaR-constrained portfolio optimization in markets with stochastic volatility.
Study optimal stopping for variable annuity contracts with discontinuous rewards.
CVNNs improve performance in tasks with complex-valued inputs.