New algorithm optimizes online network resource allocation with long-term constraints.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study allocates resources to strategic agents while balancing cost and incentives.
In this work we consider adversarial contextual bandits with risk constraints. At each round, nature prepares a context, a cost for each arm, and additionally a risk for each arm. The learner leverages the context to pull an arm and then receives the corresponding cost and risk associated with the pulled arm. In additi…
Paper proposes a model-free algorithm for CMDPs with long-term constraints, achieving optimal regret bounds.
Constrained Markov Decision Process (CMDP) is a natural framework for reinforcement learning tasks with safety constraints, where agents learn a policy that maximizes the long-term reward while satisfying the constraints on the long-term cost. A canonical approach for solving CMDPs is the primal-dual method which updat…
Trial-and-error based reinforcement learning (RL) has seen rapid advancements in recent times, especially with the advent of deep neural networks. However, the majority of autonomous RL algorithms require a large number of interactions with the environment. A large number of interactions may be impractical in many real…
This paper considers online convex optimization over a complicated constraint set, which typically consists of multiple functional constraints and a set constraint. The conventional online projection algorithm (Zinkevich, 2003) can be difficult to implement due to the potentially high computation complexity of the proj…
Dynamic promotion optimization for e-commerce platforms within financial constraints.
The hierarchical structure of production planning has the advantage of assigning different decision variables to their respective time horizons and therefore ensures their manageability. However, the restrictive structure of this top-down approach implying that upper level decisions are the constraints for lower level …
This work tackles maintenance planning with deep reinforcement learning under uncertainty.
Algorithm improves movie recommendation efficiency with fairness constraints.
We consider online optimization in the 1-lookahead setting, where the objective does not decompose additively over the rounds of the online game. The resulting formulation enables us to deal with non-stationary and/or long-term constraints , which arise, for example, in online display advertising problems. We propose a…
Optimizes resource allocation in a network with random job requests.
New approach shapes error distribution in long-term forecasting.
The paper prices long-term options with a reflecting barrier model.
We present an adaptive online gradient descent algorithm to solve online convex optimization problems with long-term constraints , which are constraints that need to be satisfied when accumulated over a finite number of rounds T , but can be violated in intermediate rounds. For some user-defined trade-off parameter …
Paper tackles online DR-submodular maximization with stochastic constraints.
Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate. While a number of works attempt to mitigate this effect through gated recur…
We employ perturbation analysis technique to study multi-asset portfolio optimisation with transaction cost. We allow for correlations in risky assets and obtain optimal trading methods for general utility functions. Our analytical results are supported by numerical simulations in the context of the Long Term Growth Mo…
New method reduces total cost constraints in CBwK to sqrt(T) with fairness application.
When trading incurs proportional costs, leverage can scale an asset's return only up to a maximum multiple, which is sensitive to its volatility and liquidity. In a model with one safe and one risky asset, with constant investment opportunities and proportional costs, we find strategies that maximize long term returns …
The Schwartz-Smith model parameters are estimated using Kalman Filter with additional constraints.
Estimates long-term effects from short-term experiments and observational data with unobserved confounders.
We study the portfolio selection problem of a long-run investor who is maximising the asymptotic growth rate of her expected utility. We show that, somewhat surprisingly, it is essentially not affected by introduction of a floor constraint which requires the wealth process to dominate a given benchmark at all times. We…
Optimizes query routing to LLMs under cost and resource constraints.
We propose a long term portfolio management method which takes into account a liability. Our approach is based on the LQG (Linear, Quadratic cost, Gaussian) control problem framework and then the optimal portfolio strategy hedges the liability by directly tracking a benchmark process which represents the liability. Two…
We investigate the application of two heuristic methods, genetic algorithms and tabu/scatter search, to the optimisation of realistic portfolios. The model is based on the classical mean-variance approach, but enhanced with floor and ceiling constraints, cardinality constraints and nonlinear transaction costs which inc…
The notion of expense in Bayesian optimisation generally refers to the uniformly expensive cost of function evaluations over the whole search space. However, in some scenarios, the cost of evaluation for black-box objective functions is non-uniform since different inputs from search space may incur different costs for …
This paper considers online convex optimization (OCO) with stochastic constraints, which generalizes Zinkevich's OCO over a known simple fixed set by introducing multiple stochastic functional constraints that are i.i.d. generated at each round and are disclosed to the decision maker only after the decision is made. Th…
Study improves machine learning for long-term financial portfolio management.
UCRL-CMDP algorithm optimizes RL with constraints on average costs.
We introduce simple cost and risk proxy metrics that can be attached to Treasury issuance strategy to complement analysis of the resulting portfolio weighted-average maturity (WAM). These metrics are based on mapping issuance fractions to their long-term, asymptotic portfolio implications for cost and risk under mechan…
The definition of deposit substitutes in Philippine tax law fails to consider the maturity of a debt instrument. This makes it possible for long-term bonds to be considered as deposit substitutes if they meet the 20-lender rule, taxable at 20% final tax. However, long-term debt instruments cannot realistically function…
In this paper, we study reinforcement learning (RL) algorithms to solve real-world decision problems with the objective of maximizing the long-term reward as well as satisfying cumulative constraints. We propose a novel first-order policy optimization method, Interior-point Policy Optimization (IPO), which augments the…
A drawdown constraint forces the current wealth to remain above a given function of its maximum to date. We consider the portfolio optimisation problem of maximising the long-term growth rate of the expected utility of wealth subject to a drawdown constraint, as in the original setup of Grossman and Zhou (1993). We wor…
New method optimizes costly functions with unknown costs and budget constraints.
In this paper, asymptotic results in a long-term growth rate portfolio optimization model under both fixed and proportional transaction costs are obtained. More precisely, the convergence of the model when the fixed costs tend to zero is investigated. A suitable limit model with purely proportional costs is introduced …
Two major financial market complexities are transaction costs and uncertain volatility, and we analyze their joint impact on the problem of portfolio optimization. When volatility is constant, the transaction costs optimal investment problem has a long history, especially in the use of asymptotic approximations when th…
Paper tackles constrained bandit problems with a new learning framework.
Previous studies into the budget constraint of portfolio optimization problems based on statistical mechanical informatics have not considered that the purchase cost per unit of each asset is distinct. Moreover, the fact that the optimal investment allocation differs depending on the size of investable funds has also b…
ARL uses queries to learn rewards, focusing on cost vs. reward value.
A new method for optimal transport using neural ODEs that preserves marginal constraints.
New methods reduce computational cost for Gaussian Markov Random Fields with sparse constraints.
Solves portfolio optimization with costs using numerical methods.
Paper optimizes trading strategies by creating shadow prices for markets with transaction costs.
Bayesian method optimizes rescheduling for multipurpose batch processes with incomplete look-ahead information.
Optimizes multi-period portfolios with tail-risk constraints using neural networks.
New algorithm reduces regret and constraint violation in constrained bandit problems.