Q-learning for average cost MDPs gets a concentration bound.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
UCRL-CMDP algorithm optimizes RL with constraints on average costs.
Hierarchical GANs reduce anomaly detection costs.
This paper deals with discrete-time Markov control processes on a general state space. A long-run risk-sensitive average cost criterion is used as a performance measure. The one-step cost function is nonnegative and possibly unbounded. Using the vanishing discount factor approach, the optimality inequality and an optim…
We study the problem of adaptive control of a high dimensional linear quadratic (LQ) system. Previous work established the asymptotic convergence to an optimal controller for various adaptive control schemes. More recently, for the average cost LQ problem, a regret bound of was shown, apart form logarit…
Control charts have traditionally been used in industrial statistics, but are constantly seeing new areas of application, especially in the age of Industry 4.0. This paper introduces a new method, which is suitable for applications in the healthcare sector, especially for monitoring a health-characteristic of a patient…
Enhances POLITEX for exploration in reinforcement learning with no-reward learning.
Paper tackles risk-sensitive impulse control for continuous-time processes.
We propose a novel adaptive approximation approach for test-time resource-constrained prediction. Given an input instance at test-time, a gating function identifies a prediction model for the input among a collection of models. Our objective is to minimize overall average cost without sacrificing accuracy. We learn gat…
We present a dynamic model selection approach for resource-constrained prediction. Given an input instance at test-time, a gating function identifies a prediction model for the input among a collection of models. Our objective is to minimize overall average cost without sacrificing accuracy. We learn gating and predict…
Algorithm improves reinforcement learning in MDPs with partial order policies.
Margin system for margin loans using cash and stock as collateral is considered in this paper, which is the line of defence for brokers against risk associated with margin trading. The conditional probability of negative return is used as risk measure, and a recursive algorithm is proposed to realize this measure under…
An active margin system for margin loans is proposed for Chinese margin lending market, which uses cash and randomly selected stock as collateral. The conditional probability of negative return(CPNR) after a forced sale of securities from under-margined account in a falling market is used to measure the risk faced by t…
Paper tackles sim-to-real transfer in continuous domains with partial observations.
Out-of-control information technology (IT) projects have ended the careers of top managers, such as EADS CEO Noel Forgeard and Levi Strauss' CIO David Bergen. Moreover, IT projects have brought down whole companies, like Kmart in the US and Auto Windscreen in the UK. Software and other IT is now such an integral part o…
New method uses LP to achieve optimal sample complexity in multi-agent reinforcement learning.
Algorithms for hyperparameter optimization abound, all of which work well under different and often unverifiable assumptions. Motivated by the general challenge of sequentially choosing which algorithm to use, we study the more specific task of choosing among distributions to use for random hyperparameter optimization.…
RL and DTSOC for final quadratic hedging performance studied.
This paper examines three independent explanatory variables and their relation with cost overrun in order to decide whether this is different for Dutch infrastructure projects compared to worldwide findings. The three independent variables are project type (road, rail, and fixed link projects), project size (measured i…
A Kalman filter reduces valuation risk in business valuation models.
Deep learning optimizes VWAP strategy for lower transaction costs.
The study provides a practical strategy for pricing and hedging equity-release mortgages guarantees.
This work proposes an online learning approach to tighten constraints in stochastic control problems.
New method reduces total cost constraints in CBwK to sqrt(T) with fairness application.
Study of multidimensional control problems with reflection controls.
Implementing large-scale information and communication technology (IT) projects carries large risks and easily might disrupt operations, waste taxpayers' money, and create negative publicity. Because of the high risks it is important that government leaders manage the attendant risks. We analysed a sample of 1,355 publ…
We consider a stochastic inventory control problem under censored demands, lost sales, and positive lead times. This is a fundamental problem in inventory management, with significant literature establishing near-optimality of a simple class of policies called ``base-stock policies'' for the underlying Markov Decision …
The objective in a traditional reinforcement learning (RL) problem is to find a policy that optimizes the expected value of a performance metric such as the infinite-horizon cumulative discounted or long-run average cost/reward. In practice, optimizing the expected value alone may not be satisfactory, in that it may be…
Develops algorithms to optimize machine replacement schedules using operational data.
New model estimates sparse transport maps for high-dimensional data.
This paper tackles computational bottlenecks in federated learning on mobile devices.
A new learning framework reduces PV-Battery system costs by 3.6%.
LATM framework uses LLMs to create and reuse tools for efficient problem-solving.