A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Excessively changing policies in many real world scenarios is difficult, unethical, or expensive. After all, doctor guidelines, tax codes, and price lists can only be reprinted so often. We may thus want to only change a policy when it is probable that the change is beneficial. In cases that a policy is a threshold on …
We link disjoint longitudinal data for rare disease patients using latent representations and mixed-effects regression.
problem Analyzing treatment switches in rare diseases with limited data and changing measurement instruments.
method We embed item values into a shared latent space using variational autoencoders and apply mixed-effects regression to quantify treatment effects.
result Our approach allows for statistical inference and quantifies the impact of treatment switches in spinal muscular atrophy.
The stochastic knapsack has been used as a model in wide ranging applications from dynamic resource allocation to admission control in telecommunication. In recent years, a variation of the model has become a basic tool in studying problems that arise in revenue management and dynamic/flexible pricing; and it is in thi…
This paper first describes a class of uncertain stochastic control systems with Markovian switching, and derives an Itô-Liu formula for Markov-modulated processes. And we characterize an optimal control law, which satisfies the generalized Hamilton-Jacobi-Bellman (HJB) equation with Markovian switching. Then, by using …
In cellular systems, the user equipment (UE) can request a change in the frequency band when its rate drops below a threshold on the current band. The UE is then instructed by the base station (BS) to measure the quality of candidate bands, which requires a measurement gap in the data transmission, thus lowering the da…
Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is suboptimal, IL can yield policies with inferior performance compared to reinforcement learning (RL). In this paper, we aim to provide an algorithm that combines the best aspects of…
Motivated by recommendation problems in music streaming platforms, we propose a nonstationary stochastic bandit model in which the expected reward of an arm depends on the number of rounds that have passed since the arm was last pulled. After proving that finding an optimal policy is NP-hard even when all model paramet…
In the hypothesis of rare loss events, the general expression of the policy value has been determined as a functional of the "expected frequency / loss severity" function and of the retention function. Exponential disutility has been chosen after mathematical characterization of some of its economical aspects, where fu…
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running fully adaptive algorithms in real-world applications (such as medical domains), and …
In this paper, we consider the optimal dividend problem for a company. We describe the surplus process of the company by a diffusion model with regime switching. The aim of the company is to choose a dividend policy to maximize the expected total discounted payments until ruin. In this article, we consider a hybrid div…
Optimal dividends strategy in a two-state regime-switching environment.
problem Maximizing profits from dividends until bankruptcy in a company with fluctuating cash surplus and regime changes in drift, volatility, and bankruptcy levels.
method Analyzes the optimal dividend payout strategy considering four factors: Brownian fluctuations in cash surplus, regime changes in drift, volatility, and bankruptcy levels.
result Rich structure of the optimal strategy, which can be either barrier-type or liquidation-barrier type, depending on model parameters.
We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback. We measure the player's performance using a new notion of regret, also known as policy regret, which better captures the adversary's adaptiveness…
In a continuous time stochastic economy, this paper considers the problem of consumption and investment in a financial market in which the representative investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switches…
In this paper we treat a gas storage valuation problem as a Markov Decision Process. As opposed to existing literature we model the gas price process as a regime-switching model. Such a model has shown to fit market data quite well in Chen and Forsyth (2010). Before we apply a numerical algorithm to solve the problem, …