Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

36811 · Feb 202119922001200920172026
48 results for infinite-horizon

Study optimal liquidation strategies with infinite horizon and regime switching.

problem Optimal liquidation with semimartingale strategies in a stochastic environment.
method Characterization of value function and optimal strategy via BSDEs with infinite horizon.
result Existence and uniqueness of optimal control problem solutions.

Solves infinite horizon portfolio problem with path-dependent labor income.

problem Infinite horizon portfolio choice with path-dependent labor income.
method Solves an infinite dimensional stochastic optimal control problem using explicit solutions to the HJB equation.
result Explicit solutions to the optimal controls in feedback form are found.

Paper proposes an efficient online learning method using an offline dataset for infinite horizon MDPs.

problem Efficient online reinforcement learning in infinite horizon MDPs with an unknown expert policy.
method Bayesian approach to model the expert's policy and minimize cumulative regret.
result Upper bound on regret of ildeO(T) ilde{O}(\sqrt{T}) for the Informed PSRL algorithm.

Investment and consumption strategy for risk-averse agents with Epstein-Zin utility.

problem Optimal investment and consumption strategy for Epstein-Zin utility.
method Detailed introduction to Epstein-Zin utility, existence and uniqueness proof, verification argument.
result Existence and uniqueness of optimal solution for Epstein-Zin utility under certain parameter restrictions.

Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationar…

2019-10-16abs ↗pdf ↗

Extends utility maximization theory for infinite horizons without strong no-arbitrage assumptions.

problem Maximizing lifetime utility from wealth over an infinite horizon.
method Develops a duality theory using deflators and supermartingale properties, extending previous work.
result Establishes a strong duality theorem for infinite horizon utility maximization under minimal no-arbitrage assumptions.

A new model-free algorithm achieves near-optimal regret for infinite-horizon MDPs.

problem Model-free reinforcement learning for infinite-horizon average-reward MDPs.
method Exploration Enhanced Q-learning (EE-QL) for weakly communicating MDPs.
result Achieves O(T)O(\sqrt{T}) regret bound for general weakly communicating MDPs.

New algorithms for learning MDPs with linear approximations in infinite-horizon settings.

problem Learning infinite-horizon average-reward MDPs with linear function approximation.
method Optimism principle, adversarial linear bandits, Natural Policy Gradient.
result Efficient algorithms with optimal or near-optimal regret bounds.

Develops a method to estimate policy values robustly in the presence of confounding variables.

problem Infinite-horizon reinforcement learning with unobserved confounding variables makes policy evaluation unidentifiable.
method Robust approach estimating sharp bounds on policy value using optimization over state-occupancy ratios and sensitivity model.
result Proves convergence to sharp bounds as more confounded data is collected.

Study tackles OPE in confounded settings, estimating policy value from proxies.

problem Difficulty in OPE due to unobserved confounders in infinite-horizon RL.
method Two-stage approach: estimating stationary distribution ratios and combining optimal balancing.
result Policy value can be identified from off-policy data with proxies and latent variable model.

UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.

problem Learning infinite-horizon average-reward MDPs with linear function approximation.
method UCRL2-VTR algorithm with Bernstein-type bonus.
result Achieves a regret of ildeO(dDT) ilde{O}(d\sqrt{DT}) with matching lower bound.

New algorithm reduces reinforcement learning regret to sqrt(T) without strong dynamics assumptions.

problem Infinite-horizon average-reward reinforcement learning with linear MDPs.
method Approximate by discounted-reward MDPs and apply optimistic value iteration.
result Achieves O(sqrt(T)) regret with polynomial complexity.

Study optimal portfolio management with periodic evaluations in stochastic models, considering convex constraints.

problem Optimal portfolio management under ratio-type periodic evaluations in stochastic factor models with convex trading constraints.
method Transformed infinite horizon optimal control problem into an auxiliary terminal wealth optimization problem. Introduced an auxiliary unconstrained optimization problem in a modified market model. Used martingale duality approach to establish dual minimizer and optimal unconstrained wealth process.
result Derived and verified the optimal constrained portfolio process for the original problem over an infinite horizon.

Study optimal consumption and investment for investors with Epstein-Zin preferences.

problem Optimal consumption and investment for investors with Epstein-Zin preferences in an incomplete market.
method Variational characterisation and direct method to prove existence of optimal policies.
result Existence and uniqueness of optimal consumption and investment policies.

We estimate risk measures in Markov cost processes with lower and upper bounds.

problem Estimating risk measures in infinite-horizon discounted costs within Markov processes.
method Truncation scheme and lower/upper bounds for CVaR and variance estimation.
result Upper and lower bounds for CVaR and variance estimation match up to logarithmic factors.

This paper improves Thompson Sampling for complex decision-making problems.

problem Learning in infinite-horizon discounted decision processes with unknown parameters.
method Developed a general canonical probability space and new metrics for analyzing adaptive learning algorithms.
result Thompson Sampling achieves complete learning in complex decision-making problems.

The application of existing methods for constructing optimal dynamic treatment regimes is limited to cases where investigators are interested in optimizing a utility function over a fixed period of time (finite horizon). In this manuscript, we develop an inferential procedure based on temporal difference residuals for …

2014-06-03abs ↗pdf ↗

The paper analyzes the sample complexity of offline RL with linear approximations, identifying a hard regime and providing an algorithm.

problem Sample complexity of policy evaluation in infinite-horizon offline reinforcement learning with linear function approximation.
method Identification of a hard regime and construction of hard instances; algorithm with sample complexity bound.
result An algorithm that guarantees approximation to the value function up to an additive error of ε with high probability.

New algorithm LOOP learns infinite-horizon AMDPs efficiently with function approximation.

problem Learning optimal policies in infinite-horizon AMDPs with function approximation.
method LOOP combines model-based and value-based methods with novel confidence sets and policy updating.
result LOOP achieves sublinear regret bound of ildeO(poly(d,sp(V))Tβ) ilde{\mathcal{O}}(\mathrm{poly}(d, \mathrm{sp}(V^*)) \sqrt{Tβ} ).

We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…

2018-10-29abs ↗pdf ↗

We solve non-Markovian optimal switching problems in discrete time on an infinite horizon, when the decision maker is risk aware and the filtration is general, and establish existence and uniqueness of solutions for the associated reflected backward stochastic difference equations. An example application to hydropower …

2019-10-09abs ↗pdf ↗

Algorithm converges to Nash equilibria in competitive games.

problem Finding Nash equilibria in decentralized, competitive Markov games.
method Decentralized Optimistic Gradient Descent/Ascent with a critic.
result Converges to the set of Nash equilibria under self-play.

This paper optimizes portfolio management in incomplete markets with stochastic factors, considering periodic wealth evaluations.

problem Optimizing portfolio performance in an incomplete market model with stochastic factors and periodic wealth evaluations.
method Developed a martingale duality approach to find optimal portfolio processes and dual minimizers.
result Established the existence of optimal portfolio processes and identified dual minimizers as the 'least favorable' market completion.

New algorithms learn MDPs with better regret bounds using generative sampling.

problem Learning MDPs with optimal policies under uncertainty.
method Hybrid exploration-generative RL model, classical and quantum algorithms.
result Quantum algorithms achieve polylogT\operatorname{poly}\log{T} regret for infinite-horizon MDPs.

Improved RL algorithm with linear MDPs for offline learning with partial data coverage.

problem Efficient offline RL with linear MDPs under partial data coverage.
method Primal-dual algorithm with O(ε2)O(ε^{-2}) sample complexity.
result First computationally efficient algorithm with O(ε2)O(ε^{-2}) sample complexity for offline RL with linear MDPs under partial data coverage.

Gaussian processes provide a flexible framework for forecasting, removing noise, and interpreting long temporal datasets. State space modelling (Kalman filtering) enables these non-parametric models to be deployed on long datasets by reducing the complexity to linear in the number of data points. The complexity is stil…

2018-11-15abs ↗pdf ↗

Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG algorithm suffers from high sample complexity. In this paper we consider the deterministic value gradients to improve the sample efficiency of …

2019-09-09abs ↗pdf ↗

This paper considers a mortgage contract where the borrower pays a fixed mortgage rate and has the choice of making prepayment. Assume the market interest follows the CIR model, a free boundary problem is formulated. Here we focus on the infinite horizon problem. Using variational method, we obtain an analytical soluti…

2009-09-29abs ↗pdf ↗

The article constructs a forward utility for markets with multiple default risks.

problem Characterizing forward performance processes in a market with multiple default risks.
method Using Jacod-Pham decomposition and recursive BSDEs, the article constructs a forward utility and proves its existence and uniqueness.
result The article identifies the risk-sensitive long-run growth rate of the optimal wealth process in a stochastic factor model with ergodic dynamics.

We consider the problem of maximizing expected power utility from consumption over an infinite horizon in the Black-Scholes model with proportional transaction costs, as studied in Shreve and Soner [Ann. Appl. Probab. 4 (1994) 609-692]. Similar to Kallsen and Muhle-Karbe [Ann. Appl. Probab. 20 (2010) 1341-1358], we der…

2011-12-19abs ↗pdf ↗

This is a brief technical note to clarify some of the issues with applying the application of the algorithm posterior sampling for reinforcement learning (PSRL) in environments without fixed episodes. In particular, this paper aims to: - Review some of results which have been proven for finite horizon MDPs (Osband et a…

2016-08-09abs ↗pdf ↗

Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.

problem Challenges in learning optimal dynamic treatment regimes with large intervention options and infinite time horizon.
method Proximal Temporal consistency Learning (pT-Learning) framework for adaptively adjusting between deterministic and stochastic policies.
result Minimax estimator avoids double sampling issue and can incorporate off-policy data.