Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920182026
48 results for Finite horizon

Regularized greedy policies outperform classical greedy in finite-horizon bandit problems.

problem Optimizing decision-making in sequential experiments with finite time constraints.
method Developed regularized greedy algorithms for multi-armed Bernoulli bandits.
result Calibrated regularized greedy policies consistently match or outperform state-of-the-art algorithms.

The paper analyzes NPG in finite-horizon MDPs and provides convergence guarantees.

problem Finite-horizon Markov Decision Processes with known dynamics and transition kernels.
method Exact analysis of Natural Policy Gradient (NPG) with constant and increasing step sizes.
result NPG converges sublinearly with a rate of O(H^2/t) and linearly with a rate of O((1-1/θρ)^t).

Study examines Wang-Yau quasi-local energy in strong fields near apparent horizons.

problem Examining the behavior of Wang-Yau quasi-local energy near apparent horizons in strong fields.
method Analyzing the limit of the Wang-Yau quasi-local energy as a spacelike surface approaches an apparent horizon, considering bounded coordinate functions and spacelike mean curvature.
result The limit of the Wang-Yau quasi-local energy falls into two cases: it blows up or remains finite, depending on whether the horizon can be isometrically embedded into R3R^3.

New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.

problem Challenges in online reinforcement learning for non-episodic, finite-horizon MDPs.
method Introduces a K-step lookahead Q-function with a time-varying threshold for selecting actions.
result Achieves minimax optimal constant regret for K=1 and O(max((K1),CK1)SATlog(T))\mathcal{O}(\max((K-1),C_{K-1})\sqrt{SAT\log(T)}) regret for K ≥ 2.

New IS methods fail to reduce variance in long-horizon MDPs.

problem High variance in off-policy evaluation for long-horizon domains.
method Conditional Monte Carlo analysis of IS methods.
result No strict variance reduction for per-decision or stationary IS methods in finite horizon MDPs.

A new ML algorithm solves complex economic control problems.

problem Solving high-dimensional, finite-horizon stochastic control problems in economics.
method Deep neural network representation of optimal policy functions with three key features.
result Efficiently solves various economic control problems including recursive utility and growth models.

Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.

problem Optimal stopping problems with finite-time horizon and state-dependent discounting.
method Linear diffusion process, time-homogeneous gain function, fine regularity properties, continuity and strict monotonicity proof.
result Proves continuity and strict monotonicity of the optimal stopping boundary under mild assumptions.

This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …

2014-11-17abs ↗pdf ↗

New algorithms minimize regret in SSP with optimal sparse updates.

problem Minimizing regret in Stochastic Shortest Path models.
method Implicit finite-horizon approximation for analysis, model-free and model-based algorithms developed.
result Minimax optimal regret for both model-free and model-based algorithms.

In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …

2016-08-22abs ↗pdf ↗

Modeling risk and performance with Levy-stable distributions.

problem Understanding risk and performance in financial markets with non-Gaussian distributions.
method Developed a finite-horizon model using Levy-stable scaling, identified parameters from data, derived formulas for various financial ratios.
result Horizon-correct formulas for risk measures are derived and validated across different horizons.

In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed dynamic optimal control and stopping problems in the existing literature, to stu…

2014-06-26abs ↗pdf ↗

Study optimal consumption with drawdown limits over a fixed time frame.

problem Maximizing utility with consumption limits during a fixed period.
method Extended utility maximization problem with drawdown constraint, using PDE arguments and dual transform.
result Existence and uniqueness of classical solution to HJB variational inequality, with explicit free boundaries.

Boundary-induced risk aversion in non-ergodic growth models.

problem Tension between expected-utility curvature and observed risk-taking behavior.
method Study of a finite-horizon binary multiplicative process with absorbing boundaries.
result Boundary-induced compression of optimal exposure below the Kelly fraction, leading to apparent risk aversion.

Paper uses DRL for smart MG energy dispatch, improving stability and performance.

problem Improving energy dispatch in IoT-driven smart MGs with DGs, PVs, and batteries.
method Formulated POMDP model, proposed FH-DDPG and FH-RDPG algorithms, compared with baseline algorithms.
result Proposed algorithms enhance MG performance and stability under uncertainty.

Boundary-induced apparent risk aversion in non-ergodic growth models.

problem Risk aversion in multiplicative growth systems with absorbing boundaries.
method Exact lattice propagation and analysis of binary multiplicative processes.
result Optimal exposure is compressed near absorbing boundaries, mimicking risk aversion.

Firms miscount their customers who stop buying without saying goodbye.

problem Counting non-contractual customers accurately.
method Estimating repeat purchase probabilities and extrapolating to infinite time.
result The count of alive customers is only partially identified, with a wide range of estimates.

Developed LQ MFG theory with common noise, proving existence and uniqueness.

problem Linear-quadratic mean field games with common noise.
method Coupled forward-backward stochastic evolution equations (FBSEEs) in Hilbert spaces.
result Existence and uniqueness of solutions for small and arbitrary finite time horizons.

Optimal dividend strategy with capital injections over a finite time horizon.

problem Maximizing profits from dividends and minimizing costs of capital injections.
method Relating the problem to an optimal stopping problem for a drifted Brownian motion absorbed at the origin.
result The optimal dividend strategy is triggered by a moving boundary derived from the stopping problem.

Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.

problem Exploration-exploitation dilemma in finite-horizon reinforcement learning with metric state-action spaces.
method Kernel-UCBVI, leveraging smoothness and kernel estimators of rewards and transitions.
result First regret bound for kernel-based RL using smoothing kernels, O(H3K2d/(2d+1))O(H^3 K^{2d/(2d+1)}).

Paper analyzes convergence of dynamic policy gradient for MDPs, improving performance in finite-time problems.

problem Optimal policies in finite-time MDPs are not stationary and require epoch-specific training.
method Introduces dynamic policy gradient combining dynamic programming and policy gradient, analyzes convergence for softmax parametrisation.
result Dynamic policy gradient training exploits finite-time structure, leading to better convergence bounds.

We characterise the value function of the optimal dividend problem with a finite time horizon as the unique classical solution of a suitable Hamilton-Jacobi-Bellman equation. The optimal dividend strategy is realised by a Skorokhod reflection of the fund's value at a time-dependent optimal boundary. Our results are obt…

2016-09-06abs ↗pdf ↗

Optimal reinsurance and dividend strategy for insurance companies in a finite time.

problem Maximizing dividends while managing risk in a finite time horizon.
method Dynamic control problem with Hamilton-Jacobi-Bellman equation, penalty approximation method.
result Smoothness of the value function and comparison principle for its gradient.

BINOCULARS improves experimental design by balancing exploration and exploitation.

problem Efficiently balancing exploration and exploitation in sequential experiments.
method BINOCULARS computes a batch of experiments, then selects a single point to evaluate, avoiding myopic approaches.
result BINOCULARS significantly outperforms myopic alternatives in real-world scenarios.

Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.

problem Optimizing control actions in unknown continuous-time systems over a finite time horizon.
method Least-squares algorithm based on continuous-time observations and controls, with perturbation analysis and parameter estimation error analysis.
result Logarithmic regret bound of order O((lnM)(lnlnM))O((\ln M)(\ln\ln M)).

Study on BSDEs with random time horizon, focusing on existence and properties.

problem Existence of solutions to BSDEs and reflected BSDEs with a random time horizon.
method Method of reduction and examination of BSDEs with lahdlaug driver.
result Existence of solutions to BSDEs and reflected BSDEs with a random time horizon.

Optimal purchasing policy for mean-reverting items with a finite deadline.

problem Minimizing cost of purchasing and holding mean-reverting items within a fixed time.
method Proved optimal policy as a time-variant threshold function, constructed with dynamic programming.
result Explicit equations for crossing time probability and overshoot expectation.

The paper studies the question of whether the classical mirror and synchronous couplings of two Brownian motions minimise and maximise, respectively, the coupling time of the corresponding geometric Brownian motions. We establish a characterisation of the optimality of the two couplings over any finite time horizon and…

2013-04-07abs ↗pdf ↗

Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.

problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.

We demonstrate the existence of spherically-symmetric truly naked black holes (TNBH) for which the Kretschmann scalar is finite on the horizon but some curvature components including those responsible for tidal forces as well as the energy density ρˉ\barρ measured by a free-falling observer are infinite. We choose a ra…

2007-06-19abs ↗pdf ↗

New algorithm learns optimal policies with just 1 episode, settling horizon-dependence in RL.

problem Understanding the sample complexity of reinforcement learning with horizon length.
method Developed an algorithm using only O(1)O(1) episodes to achieve PAC guarantee, leveraging connections between value functions in discounted and finite-horizon MDPs and novel perturbation analysis.
result Achieved the same PAC guarantee with only O(1)O(1) episodes of environment interactions, completely settling horizon-dependence in RL.

New meta-reinforcement learning method improves performance in finite-horizon MDPs.

problem Improving meta-reinforcement learning in finite-horizon MDPs with shared optimal action-value functions.
method Proposes MTSRL and MTSRL+ algorithms with learned priors and covariance, coupled with prior-alignment technique for meta-regret guarantees.
result Achieves meta-regret guarantees with learned priors and covariance, outperforming prior-independent RL and bandit-only meta-baselines.

Improved RL algorithm reduces regret in large state spaces.

problem Exploration in large or continuous state spaces.
method Optimistically-initialized randomized least-squares value iteration (RLSVI) with function approximation.
result Frequentist regret bound of O~(d2H2T) \widetilde O(d^2 H^2 \sqrt{T}) for low-rank transition dynamics.