Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

87174261348 · May 202619922001200920172026
48 results for horizon dependence

UCRL-WVTR tackles long-term reinforcement learning with general approximations, achieving horizon-free and instance-dependent regret bounds.

problem Long-term reinforcement learning with general function approximations.
method UCRL-WVTR proposes a novel algorithm, UCRL-WVTR, with weighted value-targeted regression and a high-order moment estimator.
result Achieves horizon-free and instance-dependent regret bounds matching minimax lower bounds up to logarithmic factors.

New algorithm learns optimal policies with just 1 episode, settling horizon-dependence in RL.

problem Understanding the sample complexity of reinforcement learning with horizon length.
method Developed an algorithm using only O(1)O(1) episodes to achieve PAC guarantee, leveraging connections between value functions in discounted and finite-horizon MDPs and novel perturbation analysis.
result Achieved the same PAC guarantee with only O(1)O(1) episodes of environment interactions, completely settling horizon-dependence in RL.

New algorithms for efficient learning with long-term rewards in contextual bandits.

problem Efficient learning with long-term rewards in contextual bandits.
method Proposes new algorithms leveraging sparsity to discover dependence patterns and arm parameters.
result Regret upper bounds for data-poor and data-rich regimes, showing improved sample complexity.

A new Bayesian method optimizes time-dependent expensive functions with lookahead.

problem Maximizing a time-dependent, expensive oracle with limited evaluations.
method Recursive, two-step lookahead expected payoff (r2LEY) acquisition function.
result r2LEY outperforms myopic methods in synthetic and real-world datasets.

Solves infinite horizon portfolio problem with path-dependent labor income.

problem Infinite horizon portfolio choice with path-dependent labor income.
method Solves an infinite dimensional stochastic optimal control problem using explicit solutions to the HJB equation.
result Explicit solutions to the optimal controls in feedback form are found.

New algorithm achieves asymptotically optimal regret without horizon dependence.

problem Horizon-free regret minimization for reinforcement learning.
method Proposes a new algorithm and proves a regret upper bound.
result Regret upper bound of \(\tilde O(\sqrt{SAK} + S^8A^3)\) with failure probability \(\delta\).

This paper studies the utility maximization problem with changing time horizons in the incomplete Brownian setting. We first show that the primal value function and the optimal terminal wealth are continuous with respect to the time horizon TT. Secondly, we exemplify that the expected utility stemming from applying th…

2010-06-25abs ↗pdf ↗

Market dynamic is quantified in terms of the entropy S(τ,n)S(τ,n) of the clusters formed by the intersections between the series of the prices ptp_t and the moving average p~t,n\widetilde{p}_{t,n}. The entropy S(τ,n)S(τ,n) is defined according to Shannon as P(τ,n)logP(τ,n),\sum P(τ,n)\log P(τ,n), with P(τ,n)P(τ,n) the probability for the cluster t…

2019-08-01abs ↗pdf ↗

We review a resent {\em time-dependent} performance measure for economical time series -- the (optimal) investment horizon approach. For stock indices, the approach shows a pronounced gain-loss asymmetry that is {\em not} observed for the individual stocks that comprise the index. This difference may hint towards an sy…

2005-04-21abs ↗pdf ↗

Behavior cloning can achieve horizon-independent sample complexity in offline imitation learning.

problem Sample complexity in imitation learning increases with problem horizon.
method New analysis of behavior cloning with logarithmic loss.
result Behavior cloning can achieve linear dependence on horizon in offline IL under dense rewards.

This paper examines the volatility and covariance dynamics of cash and futures contracts that underlie the Optimal Hedge Ratio (OHR) across different hedging time horizons. We examine whether hedge ratios calculated over a short term hedging horizon can be scaled and successfully applied to longer term horizons. We als…

2011-03-30abs ↗pdf ↗

New algorithms learn MDPs with better regret bounds using generative sampling.

problem Learning MDPs with optimal policies under uncertainty.
method Hybrid exploration-generative RL model, classical and quantum algorithms.
result Quantum algorithms achieve polylogT\operatorname{poly}\log{T} regret for infinite-horizon MDPs.

Recently, there has been significant progress in understanding reinforcement learning in discounted infinite-horizon Markov decision processes (MDPs) by deriving tight sample complexity bounds. However, in many real-world applications, an interactive learning agent operates for a fixed or bounded period of time, for ex…

2015-10-29abs ↗pdf ↗

Heterotic horizons preserving 4 supersymmetries have sections which are T^2 fibrations over 6-dimensional conformally balanced Hermitian manifolds. We give new examples of horizons with sections S^3 X S^3 X T^2 and SU(3). We then examine the heterotic horizons which are T^4 fibrations over a Kahler 4-dimensional manifo…

2010-03-15abs ↗pdf ↗

New algorithm for RL with horizon-free reward-free exploration for linear MDPs.

problem Reward-free reinforcement learning with long planning horizons.
method Uncertainty-weighted value-targeted regression with exploration-driven pseudo-reward and moment estimator.
result Horizon-free sample complexity of O(d2ε2)O(d^2\varepsilon^{-2}) for finding an ε\varepsilon-optimal policy.

Long horizon reinforcement learning is as hard as short horizon learning.

problem Understanding the difficulty of long horizon reinforcement learning problems.
method Introduced new concepts: ε-net for optimal policies and Online Trajectory Synthesis algorithm.
result Proved that sample complexity scales logarithmically with the planning horizon, refuting the conjecture.

New algorithm reduces reinforcement learning complexity, approaching contextual bandits.

problem Episodic reinforcement learning's difficulty compared to contextual bandits.
method Proposes MVP algorithm with a new Bernstein-type bonus for episodic reinforcement learning.
result Achieves near-optimal regret bound of $O\left(\left(\sqrt{SAK} + S^2A ight) \poly\log \left(SAHK ight) ight)$, improving state-of-the-art results.

Study optimal consumption with drawdown limits over a fixed time frame.

problem Maximizing utility with consumption limits during a fixed period.
method Extended utility maximization problem with drawdown constraint, using PDE arguments and dual transform.
result Existence and uniqueness of classical solution to HJB variational inequality, with explicit free boundaries.

Optimizes investment under uncertain time horizons with non-concave utility.

problem Optimizing investment decisions with non-concave utility and uncertain time horizons.
method Established necessary and sufficient conditions for optimality, suggested recursive procedure for non-concave utility.
result Optimal investment strategies under uncertain time horizons exhibit multimodal distribution, indicating flexibility in switching between local maximizers.

Study on cryptocurrency market correlations at various time scales.

problem Understanding the hierarchical structure of cryptocurrency market dynamics.
method Analysis of MST and TMFG for 25 liquid cryptocurrencies at different time horizons.
result Cryptocurrency market correlations decrease with finer time scales and show a growing hierarchical structure with coarser scales.

Temporal aggregation reveals latent default correlation from monthly data.

problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.

Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.

problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.

Many robotic applications require the agent to perform long-horizon tasks in partially observable environments. In such applications, decision making at any step can depend on observations received far in the past. Hence, being able to properly memorize and utilize the long-term history is crucial. In this work, we pro…

2019-03-09abs ↗pdf ↗

For an investor with constant absolute risk aversion and a long horizon, who trades in a market with constant investment opportunities and small proportional transaction costs, we obtain explicitly the optimal investment policy, its implied welfare, liquidity premium, and trading volume. We identify these quantities as…

2011-10-06abs ↗pdf ↗

The paper extends spacetime topology results using codimension 2 null cut locus properties.

problem Understanding spacetime topology with and without horizons.
method Review and extension of existing literature on spacetime topology, utilizing codimension 2 null cut locus properties.
result Results for spacetimes with and without horizons, including asymptotically AdS settings.

We demonstrate the existence of spherically-symmetric truly naked black holes (TNBH) for which the Kretschmann scalar is finite on the horizon but some curvature components including those responsible for tidal forces as well as the energy density ρˉ\barρ measured by a free-falling observer are infinite. We choose a ra…

2007-06-19abs ↗pdf ↗

We establish that an optimistic variant of Q-learning applied to a fixed-horizon episodic Markov decision process with an aggregated state representation incurs regret O~(H5MK+εHK)\tilde{\mathcal{O}}(\sqrt{H^5 M K} + εHK), where HH is the horizon, MM is the number of aggregate states, KK is the number of episodes, and εε is …

2019-12-13abs ↗pdf ↗

The paper examines how loss aversion impacts multi-armed bandit decisions over long periods.

problem The impact of loss aversion on multi-armed bandit decisions over long periods.
method A new central limit theorem for measures with history-dependent variances, derived under risk aversion in gains and risk loving in losses.
result Consequences of loss aversion for asymptotic properties are derived in analytical results.

Forecastability measures predictive information across horizons.

problem How much predictive information is available at each prediction horizon?
method Develops the consequences of mutual information between future observations and information set.
result Forecastability is a profile reflecting process dependence structure, with properties like compression and truncation error.

Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.

problem Optimal stopping problems with finite-time horizon and state-dependent discounting.
method Linear diffusion process, time-homogeneous gain function, fine regularity properties, continuity and strict monotonicity proof.
result Proves continuity and strict monotonicity of the optimal stopping boundary under mild assumptions.

We characterise the value function of the optimal dividend problem with a finite time horizon as the unique classical solution of a suitable Hamilton-Jacobi-Bellman equation. The optimal dividend strategy is realised by a Skorokhod reflection of the fund's value at a time-dependent optimal boundary. Our results are obt…

2016-09-06abs ↗pdf ↗

Study forecasts sub-city real estate prices weekly using radar and news sentiment.

problem Limited availability of reliable real estate price indicators at neighborhood and long horizons.
method Combining satellite radar signals and news sentiment to forecast sub-city real estate prices.
result The multimodal model reduces mean absolute error by 35% at long horizons (26-34 weeks).

Study optimal policy regret in partially observable Markov games with adaptive opponents.

problem Optimal sequential decision-making in partially observable environments against strategic, adaptive opponents.
method An epoch-based optimistic maximum-likelihood algorithm that selects one policy per epoch using confidence sets built cumulatively from past data.
result Achieves ildeO(T) ilde{O}(\sqrt{T}) policy regret for fixed problem parameters, with explicit dependence on horizon, adversary memory, confidence radius, and aggregate Eluder dimension.

Optimal reinsurance and dividend strategy for insurance companies in a finite time.

problem Maximizing dividends while managing risk in a finite time horizon.
method Dynamic control problem with Hamilton-Jacobi-Bellman equation, penalty approximation method.
result Smoothness of the value function and comparison principle for its gradient.

Study on black hole interiors with matter fields, showing oscillation condition impacts blow-up.

problem Examining Strong Cosmic Censorship in the presence of matter fields.
method Einstein equations coupled with charged/massive scalar fields, spherically symmetric data, relaxation rate analysis.
result Oscillation condition on event horizon determines whether matter fields blow up or not.