Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

168335503670 · Jun 202019922001200920172026
48 results for horizon function

UCRL-WVTR tackles long-term reinforcement learning with general approximations, achieving horizon-free and instance-dependent regret bounds.

problem Long-term reinforcement learning with general function approximations.
method UCRL-WVTR proposes a novel algorithm, UCRL-WVTR, with weighted value-targeted regression and a high-order moment estimator.
result Achieves horizon-free and instance-dependent regret bounds matching minimax lower bounds up to logarithmic factors.

Study optimal liquidation strategies with infinite horizon and regime switching.

problem Optimal liquidation with semimartingale strategies in a stochastic environment.
method Characterization of value function and optimal strategy via BSDEs with infinite horizon.
result Existence and uniqueness of optimal control problem solutions.

This paper studies the utility maximization problem with changing time horizons in the incomplete Brownian setting. We first show that the primal value function and the optimal terminal wealth are continuous with respect to the time horizon TT. Secondly, we exemplify that the expected utility stemming from applying th…

2010-06-25abs ↗pdf ↗

Optimizes investment under uncertain time horizons with non-concave utility.

problem Optimizing investment decisions with non-concave utility and uncertain time horizons.
method Established necessary and sufficient conditions for optimality, suggested recursive procedure for non-concave utility.
result Optimal investment strategies under uncertain time horizons exhibit multimodal distribution, indicating flexibility in switching between local maximizers.

A new Bayesian method optimizes time-dependent expensive functions with lookahead.

problem Maximizing a time-dependent, expensive oracle with limited evaluations.
method Recursive, two-step lookahead expected payoff (r2LEY) acquisition function.
result r2LEY outperforms myopic methods in synthetic and real-world datasets.

New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.

problem Challenges in online reinforcement learning for non-episodic, finite-horizon MDPs.
method Introduces a K-step lookahead Q-function with a time-varying threshold for selecting actions.
result Achieves minimax optimal constant regret for K=1 and O(max((K1),CK1)SATlog(T))\mathcal{O}(\max((K-1),C_{K-1})\sqrt{SAT\log(T)}) regret for K ≥ 2.

We consider the problem of portfolio optimization in a simple incomplete market and under a general utility function. By working with the associated Hamilton-Jacobi-Bellman partial differential equation (HJB PDE), we obtain a closed-form formula for a trading strategy which approximates the optimal trading strategy whe…

2016-11-28abs ↗pdf ↗

Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.

problem Optimal stopping problems with finite-time horizon and state-dependent discounting.
method Linear diffusion process, time-homogeneous gain function, fine regularity properties, continuity and strict monotonicity proof.
result Proves continuity and strict monotonicity of the optimal stopping boundary under mild assumptions.

New algorithm learns optimal policies with just 1 episode, settling horizon-dependence in RL.

problem Understanding the sample complexity of reinforcement learning with horizon length.
method Developed an algorithm using only O(1)O(1) episodes to achieve PAC guarantee, leveraging connections between value functions in discounted and finite-horizon MDPs and novel perturbation analysis.
result Achieved the same PAC guarantee with only O(1)O(1) episodes of environment interactions, completely settling horizon-dependence in RL.

Study examines Wang-Yau quasi-local energy in strong fields near apparent horizons.

problem Examining the behavior of Wang-Yau quasi-local energy near apparent horizons in strong fields.
method Analyzing the limit of the Wang-Yau quasi-local energy as a spacelike surface approaches an apparent horizon, considering bounded coordinate functions and spacelike mean curvature.
result The limit of the Wang-Yau quasi-local energy falls into two cases: it blows up or remains finite, depending on whether the horizon can be isometrically embedded into R3R^3.

UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.

problem Learning infinite-horizon average-reward MDPs with linear function approximation.
method UCRL2-VTR algorithm with Bernstein-type bonus.
result Achieves a regret of ildeO(dDT) ilde{O}(d\sqrt{DT}) with matching lower bound.

In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run. Yet, it may be difficult (or even intractable) mathematically to learn with th…

2019-02-05abs ↗pdf ↗

The paper finds the shortest time to exploit arbitrage in multi-stock markets.

problem Finding the shortest time to exploit arbitrage in multi-stock markets.
method Characterizes the minimal time horizon for relative arbitrage in markets with 2 to 3 stocks and uses geometric flows for markets with 4 or more stocks.
result Explicit computation of minimal time horizon for 2 and 3 stocks markets, and characterization via geometric flows for markets with 4 or more stocks.

New research shows exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions.

problem Determining the minimum number of queries needed for sound planners in MDPs with linear function approximation.
method Analyzing fixed-horizon and discounted MDPs with a generative model, showing lower bounds on the number of queries required.
result Sound planners need at least exponential number of queries in both fixed-horizon and discounted settings.

Heterotic horizons preserving 4 supersymmetries have sections which are T^2 fibrations over 6-dimensional conformally balanced Hermitian manifolds. We give new examples of horizons with sections S^3 X S^3 X T^2 and SU(3). We then examine the heterotic horizons which are T^4 fibrations over a Kahler 4-dimensional manifo…

2010-03-15abs ↗pdf ↗

The paper optimizes portfolios in a financial market with correlated assets using a stochastic volatility model.

problem Optimizing portfolios in a financial market with correlated assets and stochastic volatility.
method Derive a Hamilton-Jacobi-Bellman equation, use approximation methods, analyze value function using expansion of utility function, control error with second-order terms, generate close-to-optimal portfolio.
result Close-to-optimal portfolio generated using first-order approximation of utility function with controlled error.

A new ML algorithm solves complex economic control problems.

problem Solving high-dimensional, finite-horizon stochastic control problems in economics.
method Deep neural network representation of optimal policy functions with three key features.
result Efficiently solves various economic control problems including recursive utility and growth models.

Careful tuning of the learning rate, or even schedules thereof, can be crucial to effective neural net training. There has been much recent interest in gradient-based meta-optimization, where one tunes hyperparameters, or even learns an optimizer, in order to minimize the expected loss when the training procedure is un…

2018-03-06abs ↗pdf ↗

New algorithm for RL with horizon-free reward-free exploration for linear MDPs.

problem Reward-free reinforcement learning with long planning horizons.
method Uncertainty-weighted value-targeted regression with exploration-driven pseudo-reward and moment estimator.
result Horizon-free sample complexity of O(d2ε2)O(d^2\varepsilon^{-2}) for finding an ε\varepsilon-optimal policy.

Using high frequency data, we have studied empirically the change of volatility, also called volatility derivative, for various time horizons. In particular, the correlation between the volatility derivative and the volatility realized in the next time period is a measure of the response function of the market particip…

2001-05-08abs ↗pdf ↗

We present a fully nonparametric method to estimate the value function, via simulation, in the context of expected infinite-horizon discounted rewards for Markov chains. Estimating such value functions plays an important role in approximate dynamic programming and applied probability in general. We incorporate "soft in…

2013-12-26abs ↗pdf ↗

Modeling risk and performance with Levy-stable distributions.

problem Understanding risk and performance in financial markets with non-Gaussian distributions.
method Developed a finite-horizon model using Levy-stable scaling, identified parameters from data, derived formulas for various financial ratios.
result Horizon-correct formulas for risk measures are derived and validated across different horizons.

In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …

2016-08-22abs ↗pdf ↗

Langevin dynamics fails to produce accurate samples even with small score function errors.

problem Robustness of Langevin dynamics to score function errors.
method Analysis of Langevin dynamics and score function errors.
result Langevin dynamics produces a distribution far from the target distribution in TV distance even with small L2L^2 errors in the score function.

Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationar…

2019-10-16abs ↗pdf ↗

Optimizes portfolio in volatile markets with jumps, providing accurate formulas.

problem Optimizing wealth in a volatile financial market with jumps.
method Analyzes an incomplete stochastic volatility model, derives closed-form portfolio formulas using HJB equation and super-solution/sub-solution.
result Proves accuracy of derived portfolio formulas for both small and finite time horizons.

New algorithms reduce contextual bandits' regret without knowing reward noise variances.

problem Reducing regret in contextual bandits with unknown reward noise variances.
method Developed new algorithms based on the optimism principle.
result Regret scales as the square root of the sum of measurement variances, not the time horizon.

Scalar-tensor gravitation theories, such as the Brans-Dicke family of theories, are commonly partly described by a modified Einstein equation in which the Ricci tensor is replaced by the Bakry-Émery-Ricci tensor of a Lorentzian metric and scalar field. In physics this formulation is sometimes referred to as the "Jordan…

2013-10-15abs ↗pdf ↗

In this paper, we study optimal liquidation problems in a randomly-terminated horizon. We consider the liquidation of a large single-asset portfolio with the aim of minimizing a combination of volatility risk and transaction costs arising from permanent and temporary market impact. Three different scenarios are analyze…

2017-09-18abs ↗pdf ↗

New algorithms for learning MDPs with linear approximations in infinite-horizon settings.

problem Learning infinite-horizon average-reward MDPs with linear function approximation.
method Optimism principle, adversarial linear bandits, Natural Policy Gradient.
result Efficient algorithms with optimal or near-optimal regret bounds.

We analyze (the harmonic map representation of) static solutions of the Einstein Equations in dimension three from the point of view of comparison geometry. We find simple monotonic quantities capturing sharply the influence of the Lapse function on the focussing of geodesics. This allows, in particular, a sharp estima…

2011-03-24abs ↗pdf ↗

Study optimal consumption and investment for investors with Epstein-Zin preferences.

problem Optimal consumption and investment for investors with Epstein-Zin preferences in an incomplete market.
method Variational characterisation and direct method to prove existence of optimal policies.
result Existence and uniqueness of optimal consumption and investment policies.

Optimal reinsurance and dividend strategy for insurance companies in a finite time.

problem Maximizing dividends while managing risk in a finite time horizon.
method Dynamic control problem with Hamilton-Jacobi-Bellman equation, penalty approximation method.
result Smoothness of the value function and comparison principle for its gradient.

We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…

2018-10-29abs ↗pdf ↗