Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

23456890 · Jun 202019922001200920172026
48 results for Instantaneous Regret

Improved prediction algorithm for 'easy' sequences with reduced regret.

problem Prediction with expert advice for 'easy' sequences.
method Variant of NormalHedge algorithm using second-order εε-quantile regret bound.
result Second-order εε-quantile regret bound of O(VTlog(VT/ε))O\big(\sqrt{V_T \log(V_T/ε)}\big) for VT>logNV_T > \log N.

New RL algorithm tackles adversarial RMAB with unknown transitions and bandit feedback.

problem Learning in episodic RMAB with unknown transition functions and adversarial rewards.
method Developed a novel RL algorithm with a biased reward estimator and an index policy.
result Achieved ildeO(HT) ilde{\mathcal{O}}(H\sqrt{T}) regret bound for adversarial RMAB.

SMEs provide a transparent testbed for RL evaluation.

problem Lack of precise, white-box diagnostics in RL environments.
method Synthetic Monitoring Environments (SMEs) with fully configurable task characteristics and known optimal policies.
result SMEs allow for precise evaluation of RL algorithms, revealing the impact of specific environmental properties.

TVBO optimizes time-varying functions with asymptotically vanishing regret.

problem Understanding the asymptotic performance of TVBO for time-varying black-box functions.
method Provided upper and lower bounds for cumulative regret of TVBO algorithms.
result TVBO algorithms can achieve asymptotically vanishing regret under certain conditions.

This paper introduces a new metric, ULI, for RL that ensures both cumulative and instantaneous performance.

problem High-stakes applications require RL algorithms to avoid playing bad policies.
method Introduces uniform last-iterate (ULI) guarantee, a stronger metric capturing both cumulative and instantaneous performance.
result ULI directly implies near-optimal cumulative performance across various metrics, but not the other way around.

New algorithm reduces regret in stochastic shortest path problems.

problem Planning and control in environments with unknown dynamics and variable episode lengths.
method Developed an algorithm with a new regret bound of O(BSAK)O(B_\star |S| \sqrt{|A| K}).
result Guaranteed a significant reduction in regret compared to previous methods.

We consider a collaborative online learning paradigm, wherein a group of agents connected through a social network are engaged in playing a stochastic multi-armed bandit game. Each time an agent takes an action, the corresponding reward is instantaneously observed by the agent, as well as its neighbours in the social n…

2016-02-29abs ↗pdf ↗

We study online aggregation of the predictions of experts, and first show new second-order regret bounds in the standard setting, which are obtained via a version of the Prod algorithm (and also a version of the polynomially weighted average algorithm) with multiple learning rates. These bounds are in terms of excess l…

2014-02-10abs ↗pdf ↗

Two new algorithms optimize rewards while respecting safety constraints in sequential decisions.

problem Optimizing rewards with safety constraints in sequential decisions.
method Stage-wise conservative linear Thompson Sampling (SCLTS) and stage-wise conservative linear UCB (SCLUCB).
result Probabilistic regret bounds of order O(\sqrt{T} \log^{3/2}T) and O(\sqrt{T} \log T).

We prove uniqueness of instantaneously complete Ricci flows on surfaces. We do not require any bounds of any form on the curvature or its growth at infinity, nor on the metric or its growth (other than that implied by instantaneous completeness). Coupled with earlier work, particularly [23, 11], this completes the well…

2013-05-08abs ↗pdf ↗

Study cooperative bandit learning with imperfect communication, achieving near-optimal performance.

problem Real-world distributed decision-making with imperfect communication.
method Proposed decentralized algorithms for three communication scenarios: stochastic networks, random delays, and adversarially corrupted rewards.
result Achieved competitive performance and near-optimal guarantees on group regret.

New framework IDOL identifies latent causal processes with instantaneous relations from time series data.

problem Identifying latent causal processes with instantaneous relations from time series data.
method Sparse influence constraint and variational inference architecture with sparsity regularization.
result Our method can identify latent causal processes with instantaneous relations.

This paper studies the concept of instantaneous arbitrage in continuous time and its relation to the instantaneous CAPM. Absence of instantaneous arbitrage is equivalent to the existence of a trading strategy which satisfies the CAPM beta pricing relation in place of the market. Thus the difference between the arbitrag…

2019-01-16abs ↗pdf ↗

iCITRIS learns causal variables from interactive systems with instantaneous effects.

problem Identifying causal variables from temporal sequences with instantaneous effects.
method iCITRIS method for causal representation learning that handles instantaneous effects in intervened temporal sequences.
result iCITRIS accurately identifies causal variables and their causal graph from three interactive system datasets.

Online algorithms for identifying river pollution sources.

problem Real-time estimation of river pollution sources from downstream data.
method Gradient-based online learning algorithms with adaptive step sizes and escaping from saddle points module.
result High estimation accuracy in three dimensions, superior to existing methods.

The Ricci flow preserves product structures with instantaneous curvature bounds.

problem Preserving product structures under Ricci flow with curvature constraints.
method Proving a constant ε exists such that if a solution splits as a product at time 0 and has bounded curvature, it splits for all time.
result A constant ε exists depending on dimension such that if a solution splits as a product at time 0 and has curvature bounded by ε/t, it splits for all time.

This work improves wireless network learning by using side-information about interference.

problem Improving online learning algorithms in wireless networks.
method Exploiting side-information like interference levels to improve learning algorithms.
result Improved learning algorithms achieve higher throughput with fewer samples.

This paper presents a novel one-factor stochastic volatility model where the instantaneous volatility of the asset log-return is a diffusion with a quadratic drift and a linear dispersion function. The instantaneous volatility mean reverts around a constant level, with a speed of mean reversion that is affine in the in…

2019-08-20abs ↗pdf ↗

Study cryptocurrency price dynamics using adaptive EMD and spectral analysis.

problem Analyze the time-varying volatility of cryptocurrency prices.
method Adaptive complementary ensemble empirical mode decomposition (ACE-EMD) and Hilbert spectral analysis.
result Reveal the properties of various timescales in cryptocurrency price dynamics.

New algorithm resists corruption in linear contextual bandits.

problem Adversarial corruption in linear contextual bandits.
method Variance-aware algorithm with multi-level partition and adaptive confidence sets.
result Regret bound of ildeO(C2dt=1Tσt2+C2RdT) ilde{O}(C^2d\sqrt{\sum_{t = 1}^T σ_t^2} + C^2R\sqrt{dT}).

We study hedging and pricing of unattainable contingent claims in a non-Markovian regime-switching financial model. Our financial market consists of a bank account and a risky asset whose dynamics are driven by a Brownian motion and a multivariate counting process with stochastic intensities. The interest rate, drift, …

2013-03-17abs ↗pdf ↗

Working on different aspects of algorithmic trading we empirically discovered a new market invariant. It links together the volatility of the instrument with its traded volume, the average spread and the volume in the order book. The invariant has been tested on different markets and different asset classes. In all cas…

2019-08-07abs ↗pdf ↗

Collective behaviours taking place in financial markets reveal strongly correlated states especially during a crisis period. A natural hypothesis is that trend reversals are also driven by mutual influences between the different stock exchanges. Using a maximum entropy approach, we find coordinated behaviour during tre…

2013-10-30abs ↗pdf ↗

Study optimal execution in a transient price impact model with multiple traders.

problem Optimal execution among multiple traders with transient price impact.
method Analyzed NN-player optimal execution games in an Obizhaeva--Wang model with and without regularization. Derived equilibrium solutions and explained their behavior.
result Existence of equilibrium restored with a specific time-dependent cost on block trades, and equilibrium is tractable.

Estimates chirp signal frequencies using probabilistic models.

problem Estimating instantaneous frequencies of chirp signals when true forms are unknown.
method Non-linear Gaussian processes and stochastic filters/smothers for posterior estimation.
result The method outperforms state-of-the-art methods on synthetic and real-world datasets.

A new principle minimizes residual and introduces momentum to improve PDE solution dynamics.

problem Ill-conditioning in Dirac-Frenkel residual minimization leads to non-unique parameter dynamics.
method Introduces a history variable (momentum) to select better-conditioned parameter velocities, preserving residual minimization while promoting smooth parameter evolutions.
result The approach leads to increased robustness in singular and near-singular PDE solution regimes.

New model identifies regimes in non-stationary data.

problem Identifying latent regimes in non-stationary systems with instantaneous effects.
method Identifiable Markov Switching Models with exponential family noise.
result Established identifiability of latent regimes and causal structures.

Unified framework for optimal liquidation with small market impact and semimartingale strategies.

problem Optimal liquidation under small market impact and portfolio liquidation.
method Semimartingale strategies and convergence results for BSDEs with singular terminal conditions.
result Unified framework for embedding two common liquidation models and microscopic foundation for semimartingale strategies.

To convert standard Brownian motion ZZ into a positive process, Geometric Brownian motion (GBM) eβZt,β>0e^{βZ_t}, β>0 is widely used. We generalize this positive process by introducing an asymmetry parameter α0 α\geq 0 which describes the instantaneous volatility whenever the process reaches a new low. For our new process, …

2018-09-06abs ↗pdf ↗

We explore the effect of past market movements on the instantaneous correlations between assets within the futures market. Quantifying this effect is of interest to estimate and manage the risk associated to portfolios of futures in a non-stationary context. We apply and extend a previously reported method called the P…

2019-12-27abs ↗pdf ↗

Bayesian optimization is a sample-efficient method for finding a global optimum of an expensive-to-evaluate black-box function. A global solution is found by accumulating a pair of query point and its function value, repeating these two procedures: (i) modeling a surrogate function; (ii) maximizing an acquisition funct…

2019-01-24abs ↗pdf ↗

HHT feature generation enhances financial time series forecasting.

problem Forecasting nonstationary financial time series.
method CEEMD and HHT for decomposition, machine learning integration.
result HHT-enhanced models outperform traditional models in forecasting.