Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4795142189 · Jun 202019922001200920182026
48 results for stationary action

New method estimates state-action stationary distribution for better off-policy policy evaluation.

problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.

Study stationary measures and orbit closures for non-abelian actions on surfaces.

problem Classify stationary measures and orbit closures for non-abelian action on a surface.
method Use a finite verifiable average growth condition and results from Brown and Rodriguez Hertz.
result Show that under certain conditions, the only nonatomic stationary measure is the given smooth invariant measure, and every orbit closure is either finite or dense.

KeRNS tackles non-stationary reinforcement learning in metric spaces.

problem Non-stationary reinforcement learning in metric spaces.
method KeRNS uses time-dependent kernels to model non-stationary Markov Decision Processes (MDPs).
result KeRNS achieves a regret bound that scales with the covering dimension and total variation of the MDP.

NVMDP framework tackles non-stationary MDPs with varying discount rates.

problem Challenges in non-stationary environments and infinite-horizon formulations for reinforcement learning.
method Introduces NVMDP framework that accommodates non-stationarity and varying discount rates.
result NVMDPs provide a flexible mechanism to shape optimal policies without altering state or action spaces.

Non-stationary reinforcement learning is challenging due to the complexity of updating value functions.

problem Challenges in non-stationary reinforcement learning, especially in updating value functions.
method Proved a worst-case complexity result for modifying reinforcement learning problems.
result Modifying reinforcement learning problems requires an amount of time almost as large as the number of states.

Optimal algorithms for non-stationary bandits minimize regret over time.

problem Designing efficient algorithms for non-stationary bandit problems.
method Kiefer Wolfowitz (KW) algorithm with fixed and sliding window step sizes.
result KW algorithms achieve asymptotically optimal regret under certain conditions.

Predictive sampling improves on Thompson sampling for non-stationary bandit environments.

problem Thompson sampling fails in non-stationary bandit environments.
method Proposes predictive sampling, which deprioritizes actions based on information loss rate.
result Predictive sampling outperforms Thompson sampling in all tested non-stationary environments.

The paper studies limit sets on P(R3)\mathbb{P}(\mathbb{R}^3) using stationary measures.

problem Investigating the Hausdorff dimension of limit sets on P(R3)\mathbb{P}(\mathbb{R}^3) for SL3(R)\mathrm{SL}_3(\mathbb{R}).
method Using stationary measures to generalize the Patterson-Sullivan formula and establish dimension formulas.
result Sharp lower bounds and Hausdorff dimensions for Anosov representations and the Rauzy gasket.

New algorithm tackles non-stationary reinforcement learning with general function approximation.

problem Understanding non-stationary MDPs with function approximation.
method Dynamic Bellman Eluder (DBE) dimension for complexity, sliding window mechanism, confidence set design.
result Upper bound on dynamic regret for proposed SW-OPEA algorithm.

New method stabilizes FQE by reweighting Bellman targets.

problem Stability guarantees for FQE often rely on Bellman completeness, which can fail with function approximation.
method Proposes stationary-weighted FQE, reweighting Bellman targets by stationary target-to-behavior density ratio.
result Proves finite-sample linear convergence to stationary projected Bellman fixed point without Bellman completeness.

The paper suggests asset prices follow physical laws, allowing for accurate price movement forecasts.

problem Predicting extreme price movements in financial markets.
method Modeling asset price dynamics as a harmonic oscillator and applying the principle of stationary action.
result The theory can make accurate forecasts of price movements during market crashes and specific price displacements at other times.

Study tackles non-stationary bandit convex optimization with new algorithms.

problem Minimizing regret in non-stationary environments with various measures of non-stationarity.
method Proposed Tilted Exponentially Weighted Average with Sleeping Experts (TEWA-SE) for strongly convex losses and clipped Exploration by Optimization (cExO) for general convex losses.
result Proved minimax-optimality of TEWA-SE for strongly convex losses and introduced cExO for general convex losses.

Internal Lagrangians derived from variational principles.

problem Reproducing the principle of stationary action in variational geometry.
method Introducing stationary points of internal Lagrangians, establishing connections with symmetries and conservation laws, and investigating relations between non-degenerate and internal Lagrangians.
result Noether's theorem reformulated in terms of internal Lagrangians.

Memory-based models can learn to approximate Bayes-optimal predictors for non-stationary data.

problem Learning from non-stationary data with unobserved switching points.
method Memory-based neural models, including Transformers, LSTMs, and RNNs, trained to minimize log loss.
result Memory-based models can accurately approximate known Bayes-optimal algorithms and perform Bayesian inference over latent switching points.

New algorithm optimizes multi-objective outcomes in uncertain environments.

problem Optimizing global concave rewards in online Markov decision processes with multiple actions.
method No-regret algorithm based on online convex optimization and UCRL2, with a gradient threshold procedure.
result Non-stationary policy diversifies outcomes to optimize the global concave reward.

Study tackles OPE in confounded settings, estimating policy value from proxies.

problem Difficulty in OPE due to unobserved confounders in infinite-horizon RL.
method Two-stage approach: estimating stationary distribution ratios and combining optimal balancing.
result Policy value can be identified from off-policy data with proxies and latent variable model.

New RL algorithm tackles non-stationary environments with flexible policy updates.

problem Non-stationary reinforcement learning with time-varying rewards and transition probabilities.
method Model-free policy-based algorithm NS-NAC with restart-based exploration and dynamic learning rates.
result Dynamic regret of ildeO(S1/2A1/2ΔT1/6T5/6) ilde{\mathscr O}(|S|^{1/2}|A|^{1/2}Δ_T^{1/6}T^{5/6}) for both algorithms.

The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…

2013-11-07abs ↗pdf ↗

We consider the reduced Allen-Cahn action functional, which appears as the sharp interface limit of the Allen-Cahn action functional and can be understood as a formal action functional for a stochastically perturbed mean curvature flow. For suitable evolutions of generalized hypersurfaces this functional consists of th…

2013-04-07abs ↗pdf ↗

Develops a new method for optimizing policies in hierarchical models.

problem Optimizing complex policies in hierarchical models.
method Applies second-order methods in the space of state-action paths.
result The natural path gradient method can be computed exactly and reflects state-space hierarchy.

A new algorithm for non-stationary linear bandits with improved regret bound.

problem Non-stationary linear bandit problem with time-varying rewards.
method D-LinUCB, a discounted linear regression algorithm with exponential weights.
result Upper bound on dynamic regret of order d^{2/3} B_T^{1/3}T^{2/3}, optimal in slowly-varying and abruptly-changing environments.

A new method optimizes in nonstationary environments with many arms efficiently.

problem Optimizing in nonstationary environments with a large number of arms.
method Gaussian interpolation to learn continuous Lipschitz reward functions in nonstationary environments.
result Efficiently learns continuous Lipschitz reward functions with O(T)\mathcal{O}^*(\sqrt{T}) cumulative regret.

Study extreme values of stable random fields on geometric spaces.

problem Understanding extreme values of stable random fields on various geometric spaces.
method Analyzing extreme values through Patterson-Sullivan measures and extremal cocycle growth.
result Established a dichotomy for the growth-rate of maxima sequences of stable random fields.

The paper studies critical points and flows of a G2G_2-Hilbert functional on manifolds with circle actions.

problem Critical points and flows of the G2G_2-Hilbert functional on manifolds with S1\mathbb S^1-actions.
method Analysis of S1\mathbb S^1-invariant G2G_2-structures, reduction to a 6-dimensional quotient, and derivation of a negative L2L^2-gradient flow.
result The unnormalized flow admits only trivial stationary configurations: flat connection, scalar-flat base metric, and constant fiber length.

Paper introduces Decentralized Non-stationary Competing Bandits ( exttt{DNCB}) for dynamic matching markets.

problem Understanding dynamic two-sided matching markets with competing agents.
method Proposes a decentralized asynchronous learning algorithm ( exttt{DNCB}) for non-stationary environments.
result Obtains sub-linear (logarithmic) regret of exttt{DNCB} in dynamic settings.

Two randomized algorithms improve performance in non-stationary linear bandits.

problem Conservatism in optimistic algorithms for non-stationary linear bandits.
method Two perturbation approaches: randomization and random perturbations.
result D-RandLinUCB and D-LinTS achieve optimal dynamic regret and are oracle-efficient.

A new kernel improves Gaussian process performance for non-stationary data.

problem Poor prediction and uncertainty quantification with standard GPs.
method Study and comparison of non-stationary kernels, propose a new combined kernel.
result A new kernel outperforms existing stationary and non-stationary kernels.

New findings on universal learning in contextual bandits with adversarial rewards.

problem Learning in contextual bandits with time-varying, adversarial rewards.
method Characterization of learnable processes and necessary/sufficient conditions for universal learning.
result Optimistic universal learning for contextual bandits with adversarial rewards is impossible in general.

Formula derived for spectral determinant of sphere with conical singularities.

problem Calculating the spectral determinant of a sphere with conical singularities.
method Explicit closed formula derived using zeta regularization and Liouville action.
result Metrics with equal conical angles are a stationary point of the determinant, and a minimum if surface area is small.

Unified approach for non-stationary linear bandits with dynamic regret.

problem Non-stationary linear bandits with round-specific feasible actions and drifting reward models.
method Unified misspecification-reduction viewpoint, restarting algorithms with misspecification-dependent regret guarantees.
result Optimal \(T^{2/3}P_T^{1/3}\) dynamic-regret dependence for both linear bandits and contextual linear bandits.

The paper deals with the Weyl equation which is the massless Dirac equation. We study the Weyl equation in the stationary setting, i.e. when the spinor field oscillates harmonically in time. We suggest a new geometric interpretation of the stationary Weyl equation, one which does not require the use of spinors, Pauli m…

2010-01-26abs ↗pdf ↗

The second order differential equation Dγ˙dt(t)=Fγ(t)(γ˙(t))V(γ(t))\frac{D\dotγ}{dt}(t) = F_{γ(t)}(\dotγ(t)) - \nabla V(γ(t)) on a Lorentzian manifold describes, in particular, the dynamics of particles under the action of a electromagnetic field FF and a conservative force V-\nabla V. We provide a first study on the extendability of its solu…

2012-11-09abs ↗pdf ↗

If a differential equation in a Banach manifold is invariant or quasi-invariant under the action of one or more Lie groups, then its stationary points cannot be isolated, so that classical linearized stability theorem does not apply to it. The first main purpose of this paper is to establish a linearized stability theo…

2016-06-30abs ↗pdf ↗