Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

61122182243 · Jun 202019922001200920172026
48 results for conservative exploration

StepMix algorithm ensures safe exploration in reinforcement learning with near-optimal performance.

problem Conservative exploration in reinforcement learning under episode-wise constraints.
method StepMix algorithm balances exploitation and exploration while ensuring episode-wise conservative constraint.
result StepMix achieves near-optimal regret order as in the constraint-free setting.

The paper tackles safe exploration in RL by a conservative safety critic.

problem Safe exploration in reinforcement learning (RL) when partially trained policies are deployed.
method Learning a conservative safety estimate through a critic, provably bounding catastrophic failures.
result The approach provably converges to competitive task performance with significantly lower catastrophic failure rates.

Proposes a method to avoid excessive exploration in reinforcement learning.

problem Avoiding excessive exploration in reinforcement learning to deploy it in practice.
method Designs a novel algorithm using UCB reinforcement learning policy with adaptive exploration constraints.
result Proves that the approach remains conservative while minimizing regret in tabular settings and validates on real-world tasks.

While learning in an unknown Markov Decision Process (MDP), an agent should trade off exploration to discover new information about the MDP, and exploitation of the current knowledge to maximize the reward. Although the agent will eventually learn a good or optimal policy, there is no guarantee on the quality of the in…

2020-02-08abs ↗pdf ↗

A challenging problem in complex networks is the network reconstruction problem from data. This work deals with a class of networks denoted as conserved networks, in which a flow associated with every edge and the flows are conserved at all non-source and non-sink nodes. We propose a novel polynomial time algorithm to …

2019-05-21abs ↗pdf ↗

Paper proposes an RL algorithm to ensure policy performance guarantees.

problem Lack of performance guarantees for RL policies compared to baselines.
method Online model-free algorithm that ensures conservative exploration.
result Regret bound of ildeO(T) ilde{\mathcal{O}}(\sqrt{T}) for both discrete and continuous spaces.

The paper explores symmetries and conserved charges on pre-symplectic manifolds.

problem Analyzing conserved charges on solutions of Hamiltonian field theories.
method Using pre-symplectic structures and Gotay's coisotropic embedding theorem, the paper deals with gauge theories and examples like Electrodynamics and Klein-Gordon theory.
result Emergence of the energy-momentum tensor algebra of conserved currents.

We begin an exploration of parametric Backlund transformations for hyperbolic Monge-Ampere systems. We compute invariants for such transformations and explore the behavior of four examples regarding their invariants, symmetries, and conservation laws. We prove some preliminary results and indicate directions for furthe…

2002-08-05abs ↗pdf ↗

The problem of multi-armed bandits (MAB) asks to make sequential decisions while balancing between exploitation and exploration, and have been successfully applied to a wide range of practical scenarios. Various algorithms have been designed to achieve a high reward in a long term. However, its short-term performance m…

2019-11-26abs ↗pdf ↗

Safe exploration method for RL under disturbance ensures safety with probabilistic guarantees.

problem Safe reinforcement learning in real environments with disturbance.
method Uses partial prior knowledge and conservative inputs to ensure state constraint satisfaction.
result Guaranteed safety with pre-specified probability in the presence of stochastic disturbance.

In many fields such as digital marketing, healthcare, finance, and robotics, it is common to have a well-tested and reliable baseline policy running in production (e.g., a recommender system). Nonetheless, the baseline policy is often suboptimal. In this case, it is desirable to deploy online learning algorithms (e.g.,…

2020-02-08abs ↗pdf ↗

VaR-CPO optimizes VaR-constrained RL problems with conservative policy updates.

problem Optimizing VaR-constrained reinforcement learning problems.
method Combines Cantelli's inequality and trust-region framework for efficient and conservative optimization.
result Achieves zero constraint violations during training in feasible environments.

We explore the geometry of the Nahm-Schmid equations, a version of Nahm's equations in split signature. Our discussion ties up different aspects of their integrable nature: dimensional reduction from the Yang--Mills anti-self-duality equations, explicit solutions, Lax-pair formulation, conservation laws and spectral cu…

2017-11-07abs ↗pdf ↗

Paper explores weak solutions' regularity in critical dimensions without conservation law.

problem Regularity of weak solutions to higher order elliptic systems in critical dimensions.
method Elementary and unified treatment, without conservation law.
result Interior Hölder continuity for solutions in critical dimensions.

BRAID fine-tunes diffusion models to optimize reward models in offline scenarios.

problem Combining generative modeling and model-based optimization in offline scenarios.
method Conservative fine-tuning of diffusion models using RL to optimize reward models.
result BRAID outperforms existing methods in offline data, avoiding invalid designs.

We present a class of macroscopic models of the Limit Order Book to simulate the aggregate behaviour of market makers in response to trading flows. The resulting models are solved numerically and asymptotically, and a class of similarity solutions linked to order book formation and recovery is explored. The main result…

2019-10-21abs ↗pdf ↗

The abstract explores a new wave equation linking quantum mechanics and complex adaptive systems.

problem Understanding the underlying mechanism of distribution formation in complex quantum entanglement.
method Exploring the logical relationship between Schrödinger's wave equation and Shi's trading volume-price wave equation in finance.
result A non-localized wave equation in quantum mechanics reveals the invariance of interaction as a universal law.

Theoretical analysis confirms non-conservative algorithms can converge to optimal policies.

problem Theoretical guarantees for non-conservative reinforcement learning algorithms.
method Theoretical analysis of Peng's Q(λλ) algorithm.
result Peng's Q(λλ) converges to an optimal policy under certain conditions.

Study finds conserved quantities for two types of curves on conformal sphere.

problem Identifying conserved quantities for specific types of curves on a conformal sphere.
method Used parallel tractor and Lagrangian formalism to compute conserved quantities.
result Found relation between conserved quantities of two curve types.

The paper optimizes LLM inference systems through queueing theory.

problem Efficient LLM inference for AI agents under various routing topologies.
method Developed a fluid-limit framework for multi-class batched processing networks under K-FCFS scheduling.
result Proved that work-conserving scheduling algorithms maximize throughput for LLM inference.

Paper presents a reduction-based framework for conservative bandits and RL with improved lower and upper bounds.

problem Conservative bandits and reinforcement learning problems.
method Reduction technique to calculate necessary and sufficient budget from baseline policy.
result Improved lower and upper bounds for various conservative settings.

Auto-CEI improves LLM reasoning by balancing assertiveness and conservativeness.

problem Hallucinations and laziness in LLM reasoning tasks.
method Expert Iteration explores reasoning trajectories, guiding incorrect paths back on track and promoting appropriate 'I don't know' responses.
result Auto-CEI achieves superior alignment in logical reasoning, mathematics, and planning tasks.

Real-world applications require RL algorithms to act safely. During learning process, it is likely that the agent executes sub-optimal actions that may lead to unsafe/poor states of the system. Exploration is particularly brittle in high-dimensional state/action space due to increased number of low-performing actions. …

2019-02-23abs ↗pdf ↗

Given a vector field on a manifold M, we define a globally conserved quantity to be a differential form whose Lie derivative is exact. Integrals of conserved quantities over suitable submanifolds are constant under time evolution, the Kelvin circulation theorem being a well-known special case. More generally, conserved…

2016-10-18abs ↗pdf ↗

We study higher-order conservation laws of the non-linearizable elliptic Poisson equation 2uzzˉ=f(u) \frac{{\partial}^2 u}{\partial z \partial \bar{z}} = -f(u) as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…

2009-06-17abs ↗pdf ↗

New conservation laws found for polyharmonic maps in critical dimension.

problem Existence of conservation laws for polyharmonic maps in critical dimension.
method Small perturbation of Uhlenbeck's gauge fixing matrix.
result Existence of conservation laws for elliptic systems of even order in critical dimension.

The paper studies symmetries and conservation laws of non-diagonalisable hydrodynamic systems.

problem Integrating non-diagonalisable hydrodynamic systems of partial differential equations.
method Analysis of gl-regular Nijenhuis operators, splitting Theorem for symmetries and conservation laws, relationship between symmetries and conservation laws.
result The system of partial differential equations is integrable in quadratures.

We present a connection between the Killing fields that arise in the loop-group approach to integrable systems and conservation laws viewed as elements of the characteristic cohomology. We use the connection to generate the complete set of conservation laws (as elements of the characteristic cohomology) for the Tzitzei…

2012-08-13abs ↗pdf ↗

New neural network enforces mass conservation for better ice flow predictions.

problem Reliably project future sea level rise by improving ice sheet model inputs.
method Proposes divergence-free neural networks (dfNNs) enforcing local mass conservation.
result dfNNs yield more reliable ice flux estimates compared to other models.

The study finds resonance points in polarised curves with polynomial conserved quantities.

problem Finding resonance points in polarised curves with polynomial conserved quantities.
method Using the non-orthogonality assumption on the conserved quantity, the study deduces the existence of resonance points.
result Every finite type polarised curve in the conformal 2-sphere with a polynomial conserved quantity admits a resonance point.

Following an approach of the second author for conformally invariant variational problems in two dimensions, we show in four dimensions the existence of a conservation law for fourth order systems, which includes both intrinsic and extrinsic biharmonic maps. With the help of this conservation law we prove the continuit…

2006-07-20abs ↗pdf ↗