Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

82164245327 · Jun 202019922001200920172026
48 results for history-dependent systems

Volterra signature provides a clear, interpretable feature for history-dependent systems.

problem Learning from non-Markovian time series with implicit memory mechanisms.
method Develops Volterra signature as a tensor algebra representation weighted by a temporal kernel, proving injectivity and universal approximation.
result Volterra signature leads to linear functionals and universal approximation, improving dynamic learning tasks.

The paper explains why estimating a history-dependent policy can reduce MSE in reinforcement learning.

problem Understanding why history-dependent policies can improve MSE in off-policy evaluation.
method The paper derives a bias-variance decomposition of MSE for various OPE estimators, showing how history-dependent policies can decrease variance and increase bias.
result History-dependent policies can decrease the variance of importance sampling estimators, leading to lower MSE.

Study designs incentives for adapting multi-agent systems without knowing their learning dynamics.

problem Designing incentives for an adapting population in multi-agent systems without prior knowledge of their learning dynamics.
method Introduces a model-based non-episodic Reinforcement Learning (RL) formulation for steering Markovian agents towards desired policies, focusing on history-dependent strategies to handle model uncertainty.
result Identifies conditions for the existence of steering strategies to guide agents to desired policies and provides empirical algorithms to approximately solve the objective.

Neural Laplace models diverse DEs in the Laplace domain for better dynamics.

problem Inadequate ODEs for long-range dependencies and discontinuities.
method Unified framework in Laplace domain, using stereographic map for smoothness.
result Superior performance in diverse DEs, including complex history dependency and abrupt changes.

Rhino learns causal relationships from time series data with history-dependent noise.

problem Discovering causal relationships from time series data with non-linear relations, instantaneous effects, and history-dependent noise.
method Combines vector auto-regression, deep learning, and variational inference.
result Demonstrates better causal relationship discovery performance compared to baselines.

SRMC framework reduces Monte Carlo variance by history-based sampling in high-dimensional spaces.

problem Efficient sampling in high-dimensional discrete or continuous state spaces.
method Score-Repellent Monte Carlo (SRMC) framework that summarizes history through running average of score evaluations.
result Improves estimator variance and mode coverage with constant memory usage.

Paper proposes a policy gradient method for confounded POMDPs.

problem Estimating policy gradients for confounded POMDPs with continuous state and observation spaces.
method Developed a novel identification result to estimate policy gradients using offline data, solved conditional moment restrictions, and applied min-max learning with function approximation.
result Showed global convergence of the proposed algorithm in finding the optimal policy.

Unified analytical tool for non-Markovian jump processes.

problem Analyzing history-dependent jump processes with non-Markovian behavior.
method Developed a standard form of master equations using Laplace-space embedding and asymptotic solution.
result Unified analytical toolset for general non-Markovian processes, leading to the GLE approximation.

A data-driven approach predicts morphological development under structural instability.

problem Understanding and predicting spatiotemporal complexities of morphogenesis under structural instability.
method Machine-learning framework based on physical modeling of morphogenesis.
result Identification of key bifurcation characteristics and prediction of history-dependent development.

Trust is a collective, self-fulfilling phenomenon that suggests analogies with phase transitions. We introduce a stylized model for the build-up and collapse of trust in networks, which generically displays a first order transition. The basic assumption of our model is that whereas trust begets trust, panic also begets…

2014-09-22abs ↗pdf ↗

Framework preserves emergent physics in non-equilibrium systems from particle trajectories.

problem Linking short spatiotemporal scales to emergent bulk physics in multiscale systems.
method Metriplectic bracket formalism for structure-preserving coarse-graining.
result Preservation of thermodynamic laws and conservation in machine-learned dynamics.

HDT improves MCMC on graphs with history-dependent sampling.

problem Efficient sampling from target distributions on general graphs with low computational overhead.
method History-driven target (HDT) framework that replaces the original target distribution with a history-dependent one.
result Near-zero variance performance and scalability to large graphs with memory-efficient implementation.

New framework for robust reinforcement learning policies in uncertain environments.

problem Robust reinforcement learning policies in environments with distributional shifts.
method Comprehensive modeling framework centered around robust Markov decision processes (RMDPs).
result Existence and conditions for the dynamic programming principle (DPP) in RMDPs.

The paper examines how loss aversion impacts multi-armed bandit decisions over long periods.

problem The impact of loss aversion on multi-armed bandit decisions over long periods.
method A new central limit theorem for measures with history-dependent variances, derived under risk aversion in gains and risk loving in losses.
result Consequences of loss aversion for asymptotic properties are derived in analytical results.

In many settings (e.g., robotics) demonstrations provide a natural way to specify tasks; however, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for the tasks, such as rewards or policies, can be safely composed and/or do not explicitly capture history dependen…

2019-07-26abs ↗pdf ↗

A scalable framework uses Langevin sampling to approximate neural network models of evolving processes.

problem Uncertainty quantification in neural network models of dynamic systems.
method Flexible data model based on NODE, joint learning of data model and posterior parameters, Langevin sampling.
result Demonstrated performance on chemical reaction and material physics data, compared favorably to variational inference.

The paper interprets policy-gradient algorithms using continuation theory.

problem Optimizing nonconvex functions in reinforcement learning.
method Formulates policy optimization as optimization by continuation, interprets policy-gradient algorithms as implicitly optimizing deterministic policies.
result Exploration in policy-gradient algorithms is seen as computing a continuation of the return of the policy.

L2O-CFGD meta-learns hyperparameters for FGD, improving performance.

problem Challenges in convergence and hyperparameter selection for FGD.
method Learning to Optimize Caputo Fractional Gradient Descent (L2O-CFGD).
result Meta-learned schedule outperforms static hyperparameters and achieves comparable performance to black-box meta-learners.

Flexible Hawkes model with Gaussian process self-effects for time-dependent data.

problem Modeling time-dependent point processes with history dependence and self-effects.
method Extended Hawkes process with Gaussian process self-effects for both excitatory and inhibitory types, using Bayesian inference and mean-field variational approximation.
result Efficient approximate Bayesian inference achieved via data augmentation and mean-field variational approach.

New method improves treatment effect estimation in adaptive experiments with noncompliance.

problem Estimating average treatment effect in adaptive experiments with binary instrumental variable.
method AMRIV estimator that balances outcome noise and compliance variability.
result AMRIV achieves semiparametric efficiency bound and is robust to noncompliance.

Bootstrap method for Markov chains in reinforcement learning.

problem Distributional consistency in finite controlled Markov chains with unknown control policies.
method Model-based bootstrap with novel LLN and CLT for visitation counts and transition increments.
result Asymptotically valid confidence intervals for value and QQ-functions in offline RL.

We propose a general framework to describe the impact of different events in the order book, that generalizes previous work on the impact of market orders. Two different modeling routes can be considered, which are equivalent when only market orders are taken into account. One model posits that each event type has a te…

2011-07-18abs ↗pdf ↗

Study non-rectangular robust MDPs for average-reward, finding optimal policies and transient values.

problem Non-rectangular robust Markov decision processes under average-reward criterion.
method Proves history-dependent policies are robust-optimal, introduces transient-value framework, constructs epoch-based policy.
result Existence and properties of robust optimal policies, transient value bounds.

Overview of integrable systems with symmetries, focusing on toric and semitoric systems.

problem Classifying and understanding integrable systems with symmetries.
method Using decorated polygons and controlled bifurcations in one-parameter families of systems.
result Construction of explicit semitoric systems with prescribed invariants.

Learning to control linear systems is statistically hard, especially for underactuated systems.

problem Statistical difficulty of learning to control linear systems, especially underactuated ones.
method Utilized minimax lower bounds and structural assumptions to prove learning complexity can be exponential.
result Learning complexity can be at most exponential with the controllability index of the system.

Discrete-time systems can be characterized by simple flat coordinates and their shifts.

problem Characterizing flatness of discrete-time systems.
method Developed a map from flat coordinates and their shifts to system state and input, fulfilling system equations identically.
result Derived necessary conditions for a system to be flat, without requiring differential geometry methods.

The paper explores when linear system identification is hard or easy, especially for under-actuated systems.

problem Statistical hardness of learning linear systems, especially under-actuated or under-excited systems.
method Using tools from minimax theory and recent statistical tools for finite sample analysis of system identification.
result The controllability index of linear systems affects the sample complexity of identification, making some systems hard to learn.

This paper improves system identification by reducing sample complexity for high-dimensional linear dynamical systems.

problem High sample complexity for learning partially observed linear dynamical systems in high dimensions.
method Introduces an 1\ell_1-regularized estimation method that reduces sample complexity from linear to logarithmic with system dimension.
result Markov parameters can be learned with logarithmic number of samples relative to system dimension, improving sample complexity.

In integrable hydrodynamic systems, coordinates exist where generators and symmetries are simple.

problem Existence of Riemannian invariants for integrable systems of hydrodynamic type.
method Finding coordinates where the generator and all symmetries are diagonal.
result In integrable hydrodynamic systems, there exist coordinates where the generator and all symmetries are diagonal.

This paper studies nonholonomic constraints in Hamiltonian systems, deriving equations and theorems.

problem Analyzing nonholonomic constraints in Hamiltonian systems.
method Deriving distributional RCH systems, geometric constraint conditions, and Hamilton-Jacobi theorems.
result Derives precise geometric constraint conditions and Hamilton-Jacobi theorems for nonholonomic systems.

Estimates input from output of nonlinear systems using ANN.

problem Estimating unknown compositional input from system output.
method Artificial Neural Networks (ANNs) for nonlinear system inversion.
result ANNs can compete with optimal bounds for linear systems and demonstrate promising results for nonlinear systems.

This paper considers control systems defined on Lie algebroids. After deriving basic controllability tests for general control systems, we specialize our discussion to the class of mechanical control systems on Lie algebroids. This class of systems includes mechanical systems subject to holonomic and nonholonomic const…

2004-02-26abs ↗pdf ↗