Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4488131175 · Jun 202019922001200920172026
48 results for closed-loop behavior

Analyzes how learning algorithms affect and are affected by data manipulation.

problem Characterizing the closed-loop behavior of learning algorithms in the presence of decision-dependent data.
method Analyzes repeated risk minimization as perturbed gradient flows of performative risk minimization, considering multiple local minimizers.
result Characterizes the region of attraction for various equilibria and introduces performative alignment.

Framework learns robust control policies from expert demonstrations.

problem Adversarial robustness and closed-loop generalization in feedback control policies.
method Lipschitz-constrained loss minimization for certified robustness and generalization.
result Finite sample bound on policy learning error and robust closed-loop stability.

Paper predicts recycling bin full events to reduce RVM downtime.

problem Predicting bin full events to increase RVM uptime.
method Hybrid approach combining machine learning and statistical approximation.
result Forecasting leads to less downtime and costs compared to emptying strategies.

Investor selects portfolios based on news attention in a hidden Markov model.

problem Mean-variance portfolio selection in a dynamic attention context.
method Closed-loop equilibrium strategies via extended HJB equation and Markov chain approximation.
result Equilibrium strategies found through iterative algorithm and numerical examples.

Study on parameter dynamics in exponential families under closed-loop learning.

problem Understanding and preventing convergence to biased absorbing states in model parameter estimation.
method Derived equations of motion for parameter dynamics in exponential families. Showed convergence to biased states under maximum likelihood estimation and proposed solutions.
result Closed-loop learning can lead to biased parameter estimates, but can be mitigated by using maximum a posteriori estimation or regularisation.

Geometrically characterizes virtual nonlinear nonholonomic constraints using symplectic methods.

problem Characterizing virtual nonlinear nonholonomic constraints geometrically.
method Geometric characterization using symplectic structures and Chetaev equations.
result A unique control law exists to satisfy virtual constraints, and closed-loop dynamics are projections of uncontrolled dynamics.

Formula adjusts steady-state models for control confounding.

problem Learning steady-state models from operational data can be flawed due to control confounding.
method Derives a formula to adjust for control confounding using structural dynamical causal models.
result Estimates a causal steady-state model from closed-loop operational data.

We show LLMs can be locally linear, enabling better control of activations.

problem Suboptimal control of LLM activations during generation.
method Model LLM inference as a linear dynamical system, compute feedback controllers using Jacobians, and adapt classical control theory.
result Robust, fine-grained control of LLM activations across models and tasks.

Behavior cloning training instabilities amplified by SGD noise over long horizons.

problem Training instabilities in behavior cloning with deep neural networks.
method Empirical dissection of minibatch SGD updates and their effects on long-horizon rewards.
result Exponential moving average (EMA) of iterates effectively mitigates gradient variance amplification (GVA).

The paper analyzes strategic irreversible investments with novel dynamic strategies.

problem Tradeoff between preemption incentives and option value of waiting in oligopolistic markets.
method Developed novel Markov perfect equilibrium to handle singular control of optimal investment.
result Simpler strategies lead to a 'preemption trap' with zero net present values.

Data-efficient learning in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems. In this paper, we consider one instance of this challenge, the pixels to torques problem, where an agent must learn a closed-loop control policy from pixel i…

2015-02-08abs ↗pdf ↗

New method evaluates AI stock prediction systems based on decision-making processes.

problem Lack of evaluation for AI systems' decision-making processes.
method Scores intermediate decision process using large language models and closed-loop reinforcement learning feedback.
result Composite behavioral score correlates with Sharpe ratio and reduces prediction error.

End-to-end learnable network for safer self-driving with interpretable intermediate representations.

problem Safe motion planning for self-driving vehicles.
method Differentiable semantic occupancy representation for cost calculation in motion planning.
result Significantly outperforms state-of-the-art planners in imitating human behaviors and producing safer trajectories.

This work discusses a closed-loop control strategy for complex systems utilizing scarce and streaming data. A discrete embedding space is first built using hash functions applied to the sensor measurements from which a Markov process model is derived, approximating the complex system's dynamics. A control strategy is t…

2016-04-11abs ↗pdf ↗

This paper uses NLDT to find interpretable control rules from complex DRL policies.

problem Complex, non-interpretable policies from black-box AI methods.
method Evolutionary optimization of NLDT for hierarchical control rules.
result Interpretable control rules with similar performance to black-box DRL.

Agents acting in the natural world aim at selecting appropriate actions based on noisy and partial sensory observations. Many behaviors leading to decision mak- ing and action selection in a closed loop setting are naturally phrased within a control theoretic framework. Within the framework of optimal Control Theory, o…

2014-06-27abs ↗pdf ↗

CLQT benchmarks LLM portfolio managers by evaluating their decision-making process, not just returns.

problem Most benchmarks rank LLMs by returns, ignoring their decision-making process and potential for look-ahead leakage.
method CLQT reframes evaluation as diagnosis, using a closed-loop, cost-aware, strategy-consistent environment with a five-stage cycle.
result CLQT provides a durable map of agent competencies and limitations, separating outcome from process.

DQN outperforms static policies in a dynamic fee environment for automated market makers.

problem How automated market makers (AMMs) perform under dynamic fees is unknown.
method Constructed a closed-loop simulator with dynamic fees, noise flow, and arbitrage.
result A small DQN policy outperforms static policies in a dynamic fee environment.

Paper tackles stochastic control with mean and higher-order moments, finding Nash equilibria.

problem Time-inconsistent stochastic control problems with mean and higher-order moments.
method Developed closed-loop and open-loop Nash equilibrium controls using PDEs and maximum principles.
result Identical closed-loop and open-loop Nash equilibria controls, independent of state value and random path.

A geometric approach to differential game theory is illustrated. The parallel pursuit is considered as a two-player zero-sum differential game. The optimal strategies of each player is designed based on Riemann-Finsler geometry. Our approach incorporates a closed loop optimal control and the presentation is familiar wi…

2011-01-10abs ↗pdf ↗

New budget quantifies drift in closed-loop learning, improving reproducibility.

problem Characterizing statistical learning under distributional drift in closed-loop settings.
method Introduces an intrinsic drift budget CTC_T quantifying cumulative information-geometric motion of the data distribution.
result Proves a drift-feedback bound of order T1/2+CT/TT^{-1/2}+C_T/T for prequential reproducibility, up to controlled second-order remainder terms.

We present an alternative local definition of the writhe of a self-avoiding closed loop which differs from the traditional non-local definition by an integer. When studying dynamics this difference is immaterial. We employ a formula due to Aldinger, Klapper and Tabor for the change in writhe and propose a set of local,…

1997-03-13abs ↗pdf ↗

The paper tackles performative risk optimization under weak convexity assumptions.

problem Optimizing performative risk in a closed-loop prediction system with weak convexity.
method Relaxing convexity assumptions to maintain optimization feasibility.
result Iterative optimization methods remain applicable even with weakened convexity conditions.

Filling length measures the length of the contracting closed loops in a null-homotopy. The filling length function of Gromov for a finitely presented group measures the filling length as a function of length of edge-loops in the Cayley 2-complex. We give a bound on the filling length function in terms of the log of an …

2000-08-03abs ↗pdf ↗

Sequential learning of tasks using gradient descent leads to an unremitting decline in the accuracy of tasks for which training data is no longer available, termed catastrophic forgetting. Generative models have been explored as a means to approximate the distribution of old tasks and bypass storage of real data. Here …

2018-11-03abs ↗pdf ↗

Combines Gaussian processes and polynomial chaos for stochastic control.

problem Uncertainties in dynamic models lead to performance issues in predictive control.
method Combines Gaussian processes with polynomial chaos expansions to estimate probability distributions of nonlinear functions.
result Demonstrates accurate approximation and closed-loop performance in stochastic nonlinear model predictive control.

We propose a sliding surface for systems on the Lie group SO(3)×R3SO(3)\times \mathbb{R}^3 . The sliding surface is shown to be a Lie subgroup. The reduced-order dynamics along the sliding subgroup have an almost globally asymptotically stable equilibrium. The sliding surface is used to design a sliding-mode controller for t…

2019-05-14abs ↗pdf ↗

We utilize Wi-Fi communications from smartphones to predict their mobility mode, i.e. walking, biking and driving. Wi-Fi sensors were deployed at four strategic locations in a closed loop on streets in downtown Toronto. Deep neural network (Multilayer Perceptron) along with three decision tree based classifiers (Decisi…

2018-09-16abs ↗pdf ↗

Investigates time-inconsistent portfolio selection under MMV preferences.

problem Time-inconsistent optimal strategies for MMV preferences.
method Nash equilibrium controls for MMV and MV preferences, solving FBSDE and HJB equations.
result MMV optimal strategies lead to higher investment amounts than MV strategies, narrowing over time.