Collective behavior of the complex socio-economic systems is heavily influenced by the herding, group, behavior of individuals. The importance of the herding behavior may enable the control of the collective behavior of the individuals. In this contribution we consider a simple agent-based herding model modified to inc…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extends driving model to control agent behavior in simulations.
Research explores how interconnected systems synchronize and how to control their behavior.
Investigates optimal strategies for behavioral control problems with finite variation controls.
Paper uses RL for high-level character control in 3D environments.
New method disentangles perceptual uncertainty and behavioral costs in partially observable systems.
We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the well-studied areas of controllable generation of images, text, and speech, there are two questions that pose significant challenges when genera…
A new framework for adaptive behavior using reusable value profiles.
New method infers human sensorimotor costs from behavior.
Unified Bayesian model explains in-context learning and activation steering in LLMs.
Paper develops a framework for learning interpretable representations of sequential decision behavior.
Action chunking and data exploration improve behavior cloning in robotics.
What is the role of real-time control and learning in the formation of social conventions? To answer this question, we propose a computational model that matches human behavioral data in a social decision-making game that was analyzed both in discrete-time and continuous-time setups. Furthermore, unlike previous approa…
Paper proposes a method to improve off-policy reinforcement learning in batch settings.
Young investors, especially students, dominate Indonesian stock exchanges.
A decentralized deep RL controller improves hexapod locomotion learning.
This paper contains a summary of mathematical researches of stochastic properties of the long time behavior of a continuously observed (and interactively controlled) quantum--field top. Applications to interactively controlled stochastic computer-graphic dynamical systems are also discussed.
This paper provides a full controlled version of algebraic -theory. This includes a rich array of assembly maps; the controlled assembly isomorphism theorem identifying the controlled group with homology; and the stability theorem describing the behavior of the inverse limit as the control parameter goes to 0. There…
Motivated by the ubiquity of control-affine systems in optimal control theory, we investigate the geometry of point-affine control systems with metric structures in dimensions two and three. We compute local isometric invariants for point-affine distributions of constant type with metric structures for systems with 2 s…
RFC enhances humanoid control to imitate complex human motions.
We study the inverse optimal control problem in social sciences: we aim at learning a user's true cost function from the observed temporal behavior. In contrast to traditional phenomenological works that aim to learn a generative model to fit the behavioral data, we propose a novel variational principle and treat user …
Data poisoning is an attack on machine learning models wherein the attacker adds examples to the training set to manipulate the behavior of the model at test time. This paper explores poisoning attacks on neural nets. The proposed attacks use "clean-labels"; they don't require the attacker to have any control over the …
PWIL learns agent behavior from expert using Wasserstein distance.
ComiRec framework predicts user interests for personalized recommendations.
Weakly-supervised RL identifies meaningful tasks, improving performance in complex environments.
As energy markets begin clearing at sub-hourly rates, their interaction with load control systems becomes a potentially important consideration. A simple model for the control of thermal systems using market-based power distribution strategies is proposed, with particular attention to the behavior and dynamics of elect…
We study an optimal investment control problem for an insurance company. The surplus process follows the Cramer-Lundberg process with perturbation of a Brownian motion. The company can invest its surplus into a risk free asset and a Black-Scholes risky asset. The optimization objective is to minimize the probability of…
Derives time-averaged active inference from control principles.
Lane change is a challenging task which requires delicate actions to ensure safety and comfort. Some recent studies have attempted to solve the lane-change control problem with Reinforcement Learning (RL), yet the action is confined to discrete action space. To overcome this limitation, we formulate the lane change beh…
This paper provides theoretical foundations for using quantized actions in behavior cloning.
In this paper, we show the implementation of deep neural networks applied in process control. In our approach, we based the training of the neural network on model predictive control. Model predictive control is popular for its ability to be tuned by the weighting matrices and by the fact that it respects the constrain…
We make policy optimization algorithms batch size-invariant by decoupling proximal and behavior policies.
Framework for controlling multiple risks in AI models.
We investigate necessary and sufficient conditions under which a general nonlinear affine control system with outputs can be written as a gradient control system corresponding to some pseudo-Riemannian metric defined on the state space. The results rely on a suitable notion of compatibility of the system with respect t…
The paper provides theoretical guarantees for behavior cloning using generative models.
We study a risk sensitive control version of the lifetime ruin probability problem. We consider a sequence of investments problems in Black-Scholes market that includes a risky asset and a riskless asset. We present a differential game that governs the limit behavior. We solve it explicitly and use it in order to find …
Paper proposes personalized climate control for driver comfort.
In this paper, we propose a decision making algorithm intended for automated vehicles that negotiate with other possibly non-automated vehicles in intersections. The decision algorithm is separated into two parts: a high-level decision module based on reinforcement learning, and a low-level planning module based on mod…
Reinforcement learning mimics expert behavior.
Agents acting in the natural world aim at selecting appropriate actions based on noisy and partial sensory observations. Many behaviors leading to decision mak- ing and action selection in a closed loop setting are naturally phrased within a control theoretic framework. Within the framework of optimal Control Theory, o…
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise. In particular, both the faster and slower recursions have non-additive controlled Markov noise components in addition to martingale difference noise. We analyze the asymptotic…
In a voice-controlled smart-home, a controller must respond not only to user's requests but also according to the interaction context. This paper describes Arcades, a system which uses deep reinforcement learning to extract context from a graphical representation of home automation system and to update continuously its…
This paper presents a hierarchical framework for Deep Reinforcement Learning that acquires motor skills for a variety of push recovery and balancing behaviors, i.e., ankle, hip, foot tilting, and stepping strategies. The policy is trained in a physics simulator with realistic setting of robot model and low-level impeda…
We describe an approach to understand the peculiar and counterintuitive generalization properties of deep neural networks. The approach involves going beyond worst-case theoretical capacity control frameworks that have been popular in machine learning in recent years to revisit old ideas in the statistical mechanics of…
Exchange uses incentives to optimize limit order book dynamics.
We solve a class of control problems with fuel constraint by means of the log-Laplace transforms of -functionals of Dawson-Watanabe superprocesses. This solution is related to the superprocess solution of quasilinear parabolic PDEs with singular terminal condition. For the probabilistic verification proof, we develo…
Combines Lyapunov functions with controller synthesis for safe control policies.
Humans are able to perform a myriad of sophisticated tasks by drawing upon skills acquired through prior experience. For autonomous agents to have this capability, they must be able to extract reusable skills from past experience that can be recombined in new ways for subsequent tasks. Furthermore, when controlling com…