A distributed system identification method for LTI systems using reverse experience replay.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the problem of identifying the policy space of a learning agent, having access to a set of demonstrations generated by its optimal policy. We introduce an approach based on statistical testing to identify the set of policy parameters the agent can control, within a larger parametric policy space. After present…
Efficiently identifies best algorithms for game tasks.
A new algorithm for identifying the best arm in linear feedback with safety constraints.
New approach incentivizes strategic agents to explore, making exploration almost free.
Perceiving the surrounding environment in terms of objects is useful for any general purpose intelligent agent. In this paper, we investigate a fundamental mechanism making object perception possible, namely the identification of spatio-temporally invariant structures in the sensorimotor experience of an agent. We take…
Paper shows faster core identification in matching markets.
A new method for identifying the best arm in multi-armed bandits with mediator feedback.
Successful human-robot cooperation hinges on each agent's ability to process and exchange information about the shared environment and the task at hand. Human communication is primarily based on symbolic abstractions of object properties, rather than precise quantitative measures. A comprehensive robotic framework thus…
Modeling agent behavior is central to understanding the emergence of complex phenomena in multiagent systems. Prior work in agent modeling has largely been task-specific and driven by hand-engineering domain-specific prior knowledge. We propose a general learning framework for modeling agent behavior in any multiagent …
Optimal algorithm found for collaborative learning in bandits with optimal regret bounds.
A network of agents attempt to learn some unknown state of the world drawn by nature from a finite set. Agents observe private signals conditioned on the true state, and form beliefs about the unknown state accordingly. Each agent may face an identification problem in the sense that she cannot distinguish the truth in …
Quantum algorithm solves best arm identification problem faster.
New algorithm achieves near optimal sample complexity for 1-identification problem.
Bayesian algorithm improves best-arm identification within fixed budget.
New method for identifying best arm in batched multi-armed bandit problems.
Study best arm identification with safety constraints in bandit problems.
The agent-based model of stock price dynamics on a directed evolving complex network is suggested and studied by direct simulation. The stationary regime is maintained as a result of the balance between the extremal dynamics, adaptivity of strategic variables and reconnection rules. The inherent structure of node agent…
A new bandit strategy improves learning speed at the cost of accuracy.
The paper tackles best arm identification with minimal regret in experiments.
Exploration in multi-task reinforcement learning is critical in training agents to deduce the underlying MDP. Many of the existing exploration frameworks such as , , Thompson sampling assume a single stationary MDP and are not suitable for system identification in the multi-task setting. We present a nove…
We consider a novel stochastic multi-armed bandit problem called {\em good arm identification} (GAI), where a good arm is defined as an arm with expected reward greater than or equal to a given threshold. GAI is a pure-exploration problem that a single agent repeats a process of outputting an arm as soon as it is ident…
Extracts intrinsic spatial coordinates for complex agent systems to learn PDEs.
We show that there is a common mode of origin for the power laws observed in two different models: (i) the Pareto law for the distribution of money among the agents with random saving propensities in an ideal gas-like market model and (ii) the Gutenberg-Richter law for the distribution of overlaps in a fractal-overlap …
Best arm identification (or, pure exploration) in multi-armed bandits is a fundamental problem in machine learning. In this paper we study the distributed version of this problem where we have multiple agents, and they want to learn the best arm collaboratively. We want to quantify the power of collaboration under limi…
Model improves email-based conversational agents' ability to extract relevant information.
New framework recovers reward and rationality parameters from game behavior.
In this paper, we study the system identification problem for sparse linear time-invariant systems. We propose a sparsity promoting block-regularized estimator to identify the dynamics of the system with only a limited number of input-state data samples. We characterize the properties of this estimator under high-dimen…
Study non-asymptotic BPI guarantees for online RL.
In this report we examine the effectiveness of WISER in identification of a chemical culprit during a chemical based Mass Casualty Incident (MCI). We also evaluate and compare Binary Decision Tree (BDT) and Artificial Neural Networks (ANN) using the same experimental conditions as WISER. The reverse engineered set of S…
Algorithm identifies best arm in piecewise stationary linear bandits with minimal samples.
MarketSenseAI system outperforms passive benchmarks by 25.2% on S&P 500, adding value over random selection.
ADR helps LLMs find and use historical analogies for foresight analysis.
Algorithm learns new tasks efficiently from past experience.
Common event-triggered state estimation (ETSE) algorithms save communication in networked control systems by predicting agents' behavior, and transmitting updates only when the predictions deviate significantly. The effectiveness in reducing communication thus heavily depends on the quality of the dynamics models used …
We propose a generalization of the best arm identification problem in stochastic multi-armed bandits (MAB) to the setting where every pull of an arm is associated with delayed feedback. The delay in feedback increases the effective sample complexity of standard algorithms, but can be offset if we have access to partial…
Study simulates liquidity in fractional ownership markets using ABM.
AdaptOn achieves logarithmic regret in adaptive control of unknown partially observable linear systems.
In this thesis, we develop a comprehensive account of the expressive power, modelling efficiency, and performance advantages of so-called trading agents (i.e., Deep Soft Recurrent Q-Network (DSRQN) and Mixture of Score Machines (MSM)), based on both traditional system identification (model-based approach) as well as on…
AE-LSVI identifies near-optimal policies in complex systems with minimal data.
Recent advances in computing power and the potential to make more realistic assumptions due to increased flexibility have led to the increased prevalence of simulation models in economics. While models of this class, and particularly agent-based models, are able to replicate a number of empirically-observed stylised fa…
Study quantile reward identification with 1-bit feedback constraints.
Study reveals investor behavior in NFT bubbles.
CGAs estimate team performance from data, simplifying SV computation.
Toward enabling next-generation robots capable of socially intelligent interaction with humans, we present a of interactions in a social environment of multiple agents and multiple groups. The Multiagent Group Perception and Interaction (MGpi) network is a deep neural network that predi…
New exploration bonuses improve reinforcement learning efficiency.
Geometric methods solve sampling, optimisation, inference, and adaptive decision-making.
Algorithm identifies optimal stable matching in uncertain two-sided markets.