Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alternately optimizes, two policies: a fast, reactive policy (e.g., a deep neural network) deployed at t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Algorithm improves reinforcement learning policies using offline data.
MPC outperforms reactive budgeting in non-stationary return environments.
Designs a single policy for collecting data to train near-optimal policies.
Decouples critic chunk length from policy to improve policy reactivity and performance.
We introduce Recurrent Predictive State Policy (RPSP) networks, a recurrent architecture that brings insights from predictive state representations to reinforcement learning in partially observable environments. Predictive state policy networks consist of a recursive filter, which keeps track of a belief about the stat…
RL optimizes meta-order execution by adapting to market conditions.
A new framework uses stochastic optimal control to estimate rare events more accurately.
Many robotic applications require the agent to perform long-horizon tasks in partially observable environments. In such applications, decision making at any step can depend on observations received far in the past. Hence, being able to properly memorize and utilize the long-term history is crucial. In this work, we pro…
Improved queue-reactive model considers order sizes for better market simulation.
A new algorithm improves efficiency and robustness of heuristic optimization in simulation-based problems.
Enhances queue-reactive model for realistic limit order book simulation.
Pronounced variability due to the growth of renewable energy sources, flexible loads, and distributed generation is challenging residential distribution systems. This context, motivates well fast, efficient, and robust reactive power control. Real-time optimal reactive power control is possible in theory by solving a n…
We present a new volatility model, simple to implement, that includes a leverage effect whose return-volatility correlation function fits to empirical observations. This model is able to capture both the "retarded effect" induced by the specific risk, and the "panic effect", which occurs whenever systematic risk become…
We present a reactive beta model that includes the leverage effect to allow hedge fund managers to target a near-zero beta for market neutral strategies. For this purpose, we derive a metric of correlation with leverage effect to identify the relation between the market beta and volatility changes. An empirical test ba…
Unified model for market dynamics, linking price and order flow.
Electronic power inverters are capable of quickly delivering reactive power to maintain customer voltages within operating tolerances and to reduce system losses in distribution grids. This paper proposes a systematic and data-driven approach to determine reactive power inverter output as a function of local measuremen…
We propose and study a new model for reinforcement learning with rich observations, generalizing contextual bandits to sequential decision making. These models require an agent to take actions based on observations (features) with the goal of achieving long-term performance competitive with a large set of policies. To …
Neural planners for RDDL MDPs produce deep reactive policies in an offline fashion. These scale well with large domains, but are sample inefficient and time-consuming to train from scratch for each new problem. To mitigate this, recent work has studied neural transfer learning, so that a generic planner trained on othe…
A Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects. Early work in RMDPs outputs generalized (instance-independent) first-order policies or value functions as a means to solve all instanc…
AIF improves physical AI agents' performance in dynamic environments.
Paper improves volatility estimation using a Queue-Reactive model.
Accurate predictions of reactive mixing are critical for many Earth and environmental science problems. To investigate mixing dynamics over time under different scenarios, a high-fidelity, finite-element-based numerical model is built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate …
Study examines Fed's pandemic communication strategies.
Examines financial risks' impact on EU-15 economic growth.
The hybrid clustering-classification neural network is proposed. This network allows increasing a quality of information processing under the condition of overlapping classes due to the rational choice of a learning rate parameter and introducing a special procedure of fuzzy reasoning in the clustering process, which o…
Enhances diffusion-based sampling for molecular systems.
Distribution grids are currently challenged by frequent voltage excursions induced by intermittent solar generation. Smart inverters have been advocated as a fast-responding means to regulate voltage and minimize ohmic losses. Since optimal inverter coordination may be computationally challenging and preset local contr…
Paper revises power theory using classical mechanics concepts.
Turbulence is still one of the main challenges for accurately predicting reactive flows. Therefore, the development of new turbulence closures which can be applied to combustion problems is essential. Data-driven modeling has become very popular in many fields over the last years as large, often extensively labeled, da…
Analysis of reactive-diffusion simulations requires a large number of independent model runs. For each high-fidelity simulation, inputs are varied and the predicted mixing behavior is represented by changes in species concentration. It is then required to discern how the model inputs impact the mixing process. This tas…
Although the challenge of the device connection is much relieved in 5G networks, the training latency is still an obstacle preventing Federated Learning (FL) from being largely adopted. One of the most fundamental problems that lead to large latency is the bad candidate-selection for FL. In the dynamic environment, the…
Traditional energy-based learning models associate a single energy metric to each configuration of variables involved in the underlying optimization process. Such models associate the lowest energy state to the optimal configuration of variables under consideration, and are thus inherently dissipative. In this paper we…
During reactive transport modeling, the computational cost associated with chemical reaction calculations is often 10-100 times higher than that of transport calculations. Most of these costs results from chemical equilibrium calculations that are performed at least once in every mesh cell and at every time step of the…
Deep reinforcement learning improves trading performance in volatile energy markets.
RL optimizes trading algorithms to reduce market impact and costs.
Algorithm detects concept drift and adapts models in streaming data.
In this work we introduce two variants of multivariate Hawkes models with an explicit dependency on various queue sizes aimed at modeling the stochastic time evolution of a limit order book. The models we propose thus integrate the influence of both the current book state and the past order flow. The first variant cons…
New BE dimension measure reveals rich RL problems with sample-efficient algorithms.
Study develops a data-based model for in-cylinder pressure and cyclic variations in RCCI engines.
A new hierarchy quantifies agency in systems based on information processing.
A2MT learns agents to select which modalities to acquire at test time.
Agent-based model simulates market dynamics with real-time order matching.
This paper proposes a general model for synchronized crowding behavior. An order parameter is introduced to quantify the level of synchronization which is shown a function of percentage of agents in reactive state. Further, synchronization is shown to be driven by the most active agents with the highest volatility. A t…
The paper develops a method to predict the latent deterioration phase in limit order books before stress is observed.
New algorithm extracts device profiles for short-term power predictions in commercial buildings.
We study the relation between the trading behavior of agents and volatility in toy markets of adaptive inductively rational agents. We show that excess volatility, in such simplified markets, arises as a consequence of {\em i)} the neglect of market impact implicit in price taking behavior and of {\em ii)} excessive re…
Improved portfolio optimization method yields better risk-adjusted returns.