This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New research shows that the dimension gap between intrinsic and ambient dimensions affects adversarial vulnerability of machine learning models.
Non-intrusive load monitoring (NILM), also known as energy disaggregation, is a blind source separation problem where a household's aggregate electricity consumption is broken down into electricity usages of individual appliances. In this way, the cost and trouble of installing many measurement devices over numerous ho…
Hybrid RL algorithm combines offline and online data for robust and efficient policy learning.
Continuum Dropout improves neural differential equations by preventing overfitting.
Energy disaggregation in a non-intrusive way estimates appliance level electricity consumption from a single meter that measures the whole house electricity demand. Recently, with the ongoing increment of energy data, there are many data-driven deep learning architectures being applied to solve the non-intrusive energy…
TIM framework uses LLMs and domain experts to infer DeFi user transaction intents.
This study analyzes the effects of lifting lockdowns on Brazil's COVID-19 spread.
This article studies the financial integration between the six main Latin American markets and the US market in a nonlinear framework. Using the threshold cointegration techniques of Hansen and Seo (2002), we show significant threshold stock market linkages between Mexico, Chile and the US. Thus, the dynamics of these …
Investigates numerical issues in GP interpolation parameter estimation.
In the last few years, deep multi-agent reinforcement learning (RL) has become a highly active area of research. A particularly challenging class of problems in this area is partially observable, cooperative, multi-agent learning, in which teams of agents must learn to coordinate their behaviour while conditioning only…
This paper reviews off-policy evaluation methods in reinforcement learning.
In this work we use a proven model to study a dynamic duopolistic competition between an old and a new technology which, through improved technical performance - e.g. data transmission capacity - fight in order to conquer market share. The process whereby an old technology fights a new one off through own improvements …
This paper investigates the problem of online prediction learning, where learning proceeds continuously as the agent interacts with an environment. The predictions made by the agent are contingent on a particular way of behaving, represented as a value function. However, the behavior used to select actions and generate…
The paper analyzes the role of ReLU gates in deep learning networks.
We introduce a new approach to incorporate uncertainty into the decision to invest in a commodity reserve. The investment is an irreversible one-off capital expenditure, after which the investor receives a stream of cashflow from extracting the commodity and selling it on the spot market. The investor is exposed to pri…
Recent studies have shown that the aggregated dynamic flexibility of an ensemble of thermostatic loads can be modeled in the form of a virtual battery. The existing methods for computing the virtual battery parameters require the knowledge of the first-principle models and parameter values of the loads in the ensemble.…
Paper tackles policy selection in offline RL without hyperparameters.
Recurrent neural networks (RNNs) achieve cutting-edge performance on a variety of problems. However, due to their high computational and memory demands, deploying RNNs on resource constrained mobile devices is a challenging task. To guarantee minimum accuracy loss with higher compression rate and driven by the mobile r…
Emission from a class of benzene-based molecules known as Polycyclic Aromatic Hydrocarbons (PAHs) dominates the infrared spectrum of star-forming regions. The observed emission appears to arise from the combined emission of numerous PAH species, each with its unique spectrum. Linear superposition of the PAH spectra ide…
New metrics help rebuild trust in Active Learning for industry practitioners.
Despite the recent progress in hyperparameter optimization (HPO), available benchmarks that resemble real-world scenarios consist of a few and very large problem instances that are expensive to solve. This blocks researchers and practitioners not only from systematically running large-scale comparisons that are needed …
We propose to execute deep neural networks (DNNs) with dynamic and sparse graph (DSG) structure for compressive memory and accelerative execution during both training and inference. The great success of DNNs motivates the pursuing of lightweight models for the deployment onto embedded devices. However, most of the prev…
Adapts GRPO for off-policy RL, improving reward.
Machine learning models are vulnerable to adversarial inputs that induce seemingly unjustifiable errors. As automated classifiers are increasingly used in industrial control systems and machinery, these adversarial errors could grow to be a serious problem. Despite numerous studies over the past few years, the field of…
Paper tackles efficient evaluation of natural stochastic policies in offline RL.
Paper introduces risk assessment for contextual bandits without experiments.
New sketches for weighted sampling without replacement improve accuracy and efficiency.
Reinforcement learning has attracted great attention recently, especially policy gradient algorithms, which have been demonstrated on challenging decision making and control tasks. In this paper, we propose an active multi-step TD algorithm with adaptive stepsizes to learn actor and critic. Specifically, our model cons…
The paper explores gaps in curvature-related metrics and rigidity.
Study gap-dependent regret bounds for risk-sensitive RL.
Cramming method evaluates learned policies from contextual bandits efficiently.
Machine learning has recently emerged as a fruitful area for finding potential quantum computational advantage. Many of the quantum enhanced machine learning algorithms critically hinge upon the ability to efficiently produce states proportional to high-dimensional data points stored in a quantum accessible memory. Eve…
The paper explores how to extrapolate from limited data points using causal mechanisms.
CoCoPIE shows AI can run on regular devices without special hardware.
Paper analyzes origami slope gaps and their distribution, finding a unique pattern.
We seek to align agent behavior with a user's objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states. The user knows the rewards and unsafe states, but querying the user is expensive. To address this challenge, we propose an algorithm that safely and …
PaRoT simplifies robust training for deep neural networks.
Random hyperbolic surfaces have nearly optimal spectral gaps.
We present a method that trains large capacity neural networks with significantly improved accuracy and lower dynamic computational cost. We achieve this by gating the deep-learning architecture on a fine-grained-level. Individual convolutional maps are turned on/off conditionally on features in the network. To achieve…
Paper improves volume gap between minimal submanifolds and unit spheres.
The article proves a conjecture about the fundamental gap for horoconvex domains in hyperbolic space.
A new framework detects anomalous inputs to DNNs.
Kahler-Einstein metrics linked to eigenvalue gaps on Fano manifolds.
Local gaps in Ricci shrinkers depend only on dimension.
Study shows gaps in Bitcoin order book are linked to returns but only in the short term.
Researchers compute gap distributions for saddle connection directions on specific translation surfaces.
Proves gap rigidity theorem for Hermitian symmetric spaces.