The ABC algorithm's social interactions are characterized through a novel interaction network.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ePF improves PF for ITS by balancing exploration and exploitation, outperforming baselines.
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are com…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
LCBO tackles constrained optimization in high dimensions, offering a polynomial convergence rate.
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a st…
Stabilizes policy optimization with off-policy data using divergence augmentation.
The above title is the same, but with "semisimple" instead of "simple," as that of a notice by N. Kowalsky. There, she announced many theorems on the subject of actions of simple Lie groups preserving a Lorentz structure. Unfortunately, she published proofs for essentially only half of the announced results before her …
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
This paper improves test-time adaptation for distribution shifts using confidence maximization and input transformation.
A new ES method improves reinforcement learning speed and accuracy.
A new method for releasing AI workflows to avoid premature incorrect results.
AutoAlpha efficiently discovers effective alpha factors for quantitative investment.
Quantum Annealing Enhanced Reinforcement Learning for Accurate RUL Prediction
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
PFPN uses particle filtering to improve character control in physics-based simulations.
The prospects of Kahneman and Tversky, Mega Million and Powerball lotteries, St. Petersburg paradox, premature profits and growing losses criticized by Livermore are reviewed under an angle of view comparing mathematical expectations with awards received. Original prospects have been formulated as a one time opportunit…
System detects and predicts cardiac anomalies from ECG data.
This paper proposes a nonparametric Bayesian method for exploratory data analysis and feature construction in continuous time series. Our method focuses on understanding shared features in a set of time series that exhibit significant individual variability. Our method builds on the framework of latent Diricihlet alloc…
Sustaining efficiency and stability by properly controlling the equity to asset ratio is one of the most important and difficult challenges in bank management. Due to unexpected and abrupt decline of asset values, a bank must closely monitor its net worth as well as market conditions, and one of its important concerns …
Proximal Policy Optimization (PPO) is a highly popular model-free reinforcement learning (RL) approach. However, we observe that in a continuous action space, PPO can prematurely shrink the exploration variance, which leads to slow progress and may make the algorithm prone to getting stuck in local optima. Drawing insp…
After birth, extremely preterm infants often require specialized respiratory management in the form of invasive mechanical ventilation (IMV). Protracted IMV is associated with detrimental outcomes and morbidities. Premature extubation, on the other hand, would necessitate reintubation which is risky, technically challe…
We develop a novel framework for computing the total valuation adjustment (XVA) of a European claim accounting for funding costs, counterparty credit risk, and collateralization. Based on no-arbitrage arguments, we derive the nonlinear backward stochastic differential equations (BSDEs) associated with the replicating p…
In this paper, we study a stochastic optimal control problem with stochastic volatility. We prove the sufficient and necessary maximum principle for the proposed problem. Then we apply the results to solve an investment, consumption and life insurance problem with stochastic volatility, that is, we consider a wage earn…
This paper explores how enforcing equivariance constraints limits neural network expressivity and proposes compensatory model size increases.
New method stops experiments early for harm in diverse groups.
Bayesian optimization (BO) aims to minimize a given blackbox function using a model that is updated whenever new evidence about the function becomes available. Here, we address the problem of BO under partially right-censored response data, where in some evaluations we only obtain a lower bound on the function value. T…
Dynamic model pruning improves performance on deep neural networks without retraining.
A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested …
ConQUR tackles delusional bias in deep Q-learning, improving performance in Atari games.
Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces. Unlike traditional point estimate methods, randomized value functions maintain a posterior distribution over action-space values. This prevents the …
Neural architecture search (NAS) automatically finds the best task-specific neural network topology, outperforming many manual architecture designs. However, it can be prohibitively expensive as the search requires training thousands of different networks, while each can last for hours. In this work, we propose the Gra…
New framework improves classification accuracy using Pillai's trace and ULDA.
Controlled interventions provide the most direct source of information for learning causal effects. In particular, a dose-response curve can be learned by varying the treatment level and observing the corresponding outcomes. However, interventions can be expensive and time-consuming. Observational data, where the treat…
This paper considers an optimal life insurance for a householder subject to mortality risk. The household receives a wage income continuously, which is terminated by unexpected (premature) loss of earning power or (planned and intended) retirement, whichever happens first. In order to hedge the risk of losing income st…
A meta-learning approach for efficient algorithm selection in budget-limited scenarios.
Optimal dividend strategy with irreversible reinsurance constraints.
We introduce an extension to Merton's famous continuous time model of optimal consumption and investment, in the spirit of previous works by Pliska and Ye, to allow for a wage earner to have a random lifetime and to use a portion of the income to purchase life insurance in order to provide for his estate, while investi…
EviTrack improves sequential prediction in delayed disambiguation scenarios.
This paper examines the optimal annuitization, investment and consumption strategies of a utility-maximizing retiree facing a stochastic time of death under a variety of institutional restrictions. We focus on the impact of aging on the optimal purchase of life annuities which form the basis of most Defined Benefit pen…
EGFC learns from streaming data to classify power quality disturbances.
In reinforcement learning, a decision needs to be made at some point as to whether it is worthwhile to carry on with the learning process or to terminate it. In many such situations, stochastic elements are often present which govern the occurrence of rewards, with the sequential occurrences of positive rewards randoml…
Theory for RLHF generalization under reward shift and clipped KL.
Bayesian optimization method predicts high costs for unstable robot controllers.
Slow feature analysis (SFA) is an unsupervised-learning algorithm that extracts slowly varying features from a multi-dimensional time series. A supervised extension to SFA for classification and regression is graph-based SFA (GSFA). GSFA is based on the preservation of similarities, which are specified by a graph struc…
The paper analyzes and proposes a new stopping criterion for recursive Bayesian classification.
The paper uses learned prototypes to explain deep learning models for time-series data.
Enhances valuation of variable annuities with stochastic interest rate models.