ePF improves PF for ITS by balancing exploration and exploitation, outperforming baselines.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces. Unlike traditional point estimate methods, randomized value functions maintain a posterior distribution over action-space values. This prevents the …
The above title is the same, but with "semisimple" instead of "simple," as that of a notice by N. Kowalsky. There, she announced many theorems on the subject of actions of simple Lie groups preserving a Lorentz structure. Unfortunately, she published proofs for essentially only half of the announced results before her …
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are com…
Reinforcement learning with sparse rewards is still an open challenge. Classic methods rely on getting feedback via extrinsic rewards to train the agent, and in situations where this occurs very rarely the agent learns slowly or cannot learn at all. Similarly, if the agent receives also rewards that create suboptimal m…
PFPN uses particle filtering to improve character control in physics-based simulations.
Slow feature analysis (SFA) is an unsupervised-learning algorithm that extracts slowly varying features from a multi-dimensional time series. A supervised extension to SFA for classification and regression is graph-based SFA (GSFA). GSFA is based on the preservation of similarities, which are specified by a graph struc…
A new method for releasing AI workflows to avoid premature incorrect results.
Enhances valuation of variable annuities with stochastic interest rate models.
Post-hoc transforms can reverse model performance trends, especially in noisy settings.
Intricating cardiac complexities are the primary factor associated with healthcare costs and the highest cause of death rate in the world. However, preventive measures like the early detection of cardiac anomalies can prevent severe cardiovascular arrests of varying complexities and can impose a substantial impact on h…
This paper improves spectral clustering for large datasets using the Nystrom method.
The prospects of Kahneman and Tversky, Mega Million and Powerball lotteries, St. Petersburg paradox, premature profits and growing losses criticized by Livermore are reviewed under an angle of view comparing mathematical expectations with awards received. Original prospects have been formulated as a one time opportunit…
Delusional bias is a fundamental source of error in approximate Q-learning. To date, the only techniques that explicitly address delusion require comprehensive search using tabular value estimates. In this paper, we develop efficient methods to mitigate delusional bias by training Q-approximators with labels that are "…
Protein Thoughts interprets protein interactions with clear reasoning, improving prediction accuracy.
This paper proposes a nonparametric Bayesian method for exploratory data analysis and feature construction in continuous time series. Our method focuses on understanding shared features in a set of time series that exhibit significant individual variability. Our method builds on the framework of latent Diricihlet alloc…
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a st…
Sustaining efficiency and stability by properly controlling the equity to asset ratio is one of the most important and difficult challenges in bank management. Due to unexpected and abrupt decline of asset values, a bank must closely monitor its net worth as well as market conditions, and one of its important concerns …
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
Proximal Policy Optimization (PPO) is a highly popular model-free reinforcement learning (RL) approach. However, we observe that in a continuous action space, PPO can prematurely shrink the exploration variance, which leads to slow progress and may make the algorithm prone to getting stuck in local optima. Drawing insp…
After birth, extremely preterm infants often require specialized respiratory management in the form of invasive mechanical ventilation (IMV). Protracted IMV is associated with detrimental outcomes and morbidities. Premature extubation, on the other hand, would necessitate reintubation which is risky, technically challe…
We develop a novel framework for computing the total valuation adjustment (XVA) of a European claim accounting for funding costs, counterparty credit risk, and collateralization. Based on no-arbitrage arguments, we derive the nonlinear backward stochastic differential equations (BSDEs) associated with the replicating p…
In this paper, we study a stochastic optimal control problem with stochastic volatility. We prove the sufficient and necessary maximum principle for the proposed problem. Then we apply the results to solve an investment, consumption and life insurance problem with stochastic volatility, that is, we consider a wage earn…
A meta-learning approach for efficient algorithm selection in budget-limited scenarios.
This paper explores how enforcing equivariance constraints limits neural network expressivity and proposes compensatory model size increases.
New method stops experiments early for harm in diverse groups.
Bayesian optimization (BO) aims to minimize a given blackbox function using a model that is updated whenever new evidence about the function becomes available. Here, we address the problem of BO under partially right-censored response data, where in some evaluations we only obtain a lower bound on the function value. T…
Dynamic model pruning improves performance on deep neural networks without retraining.
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested …
LCBO tackles constrained optimization in high dimensions, offering a polynomial convergence rate.
Neural architecture search (NAS) automatically finds the best task-specific neural network topology, outperforming many manual architecture designs. However, it can be prohibitively expensive as the search requires training thousands of different networks, while each can last for hours. In this work, we propose the Gra…
New framework improves classification accuracy using Pillai's trace and ULDA.
Stabilizes policy optimization with off-policy data using divergence augmentation.
Controlled interventions provide the most direct source of information for learning causal effects. In particular, a dose-response curve can be learned by varying the treatment level and observing the corresponding outcomes. However, interventions can be expensive and time-consuming. Observational data, where the treat…
This paper considers an optimal life insurance for a householder subject to mortality risk. The household receives a wage income continuously, which is terminated by unexpected (premature) loss of earning power or (planned and intended) retirement, whichever happens first. In order to hedge the risk of losing income st…
The multi-factor model is a widely used model in quantitative investment. The success of a multi-factor model is largely determined by the effectiveness of the alpha factors used in the model. This paper proposes a new evolutionary algorithm called AutoAlpha to automatically generate effective formulaic alphas from mas…
Optimal dividend strategy with irreversible reinsurance constraints.
We introduce an extension to Merton's famous continuous time model of optimal consumption and investment, in the spirit of previous works by Pliska and Ye, to allow for a wage earner to have a random lifetime and to use a portion of the income to purchase life insurance in order to provide for his estate, while investi…
Computational swarm intelligence consists of multiple artificial simple agents exchanging information while exploring a search space. Despite a rich literature in the field, with works improving old approaches and proposing new ones, the mechanism by which complex behavior emerges in these systems is still not well und…
EviTrack improves sequential prediction in delayed disambiguation scenarios.
Learning robot controllers by minimizing a black-box objective cost using Bayesian optimization (BO) can be time-consuming and challenging. It is very often the case that some roll-outs result in failure behaviors, causing premature experiment detention. In such cases, the designer is forced to decide on heuristic cost…
This paper examines the optimal annuitization, investment and consumption strategies of a utility-maximizing retiree facing a stochastic time of death under a variety of institutional restrictions. We focus on the impact of aging on the optimal purchase of life annuities which form the basis of most Defined Benefit pen…
The emergence of deep learning networks raises a need for explainable AI so that users and domain experts can be confident applying them to high-risk decisions. In this paper, we leverage data from the latent space induced by deep learning models to learn stereotypical representations or "prototypes" during training to…
EGFC learns from streaming data to classify power quality disturbances.
In reinforcement learning, a decision needs to be made at some point as to whether it is worthwhile to carry on with the learning process or to terminate it. In many such situations, stochastic elements are often present which govern the occurrence of rewards, with the sequential occurrences of positive rewards randoml…
Theory for RLHF generalization under reward shift and clipped KL.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.