New framework for policy gradient methods in continuous time reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Randomised classifiers outperform deterministic ones in strategic classification.
New algorithm improves game learning with randomised optimism.
Randomized exploration in linear bandits achieves optimal regret bounds.
New method uses randomised signatures for generating financial time series data.
We extend previous large deviations results for the randomised Heston model to the case of moderate deviations. The proofs involve the Gärtner-Ellis theorem and sharp large deviations tools.
Unified high-probability regret bounds for online convex optimisation with randomised gradient estimators.
We propose a randomised version of the Heston model-a widely used stochastic volatility model in mathematical finance-assuming that the starting point of the variance process is a random variable. In such a system, we study the small-and large-time behaviours of the implied volatility, and show that the proposed random…
Improved Bayesian optimisation method using randomised Gaussian process UCB.
Improved control approach for correlated bandits with better performance.
We introduce Bayesian least-squares policy iteration (BLSPI), an off-policy, model-free, policy iteration algorithm that uses the Bayesian least-squares temporal-difference (BLSTD) learning algorithm to evaluate policies. An online variant of BLSPI has been also proposed, called randomised BLSPI (RBLSPI), that improves…
We introduce a general learning framework for private machine learning based on randomised response. Our assumption is that all actors are potentially adversarial and as such we trust only to release a single noisy version of an individual's datapoint. We discuss a general approach that forms a consistent way to estima…
Two modified tests improve the reliability of evaluating explanation methods.
New method optimizes experimental design for specific applications.
Numerous kinds of uncertainties may affect an economy, e.g. economic, political, and environmental ones. We model the aggregate impact by the uncertainties on an economy and its associated financial market by randomised mixtures of Lévy processes. We assume that market participants observe the randomised mixtures only …
We design a randomised parallel version of Adaboost based on previous studies on parallel coordinate descent. The algorithm uses the fact that the logarithm of the exponential loss is a function with coordinate-wise Lipschitz continuous gradient, in order to define the step lengths. We provide the proof of convergence …
A new method for stochastic control based on neural networks and using randomisation of discrete random variables is proposed and applied to optimal stopping time problems. The method models directly the policy and does not need the derivation of a dynamic programming principle nor a backward stochastic differential eq…
New method improves calibration of BayesCG for better uncertainty quantification.
Discrete time analogues of ergodic stochastic differential equations (SDEs) are one of the most popular and flexible tools for sampling high-dimensional probability measures. Non-asymptotic analysis in the Wasserstein distance of sampling algorithms based on Euler discretisations of SDEs has been recently develop…
Study proves value of non-Markovian games with partial, asymmetric info.
We consider the problem of link prediction, based on partial observation of a large network, and on side information associated to its vertices. The generative model is formulated as a matrix logistic regression. The performance of the model is analysed in a high-dimensional regime under a structural assumption. The mi…
New bounds for model generalization under deterministic gradient descent.
We propose and evaluate alternative ensemble schemes for a new instance based learning classifier, the Randomised Sphere Cover (RSC) classifier. RSC fuses instances into spheres, then bases classification on distance to spheres rather than distance to instances. The randomised nature of RSC makes it ideal for use in en…
American options in a multi-asset market model with proportional transaction costs are studied in the case when the holder of an option is able to exercise it gradually at a so-called mixed (randomised) stopping time. The introduction of gradual exercise leads to tighter bounds on the option price when compared to the …
We develop a new Monte Carlo variance reduction method to estimate the expectation of two commonly encountered path-dependent functionals: first-passage times and occupation times of sets. The method is based on a recursive approximation of the first-passage time probability and expected occupation time of sets of a Le…
EVILL uses randomised perturbations to improve exploration in bandit problems.
Paper develops a privacy-preserving nonparametric regression method.
ARC algorithm optimizes dynamic pricing with correlated observations.
Game (Israeli) options in a multi-asset market model with proportional transaction costs are studied in the case when the buyer is allowed to exercise the option and the seller has the right to cancel the option gradually at a mixed (or randomised) stopping time, rather than instantly at an ordinary stopping time. Allo…
Novel strategy for federated learning with privacy-preserving predictors and nonvacuous generalization bounds.
FP uses random projections to train networks without feedback, achieving comparable performance to backpropagation.
Develops a kernel-based framework for dynamic trading strategies.
We price and hedge American options robustly in continuous time.
Paper improves privacy bounds for shuffle model using novel numerical techniques.
Two firms compete in a financial market, choosing dividend strategies to avoid default and maximize profits.
In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the simplest case, is to estimate the leading eigenvector of a population covariance matrix under the assumption that this eigenvector is sparse…
In this paper we propose a Bayesian method for estimating architectural parameters of neural networks, namely layer size and network depth. We do this by learning concrete distributions over these parameters. Our results show that regular networks with a learnt structure can generalise better on small datasets, while f…
We formalise the widespread idea of interpreting neural network decisions as an explicit optimisation problem in a rate-distortion framework. A set of input features is deemed relevant for a classification decision if the expected classifier score remains nearly constant when randomising the remaining features. We disc…
New method estimates nested expectations with biased and antithetic sampling.
An explorative data analysis system should be aware of what the user already knows and what the user wants to know of the data: otherwise the system cannot provide the user with the most informative and useful views of the data. We propose a principled way to do exploratory data analysis, where the user's background kn…
Paper defines saddle points in asymmetric Dynkin games using martingale theory.
We develop a monitoring procedure to detect changes in a large approximate factor model. Letting be the number of common factors, we base our statistics on the fact that the -th eigenvalue of the sample covariance matrix is bounded under the null of no change, whereas it becomes spiked under cha…
Classical (Itô diffusions) stochastic volatility models are not able to capture the steepness of small-maturity implied volatility smiles. Jumps, in particular exponential Lévy and affine models, which exhibit small-maturity exploding smiles, have historically been proposed to remedy this (see \cite{Tank} for an overvi…
Estimates long-term effects from short-term experiments and observational data with unobserved confounders.
We consider classification in the presence of class-dependent asymmetric label noise with unknown noise probabilities. In this setting, identifiability conditions are known, but additional assumptions were shown to be required for finite sample rates, and so far only the parametric rate has been obtained. Assuming thes…
The generation of artificial data based on existing observations, known as data augmentation, is a technique used in machine learning to improve model accuracy, generalisation, and to control overfitting. Augmentor is a software package, available in both Python and Julia versions, that provides a high level API for th…
We construct two examples of shareholder networks in which shareholders are connected if they have shares in the same company. We do this for the shareholders in Turkish companies and we compare this against the network formed from the shareholdings in Dutch companies. We analyse the properties of these two networks in…
Study generalization of voting classifiers using margin-based bounds.