Second-order optimization speeds up deep hedging for complex options.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This thesis studies domain adaptation under minimal distribution similarity assumptions using moments.
Domain adaptation algorithms are designed to minimize the misclassification risk of a discriminative model for a target domain with little training data by adapting a model from a source domain with a large amount of training data. Standard approaches measure the adaptation discrepancy based on distance measures betwee…
Estimates RL data for dynamic treatment effects using GMM.
In this paper, we investigate the popular deep learning optimization routine, Adam, from the perspective of statistical moments. While Adam is an adaptive lower-order moment based (of the stochastic gradient) method, we propose an extension namely, HAdam, which uses higher order moments of the stochastic gradient. Our …
MOMENT selects and estimates mixed-effects models using moment identities.
Adaptive importance sampling is a class of techniques for finding good proposal distributions for importance sampling. Often the proposal distributions are standard probability distributions whose parameters are adapted based on the mismatch between the current proposal and a target distribution. In this work, we prese…
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
We propose moment-based variational inference as a flexible framework for approximate smoothing of latent Markov jump processes. The main ingredient of our approach is to partition the set of all transitions of the latent process into classes. This allows to express the Kullback-Leibler divergence between the approxima…
The total duration of drawdowns is shown to provide a moment-free, unbiased, efficient and robust estimator of Sharpe ratios both for Gaussian and heavy-tailed price returns. We then use this quantity to infer an analytic expression of the bias of moment-based Sharpe ratio estimators as a function of the return distrib…
The paper proposes a method to monitor deep learning predictions for retraining, reducing costs.
DOLCE improves off-policy evaluation and learning by decomposing effects.
Estimates MLDS using tensor decomposition, improving upon existing methods.
We consider a particular instance of a common problem in recommender systems: using a database of book reviews to inform user-targeted recommendations. In our dataset, books are categorized into genres and sub-genres. To exploit this nested taxonomy, we use a hierarchical model that enables information pooling across a…
New method approximates diffusion process posteriors using moment functions.
We present a detailed analysis of \emph{observable} moments based parameter estimators for the Heston SDEs jointly driving the rate of returns and the squared volatilities . Since volatilities are not directly observable, our parameter estimators are constructed from empirical moments of realized volatilitie…
Corrected moment-based methods improve inference in topic model regression.
We propose some machine-learning-based algorithms to solve hedging problems in incomplete markets. Sources of incompleteness cover illiquidity, untradable risk factors, discrete hedging dates and transaction costs. The proposed algorithms resulting strategies are compared to classical stochastic control techniques on s…
New model estimates corporate defaults using pure jump processes, capturing extreme events.
Neural Hawkes method estimates cryptocurrency market microstructure and causality.
We propose a fair principal component analysis method that balances reconstruction error and subgroup fairness.
We present a general probabilistic perspective on Gaussian filtering and smoothing. This allows us to show that common approaches to Gaussian filtering/smoothing can be distinguished solely by their methods of computing/approximating the means and covariances of joint probabilities. This implies that novel filters and …
Proposes DWMD for better matching of hidden representations across domains.
This work provides a computationally efficient and statistically consistent moment-based estimator for mixtures of spherical Gaussians. Under the condition that component means are in general position, a simple spectral decomposition technique yields consistent parameter estimates from low-order observable moments, wit…
Stochastic Kronecker graphs supply a parsimonious model for large sparse real world graphs. They can specify the distribution of a large random graph using only three or four parameters. Those parameters have however proved difficult to choose in specific applications. This article looks at method of moments estimators…
Develops a simulation-based method to translate expert knowledge into prior distributions for Bayesian models.
We develop a behavioral asset pricing model in which agents trade in a market with information friction. Profit-maximizing agents switch between trading strategies in response to dynamic market conditions. Due to noisy private information about the fundamental value, the agents form different evaluations about heteroge…
The performance of a modulation classifier is highly sensitive to channel signal-to-noise ratio (SNR). In this paper, we focus on amplitude-phase modulations and propose a modulation classification framework based on centralized data fusion using multiple radios and the hybrid maximum likelihood (ML) approach. In order…
We present an efficient algorithm for learning mixed membership models when the number of variables is much larger than the number of hidden components . This algorithm reduces the computational complexity of state-of-the-art tensor methods, which require decomposing an tensor, to factorizing…
We describe a method for parameter estimation in bipartite probabilistic graphical models for joint prediction of clinical conditions from the electronic medical record. The method does not rely on the availability of gold-standard labels, but rather uses noisy labels, called anchors, for learning. We provide a likelih…
Over the last decade, dividends have become a standalone asset class instead of a mere side product of an equity investment. We introduce a framework based on polynomial jump-diffusions to jointly price the term structures of dividends and interest rates. Prices for dividend futures, bonds, and the dividend paying stoc…
We consider the problem of identifying the parameters of an unknown mixture of two arbitrary -dimensional gaussians from a sequence of independent random samples. Our main results are upper and lower bounds giving a computationally efficient moment-based estimator with an optimal convergence rate, thus resolving a p…
A new method for density estimation using mixture discrepancy and moments.
This paper proposes a novel model of financial prices where: (i) prices are discrete; (ii) prices change in continuous time; (iii) a high proportion of price changes are reversed in a fraction of a second. Our model is analytically tractable and directly formulated in terms of the calendar time and price impact curve. …
Algorithm learns near-optimal policies for reward-mixing MDPs with few latent contexts.
Predictive state representations (PSRs) offer an expressive framework for modelling partially observable systems. By compactly representing systems as functions of observable quantities, the PSR learning approach avoids using local-minima prone expectation-maximization and instead employs a globally optimal moment-base…
Suppose centers are fit to points by heuristically minimizing the -means cost; what is the corresponding fit over the source distribution? This question is resolved here for distributions with bounded moments; in particular, the difference between the sample cost and distribution cost decays with $…
The paper proposes estimators for bid-ask spreads with and without serial dependence.
Motivated by the sampling problems and heterogeneity issues common in high- dimensional big datasets, we consider a class of discordant additive index models. We propose method of moments based procedures for estimating the indices of such discordant additive index models in both low and high-dimensional settings. Our …
New protocols show 1-bit mean estimation can be order-optimal without interaction.
GeoAdaLer enhances geometric understanding of Adam for stochastic optimization.
M-L2O adapts fast to new tasks by self-adapting during test-time.
NeAda solves nonconvex minimax optimization by balancing primal and dual variables adaptively.
New adaptive step-size method for convex optimization without tuning.
Optimal nonparametric regression estimator adapts to unknown smoothness.
New adaptive methods solve weakly convex stochastic optimization problems.
Adaptive reduction scheme approximates optimal policy in regularized MDPs.
Study on distributed nonparametric function estimation with optimal rate and cost of adaptation.