Sharp inequalities for matrix means with unknown variance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New measure of robustness for estimators, with tight bounds for Gaussian mean estimation.
The asymptotic concentration of the Fr{é}chet mean of IID random variables on a Rieman-nian manifold was established with a central limit theorem by Bhattacharya \& Patrangenaru (BP-CLT) [6]. This asymptotic result shows that the Fr{é}chet mean behaves almost as the usual Euclidean case for sufficiently concentrated di…
A new trading strategy using reinforcement learning for statistical arbitrage.
The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We also show that, under some asumptions, universal consisten…
New estimator reduces kernel mean estimation error.
We introduce performance-based regularization (PBR), a new approach to addressing estimation risk in data-driven optimization, to mean-CVaR portfolio optimization. We assume the available log-return data is iid, and detail the approach for two cases: nonparametric and parametric (the log-return distribution belongs in …
The study investigates the consistency of -means clustering under finite expectation assumptions.
We solve robust optimization problems using Wasserstein balls and apply it to mean-CVaR optimization.
New method for high-dimensional linear regression using empirical Bayes.
A mean function in reproducing kernel Hilbert space, or a kernel mean, is an important part of many applications ranging from kernel principal component analysis to Hilbert-space embedding of distributions. Given finite samples, an empirical average is the standard estimate for the true kernel mean. We show that this e…
Robust portfolio optimization considers uncertainty in market probabilities.
The paper analyzes the mean field Langevin dynamics and its convergence rate.
On-line portfolio selection has attracted increasing interests in machine learning and AI communities recently. Empirical evidences show that stock's high and low prices are temporary and stock price relatives are likely to follow the mean reversion phenomenon. While the existing mean reversion strategies are shown to …
This work extends Ledoit-Wolf shrinkage to unknown mean covariance estimation.
This paper extends Median-of-Means to new learning problems involving pairwise comparisons.
Simple private estimators for mean and covariance outperform existing methods.
Transformers solve Poisson means estimation via empirical Bayes.
Estimates multiple means in high dimensions using convex combinations.
The paper assesses conditions for the uniqueness of k-means clustering.
A mean function in a reproducing kernel Hilbert space (RKHS), or a kernel mean, is central to kernel methods in that it is used by many classical algorithms such as kernel principal component analysis, and it also forms the core inference step of modern kernel methods that rely on embedding probability distributions in…
We consider a system of diffusion processes that interact through their empirical mean and have a stabilizing force acting on each of them, corresponding to a bistable potential. There are three parameters that characterize the system: the strength of the intrinsic stabilization, the strength of the external random per…
Recent studies have shown that online portfolio selection strategies that exploit the mean reversion property can achieve excess return from equity markets. This paper empirically investigates the performance of state-of-the-art mean reversion strategies on real market data. The aims of the study are twofold. The first…
RL approach for continuous-time mean-variance portfolio selection with empirical validation.
Paper shows robust estimators converge to true risk minimizers at optimal rates.
K-means clustering improved for robustness to outliers and distribution shifts.
Study shows how neural networks generalize with minimal training data.
We propose an empirical Bayes estimator based on Dirichlet process mixture model for estimating the sparse normalized mean difference, which could be directly applied to the high dimensional linear classification. In theory, we build a bridge to connect the estimation error of the mean difference and the misclassificat…
A new sequential method estimates Poisson means in streaming data, achieving optimality and efficiency.
We extend the empirical results published in article "Empirical Evidence on Arbitrage by Changing the Stock Exchange" by means of machine learning and advanced econometric methodologies based on Smooth Transition Regression models and Artificial Neural Networks.
The Normal Means problem plays a fundamental role in many areas of modern high-dimensional statistics, both in theory and practice. And the Empirical Bayes (EB) approach to solving this problem has been shown to be highly effective, again both in theory and practice. However, almost all EB treatments of the Normal Mean…
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We show that universal consistency of Empirical Risk Minimiza…
A new method of moments estimator goes beyond data reweighting.
We investigate the variety of a portfolio of stocks in normal and extreme days of market activity. We show that the variety carries information about the market activity which is not present in the single-index model and we observe that the variety time evolution is not time reversal around the crash days. We obtain th…
We study theoretical and empirical aspects of the mean exit time of financial time series. The theoretical modeling is done within the framework of continuous time random walk. We empirically verify that the mean exit time follows a quadratic scaling law and it has associated a pre-factor which is specific to the analy…
A microscopic model of aggregation and fragmentation is introduced to investigate the size distribution of businesses. In the model, businesses are constrained to comply with the market price, as expected by the customers, while customers can only buy at the prices offered by the businesses. We show numerically and ana…
Explaining how overparametrized neural networks simultaneously achieve low risk and zero empirical risk on benchmark datasets is an open problem. PAC-Bayes bounds optimized using variational inference (VI) have been recently proposed as a promising direction in obtaining non-vacuous bounds. We show empirically that thi…
Kernel -means clustering can correctly identify and extract a far more varied collection of cluster structures than the linear -means clustering algorithm. However, kernel -means clustering is computationally expensive when the non-linear feature map is high-dimensional and there are many input points. Kernel …
New algorithm for biclustering with improved performance.
Paper improves CI and CS for bounded means using betting and mixtures.
Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.
Optimizes sparse mean-reverting portfolios for higher returns.
Recent empirical studies suggest that the volatilities associated with financial time series exhibit short-range correlations. This entails that the volatility process is very rough and its autocorrelation exhibits sharp decay at the origin. Another classic stylistic feature often assumed for the volatility is that it …
Optimal benchmark design varies based on costs in financial manipulation.
New DP methods for estimating means and frequencies with varying privacy demands.
Few-shot learning aims to train efficient predictive models with a few examples. The lack of training data leads to poor models that perform high-variance or low-confidence predictions. In this paper, we propose to meta-learn the ensemble of epoch-wise empirical Bayes models (E3BM) to achieve robust predictions. "Epoch…
Study examines mean estimation in high dimensions with small data.