This work examines curvature effects on empirical mean in Riemannian and affine manifolds for small samples.
problem Understanding curvature effects on empirical mean in small samples on Riemannian and affine manifolds.
method Established Taylor expansions for the first and second moments of the empirical mean, showing curvature impacts.
result Explicit formulas for bias and modulation of empirical mean's covariance matrix due to curvature.
Sharp inequalities for matrix means with unknown variance.
problem Estimating matrix means with unknown variance.
method Empirical Bernstein inequalities for symmetric random matrices.
result Adapts to unknown variance with tight deviation bounds.
New measure of robustness for estimators, with tight bounds for Gaussian mean estimation.
problem Developing robust statistical estimators for datasets with noise or outliers.
method Introducing empirical sensitivity as a new robustness measure and proving lower bounds for Gaussian mean estimation.
result Empirical sensitivity bounds for optimal estimators are tight, showing obstructions on mean and variance.
A new trading strategy using reinforcement learning for statistical arbitrage.
problem Traditional statistical arbitrage models rely on model assumptions and price deviations from a long-term mean.
method Empirical reversion time metric, reinforcement learning framework, and state space optimization.
result Optimal mean reversion strategy identified through reinforcement learning.
The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.
problem Linear inverse problems with approximate priors.
method Maximum Entropy on the Mean (MEM) method with data-driven priors.
result Empirical mean convergence and estimates for prior differences based on epigraphical distance.
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We also show that, under some asumptions, universal consisten…
New estimator reduces kernel mean estimation error.
problem Kernel mean estimation in reproducing kernel Hilbert spaces.
method Corrupt data with known distributions and estimate kernel mean under the corrupted distribution.
result The marginalized kernel mean estimator achieves lower estimation error.
We introduce performance-based regularization (PBR), a new approach to addressing estimation risk in data-driven optimization, to mean-CVaR portfolio optimization. We assume the available log-return data is iid, and detail the approach for two cases: nonparametric and parametric (the log-return distribution belongs in …
The study investigates the consistency of k-means clustering under finite expectation assumptions.
problem Consistency of k-means clustering under finite expectation assumptions. method Investigates the conditions under which k-means clustering is consistent, considering finite expectation instead of finite variance. result Inconsistency can arise due to extreme cluster imbalance, leading to some clusters having few points.
We solve robust optimization problems using Wasserstein balls and apply it to mean-CVaR optimization.
problem Distributionally robust optimization with Wasserstein ambiguity sets.
method Transformed robust optimization into non-robust with penalty term, selecting ambiguity set size.
result Impressive results in robust mean-CVaR optimization compared to other strategies.
New method for high-dimensional linear regression using empirical Bayes.
problem Estimating prior in high-dimensional linear regression.
method Variational empirical Bayes approach with NPMLE and mean field approximation.
result Established asymptotic consistency and computational efficiency of the method.
A mean function in reproducing kernel Hilbert space, or a kernel mean, is an important part of many applications ranging from kernel principal component analysis to Hilbert-space embedding of distributions. Given finite samples, an empirical average is the standard estimate for the true kernel mean. We show that this e…
Robust portfolio optimization considers uncertainty in market probabilities.
problem Uncertainty in market probabilities in multiperiod portfolio selection.
method Robust mean-variance optimization using Wasserstein ball centered at empirical data.
result Numerical simulations show improved performance compared to other strategies.
The paper analyzes the mean field Langevin dynamics and its convergence rate.
problem The convergence property of the mean field Langevin dynamics in the context of neural networks.
method The analysis uses a proximal Gibbs distribution and techniques from convex optimization.
result A concise convergence rate analysis of the mean field Langevin dynamics in both continuous and discrete time settings.
On-line portfolio selection has attracted increasing interests in machine learning and AI communities recently. Empirical evidences show that stock's high and low prices are temporary and stock price relatives are likely to follow the mean reversion phenomenon. While the existing mean reversion strategies are shown to …
This work extends Ledoit-Wolf shrinkage to unknown mean covariance estimation.
problem Large dimensional covariance matrix estimation with unknown mean under Kolmogorov asymptotics.
method Extending Ledoit-Wolf linear shrinkage to translation-invariant estimators, proving their convergence properties.
result A new estimator outperforms other standard estimators empirically.
This paper extends Median-of-Means to new learning problems involving pairwise comparisons.
problem Learning from pairwise comparisons in machine learning.
method Segmenting data into blocks, comparing pairs of decision rules, and declaring the winner based on majority performance.
result The Median-of-Means approach maintains robustness and performance under various sampling schemes.
Simple private estimators for mean and covariance outperform existing methods.
problem Private estimation of mean and covariance at small sample sizes.
method Differentially private estimators for multivariate sub-Gaussian data.
result Asymptotic error rates match theoretical bounds and outperform previous methods.
Transformers solve Poisson means estimation via empirical Bayes.
problem Estimating Poisson means under empirical Bayes setting.
method Pre-trained transformer learns to adapt to unknown prior and do in-context learning.
result Transformers achieve vanishing regret with large models and outperform classical algorithms.
Estimates multiple means in high dimensions using convex combinations.
problem Estimating multiple multi-dimensional means from samples.
method Convex combinations of empirical means with data-dependent weights.
result Our methods asymptotically approach oracle (minimax) improvement.
The paper assesses conditions for the uniqueness of k-means clustering.
problem Conditions for the uniqueness of k-means clustering.
method Analyzes the choice of k and provides necessary and sufficient conditions for uniqueness.
result Determines the asymptotic distribution of the within cluster sum of squares (WCSS) and provides a bootstrap test for uniqueness.
A mean function in a reproducing kernel Hilbert space (RKHS), or a kernel mean, is central to kernel methods in that it is used by many classical algorithms such as kernel principal component analysis, and it also forms the core inference step of modern kernel methods that rely on embedding probability distributions in…
We consider a system of diffusion processes that interact through their empirical mean and have a stabilizing force acting on each of them, corresponding to a bistable potential. There are three parameters that characterize the system: the strength of the intrinsic stabilization, the strength of the external random per…
RL approach for continuous-time mean-variance portfolio selection with empirical validation.
problem Continuous-time mean-variance portfolio selection in unknown market coefficients.
method Reinforcement learning for diffusion processes, sublinear regret bound derivation.
result RL strategy consistently outperforms model-based counterparts, especially in volatile markets.
Study shows mean-field approximation fails to improve PAC-Bayes bounds for neural networks.
problem Understanding why overparametrized neural networks achieve low risk and zero empirical risk.
method Optimized PAC-Bayes bounds using variational inference (VI), investigating mean-field approximation.
result Mean-field approximation does not provide significant improvements in PAC-Bayes bounds for neural networks.
Study finds mean reversion strategies perform well on historical data but fail in recent market conditions.
problem Performance of mean reversion strategies in recent market data.
method Empirical investigation of three mean reversion strategies (PAMR, OLMAR, TCO) on historical S&P 500 data and benchmark datasets.
result Mean reversion strategies may fail in recent market conditions, especially with transaction costs.
Paper shows robust estimators converge to true risk minimizers at optimal rates.
problem Understanding asymptotic properties of robust risk minimizers.
method Investigates robust analogues of empirical risk minimization, focusing on median of means estimator.
result Robust minimizers converge to true minimizers at optimal rates and have similar asymptotic variance.
K-means clustering improved for robustness to outliers and distribution shifts.
problem K-means is brittle to outliers, distribution shifts, and limited samples.
method Developed a distributionally robust variant using Wasserstein-2 ball around the empirical distribution.
result Substantial gains in outlier detection and robustness to noise demonstrated.
Study shows how neural networks generalize with minimal training data.
problem Understanding how neural networks generalize with limited data.
method Mean-field analysis of KL-regularized empirical risk minimization.
result Generalization error rate is O(1/n) for large n. We propose an empirical Bayes estimator based on Dirichlet process mixture model for estimating the sparse normalized mean difference, which could be directly applied to the high dimensional linear classification. In theory, we build a bridge to connect the estimation error of the mean difference and the misclassificat…
A new sequential method estimates Poisson means in streaming data, achieving optimality and efficiency.
problem Estimating Poisson means in a streaming, or online, framework.
method A quasi-Bayesian approach based on Newton's algorithm for a sequential estimate.
result Established frequentist guarantees including consistency and asymptotic optimality.
We extend the empirical results published in article "Empirical Evidence on Arbitrage by Changing the Stock Exchange" by means of machine learning and advanced econometric methodologies based on Smooth Transition Regression models and Artificial Neural Networks.
The Normal Means problem plays a fundamental role in many areas of modern high-dimensional statistics, both in theory and practice. And the Empirical Bayes (EB) approach to solving this problem has been shown to be highly effective, again both in theory and practice. However, almost all EB treatments of the Normal Mean…
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We show that universal consistency of Empirical Risk Minimiza…
A new method of moments estimator goes beyond data reweighting.
problem Estimation of moment restrictions and conditional moment restrictions.
method Kernel Method of Moments (KMM) based on maximum mean discrepancy.
result KMM achieves competitive performance on conditional moment restriction tasks.
We investigate the variety of a portfolio of stocks in normal and extreme days of market activity. We show that the variety carries information about the market activity which is not present in the single-index model and we observe that the variety time evolution is not time reversal around the crash days. We obtain th…
We study theoretical and empirical aspects of the mean exit time of financial time series. The theoretical modeling is done within the framework of continuous time random walk. We empirically verify that the mean exit time follows a quadratic scaling law and it has associated a pre-factor which is specific to the analy…
A microscopic model of aggregation and fragmentation is introduced to investigate the size distribution of businesses. In the model, businesses are constrained to comply with the market price, as expected by the customers, while customers can only buy at the prices offered by the businesses. We show numerically and ana…
Kernel k-means clustering can correctly identify and extract a far more varied collection of cluster structures than the linear k-means clustering algorithm. However, kernel k-means clustering is computationally expensive when the non-linear feature map is high-dimensional and there are many input points. Kernel …
New algorithm for biclustering with improved performance.
problem Simultaneous clustering of rows and columns with similar patterns.
method Formulated new biclustering problem, developed alternating k-means algorithm.
result Our algorithm finds local minima efficiently and outperforms other methods.
Paper improves CI and CS for bounded means using betting and mixtures.
problem Estimating means of bounded random variables.
method Composite nonnegative martingales, testing by betting, method of mixtures.
result Empirically outperforms existing CI and CS methods.
Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.
problem Understanding variations in the 80/20 rule across different contexts.
method Identifying the statistical distribution of the 80/20 rule and its variations.
result The 80/20 rule follows a Gaussian distribution with a standard deviation twice the mean.
The paper investigates how class mean vectors enhance neural network classification.
problem Improving neural network performance in classification tasks.
method Exploring the role of class mean vectors in neural networks, including direct computation of weights, performance monitoring, and self-training.
result Empirical evidence suggests that using class mean vectors can significantly improve neural network performance on classification tasks.
Optimizes sparse mean-reverting portfolios for higher returns.
problem Finding optimal stock weights for mean-reverting portfolios.
method Transformed optimization problem into SDP, added constraints.
result Sparse mean-reverting portfolios provide higher returns with transaction costs.
Recent empirical studies suggest that the volatilities associated with financial time series exhibit short-range correlations. This entails that the volatility process is very rough and its autocorrelation exhibits sharp decay at the origin. Another classic stylistic feature often assumed for the volatility is that it …
New DP methods for estimating means and frequencies with varying privacy demands.
problem Estimating statistics with users having different privacy requirements.
method Proposes algorithms for empirical mean and frequency estimation under heterogeneous privacy constraints, considering both correlated and permuted datasets.
result Establishes theoretical performance guarantees for algorithms, achieving minimax optimality.
Optimal benchmark design varies based on costs in financial manipulation.
problem Manipulation of price benchmarks in finance.
method Analyzes empirical pattern and cost structures to determine optimal benchmark design.
result The optimal benchmark depends on the relative sizes of fixed and variable costs.
Few-shot learning aims to train efficient predictive models with a few examples. The lack of training data leads to poor models that perform high-variance or low-confidence predictions. In this paper, we propose to meta-learn the ensemble of epoch-wise empirical Bayes models (E3BM) to achieve robust predictions. "Epoch…