Optimizes sparse mean-reverting portfolios for higher returns.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A faster Wasserstein k-means algorithm for histogram data reduces computation and maintains clustering quality.
New algorithm reduces runtime for robust sparse mean estimation.
New algorithm recovers sparse measures in polynomial time.
Efficiently estimates sparse mean from heavy-tailed data.
New method for estimating sparse means in noisy data.
New method estimates sparse mean from noisy data without knowing sparsity level.
New model approximates sparse mean-CVaR portfolio optimization efficiently.
Simple, scalable sparse k-means for high-dimensional data.
Proposes ARSK for robust and sparse clustering.
Robust estimators for Gaussian sparse tasks with optimal error under contamination.
We compute an approximate Fréchet mean for sets of sparse graphs.
In many situations where the interest lies in identifying clusters one might expect that not all available variables carry information about these groups. Furthermore, data quality (e.g. outliers or missing entries) might present a serious and sometimes hard-to-assess problem for large and complex datasets. In this pap…
Kernel means are frequently used to represent probability distributions in machine learning problems. In particular, the well known kernel density estimator and the kernel mean embedding both have the form of a kernel mean. Unfortunately, kernel means are faced with scalability issues. A single point evaluation of the …
We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably dominates its conventional counterpart in terms of mean square deviations. We es…
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to…
New methods show sparse portfolios offer no advantage over mean-variance in diversification.
We create interpretable word embeddings through sparse coding.
New DP optimization methods for sparse gradients, improving on existing algorithms.
New methods solve sparse estimation robustly, even with outliers.
IVF k-means algorithm improves performance on large sparse data sets.
We propose the Lasso Weighted -means (--means) algorithm as a simple yet efficient sparse clustering procedure for high-dimensional data where the number of features () can be much larger compared to the number of observations (). In the --means algorithm, we introduce a lasso-based penalty term,…
New algorithms reduce communication for sparse mean estimation in noisy distributed systems.
Sparse clustering, which aims to find a proper partition of an extremely high-dimensional data set with redundant noise features, has been attracted more and more interests in recent years. The existing studies commonly solve the problem in a framework of maximizing the weighted feature contributions subject to a $\ell…
We propose to align distributional data from the perspective of Wasserstein means. We raise the problem of regularizing Wasserstein means and propose several terms tailored to tackle different problems. Our formulation is based on the variational transportation to distribute a sparse discrete measure into the target do…
We propose a version of least-mean-square (LMS) algorithm for sparse system identification. Our algorithm called online linearized Bregman iteration (OLBI) is derived from minimizing the cumulative prediction error squared along with an l1-l2 norm regularizer. By systematically treating the non-differentiable regulariz…
Proposes a semi-supervised K-Means algorithm for better feature selection.
SIVF k-means algorithm speeds up sparse data clustering.
Privacy improves robustness in statistical estimation.
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
Paper accelerates K-means clustering for large sparse document data.
Flexible Bayesian approach for generalized linear models, especially for sparse logistic regression.
Sparse matrices simplify computation of GP variances and likelihoods.
While several papers have investigated computationally and statistically efficient methods for learning Gaussian mixtures, precise minimax bounds for their statistical performance as well as fundamental limits in high-dimensional settings are not well-understood. In this paper, we provide precise information theoretic …
Consider the problem of sparse clustering, where it is assumed that only a subset of the features are useful for clustering purposes. In the framework of the COSA method of Friedman and Meulman, subsequently improved in the form of the Sparse K-means method of Witten and Tibshirani, a natural and simpler hill-climbing …
Wide neural networks learn features under P, identifying weights and decomposing support.
Attention mechanisms and non-local mean operations in general are key ingredients in many state-of-the-art deep learning techniques. In particular, the Transformer model based on multi-head self-attention has recently achieved great success in natural language processing and computer vision. However, the vanilla algori…
Sparse coding has shown its power as an effective data representation method. However, up to now, all the sparse coding approaches are limited within the single domain learning problem. In this paper, we extend the sparse coding to cross domain learning problem, which tries to learn from a source domain to a target dom…
BPASGM uses sparse graphical models to optimize portfolio selection.
We study the problem of high-dimensional sparse mean estimation in the presence of an -fraction of adversarial outliers. Prior work obtained sample and computationally efficient algorithms for this task for identity-covariance subgaussian distributions. In this work, we develop the first efficient algorithms for rob…
Paper develops a new algorithm for sparse signal recovery.
Many conventional statistical procedures are extremely sensitive to seemingly minor deviations from modeling assumptions. This problem is exacerbated in modern high-dimensional settings, where the problem dimension can grow with and possibly exceed the sample size. We consider the problem of robust estimation of sparse…
We study the tradeoff between the statistical error and communication cost of distributed statistical estimation problems in high dimensions. In the distributed sparse Gaussian mean estimation problem, each of the machines receives data points from a -dimensional Gaussian distribution with unknown mean w…
New method for community detection in sparse directed SBMs with exact recovery guarantees.
Paper learns Koopman operator from sparse data, escaping function space constraints.
The paper develops adaptive deep learning methods for nonlinear time series models.
Signal processing tasks as fundamental as sampling, reconstruction, minimum mean-square error interpolation and prediction can be viewed under the prism of reproducing kernel Hilbert spaces. Endowing this vantage point with contemporary advances in sparsity-aware modeling and processing, promotes the nonparametric basi…
We study a mean-field spike and slab variational Bayes (VB) approximation to Bayesian model selection priors in sparse high-dimensional linear regression. Under compatibility conditions on the design matrix, oracle inequalities are derived for the mean-field VB approximation, implying that it converges to the sparse tr…