The statistical leverage scores of a complex matrix record the degree of alignment between col and the coordinate axes in . These score are used in random sampling algorithms for solving certain numerical linear algebra problems. In this paper we present a max-plus algebr…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficiently approximates statistical leverage scores for faster KRR.
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion tha…
New algorithms estimate matrix leverage scores using rank revealing and randomization.
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
In this paper, we consider the problem of column subset selection. We present a novel analysis of the spectral norm reconstruction for a simple randomized algorithm and establish a new bound that depends explicitly on the sampling probabilities. The sampling dependent error bound (i) allows us to better understand the …
In this paper, we consider a statistical problem of learning a linear model from noisy samples. Existing work has focused on approximating the least squares solution by using leverage-based scores as an importance sampling distribution. However, no finite sample statistical guarantees and no computationally efficient o…
A new robust PCA method uses Innovation Search and Leverage Scores.
Regularizes attention scores in vision transformers using bootstrapping.
OTCP extends conformal prediction to multivariate data using optimal transport.
Ridge leverage scores provide a balance between low-rank approximation and regularization, and are ubiquitous in randomized linear algebra and machine learning. Deterministic algorithms are also of interest in the moderately big data regime, because deterministic algorithms provide interpretability to the practitioner …
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…
One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the method of algorithmic leveraging samples and rescales rows/columns of data matrices to reduce the data size before performing computations on…
Proposes SD-KDE for density estimation using debiased kernel density with score-based adjustments.
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
Although deep learning has been applied to successfully address many data mining problems, relatively limited work has been done on deep learning for anomaly detection. Existing deep anomaly detection methods, which focus on learning new feature representations to enable downstream anomaly detection methods, perform in…
We consider estimating the parametric components of semi-parametric multiple index models in a high-dimensional and non-Gaussian setting. Such models form a rich class of non-linear models with applications to signal processing, machine learning and statistics. Our estimators leverage the score function based first and…
Credit scoring is without a doubt one of the oldest applications of analytics. In recent years, a multitude of sophisticated classification techniques have been developed to improve the statistical performance of credit scoring models. Instead of focusing on the techniques themselves, this paper leverages alternative d…
One approach to improving the running time of kernel-based machine learning methods is to build a small sketch of the input and use it in lieu of the full kernel matrix in the machine learning task of interest. Here, we describe a version of this approach that comes with running time guarantees as well as improved guar…
Binary testing for softmax models requires many samples, similar to leverage score models.
Universal approach combines OOD detection scores for robustness.
New method detects and locates changes in spatio-temporal point processes.
Improved sample complexity for diffusion models without needing empirical risk minimizers.
Improves generative model coverage of underrepresented modes.
Extends importance sampling to nonlinear models using adjoint operators.
SALSA efficiently approximates leverage scores for big data, improving ARMA model fitting.
Random features provide a practical framework for large-scale kernel approximation and supervised learning. It has been shown that data-dependent sampling of random features using leverage scores can significantly reduce the number of features required to achieve optimal learning bounds. Leverage scores introduce an op…
We derive optimal statistical and computational complexity bounds for exp-concave stochastic minimization in terms of the effective dimension. For common eigendecay patterns of the population covariance matrix, this quantity is significantly smaller than the ambient dimension. Our results reveal interesting connections…
Paper develops fine-grain spatiotemporal risk scores using high-resolution mobility data.
We consider the problem of exact recovery of any matrix of rank from a small number of observed entries via the standard nuclear norm minimization framework. Such low-rank matrices have degrees of freedom . We show that any arbitrary low-rank matrices can be recovered exa…
RealStats detects fake images rigorously, combining multiple detectors for robustness.
Latent DiTs improve data distribution recovery and inference efficiency under low-dimensional latent space.
We give a fast oblivious L2-embedding of to satisfying Our embedding dimension equals , a constant independent of the distortion . We use as a black-box any L2-embedding $Π…
In order to mimic the human ability of continual acquisition and transfer of knowledge across various tasks, a learning system needs the capability for continual learning, effectively utilizing the previously acquired skills. As such, the key challenge is to transfer and generalize the knowledge learned from one task t…
Detects changes in classifier scores to identify shifts in class priors.
In statistical analysis, measuring a score of predictive performance is an important task. In many scientific fields, appropriate scores were tailored to tackle the problems at hand. A proper score is a popular tool to obtain statistically consistent forecasts. Furthermore, a mathematical characterization of the proper…
Optimal sampling reduces power grid data analysis costs.
New algorithm converts data into sub-gaussian designs efficiently.
Random Fourier features is a widely used, simple, and effective technique for scaling up kernel methods. The existing theoretical analysis of the approach, however, remains focused on specific learning tasks and typically gives pessimistic bounds which are at odds with the empirical results. We tackle these problems an…
A new algorithm approximates logistic regression probabilities efficiently.
GAS-Norm improves deep learning time series forecasting in non-stationary settings.
New method combines score lists using joint CDFs, improving computation.
Score attack method provides a lower bound on privacy-constrained minimax risk.
New methods ensure feature importance rankings are correct with high probability.
A new method uses LLMs to discover causal pathways that affect fairness in machine learning.
This paper improves matrix completion by leveraging element importance and non-uniform sampling.