The statistical leverage scores of a complex matrix record the degree of alignment between col and the coordinate axes in . These score are used in random sampling algorithms for solving certain numerical linear algebra problems. In this paper we present a max-plus algebr…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
Efficiently approximates statistical leverage scores for faster KRR.
Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
New algorithms estimate matrix leverage scores using rank revealing and randomization.
Binary testing for softmax models requires many samples, similar to leverage score models.
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion tha…
Improves generative model coverage of underrepresented modes.
Extends importance sampling to nonlinear models using adjoint operators.
SALSA efficiently approximates leverage scores for big data, improving ARMA model fitting.
Random features provide a practical framework for large-scale kernel approximation and supervised learning. It has been shown that data-dependent sampling of random features using leverage scores can significantly reduce the number of features required to achieve optimal learning bounds. Leverage scores introduce an op…
A new robust PCA method uses Innovation Search and Leverage Scores.
We consider the problem of exact recovery of any matrix of rank from a small number of observed entries via the standard nuclear norm minimization framework. Such low-rank matrices have degrees of freedom . We show that any arbitrary low-rank matrices can be recovered exa…
This paper improves matrix completion by leveraging element importance and non-uniform sampling.
The power of randomized algorithms in numerical methods have led to fast solutions which use the Singular Value Decomposition (SVD) as a core routine. However, given the large data size of modern and the modest runtime of SVD, most practical algorithms would require some form of approximation, such as sketching, when r…
In this paper, we consider the problem of column subset selection. We present a novel analysis of the spectral norm reconstruction for a simple randomized algorithm and establish a new bound that depends explicitly on the sampling probabilities. The sampling dependent error bound (i) allows us to better understand the …
Ridge leverage scores provide a balance between low-rank approximation and regularization, and are ubiquitous in randomized linear algebra and machine learning. Deterministic algorithms are also of interest in the moderately big data regime, because deterministic algorithms provide interpretability to the practitioner …
OTCP extends conformal prediction to multivariate data using optimal transport.
BSAC improves credit scoring models by leveraging autoencoders and addressing imbalanced datasets.
We apply methods from randomized numerical linear algebra (RandNLA) to develop improved algorithms for the analysis of large-scale time series data. We first develop a new fast algorithm to estimate the leverage scores of an autoregressive (AR) model in big data regimes. We show that the accuracy of approximations lies…
Improves score estimation for noised targets using known clean scores.
New algorithm samples matrix rows proportional to their ℓ_p norm in a turnstile data stream.
For any matrix A in R^(m x n) of rank ρ, we present a probability distribution over the entries of A (the element-wise leverage scores of equation (2)) that reveals the most influential entries in the matrix. From a theoretical perspective, we prove that sampling at most s = O ((m + n) ρ^2 ln (m + n)) entries of the ma…
ScoreMatchingRiesz improves debiased machine learning and policy effects estimation.
We are interested in the task of generating multi-instrumental music scores. The Transformer architecture has recently shown great promise for the task of piano score generation; here we adapt it to the multi-instrumental setting. Transformers are complex, high-dimensional language models which are capable of capturing…
Given a matrix and a vector , we show how to compute an -approximate solution to the regression problem in time where …
Neural score matching improves high-dimensional causal inference by using neural networks for balancing scores.
Although deep learning has been applied to successfully address many data mining problems, relatively limited work has been done on deep learning for anomaly detection. Existing deep anomaly detection methods, which focus on learning new feature representations to enable downstream anomaly detection methods, perform in…
Paper proposes a new method for training diffusion models using Markov operators.
Paper introduces a fast, robust, scalable method for detecting changes in data streams.
Study improves denoising score matching under relaxed manifold assumptions.
Proposes SD-KDE for density estimation using debiased kernel density with score-based adjustments.
Regularizes attention scores in vision transformers using bootstrapping.
A new method for Bayesian inference using diffusion models.
Novel estimator reduces diffusion model variance.
Knowledge graph based simple question answering (KBSQA) is a major area of research within question answering. Although only dealing with simple questions, i.e., questions that can be answered through a single knowledge base (KB) fact, this task is neither simple nor close to being solved. Targeting on the two main ste…
Estimates score function from data with optimal rate in high dimensions.
Paper proposes a method to reduce hallucinations in diffusion models using Laplacian score sharpening.
Semi-Implicit Variational Inference (SIVI) is improved with SIVI-SM using score matching.
We study the problem of interpreting trained classification models in the setting of linguistic data sets. Leveraging a parse tree, we propose to assign least-squares based importance scores to each word of an instance by exploiting syntactic constituency structure. We establish an axiomatic characterization of these i…
Paper proposes a new method for estimating treatment effects using interpretable deep learning models.
Paper tackles transparency and auditability of machine learning in credit scoring.
We give the first algorithm for kernel Nyström approximation that runs in *linear time in the number of training points* and is provably accurate for all kernel matrices, without dependence on regularity or incoherence conditions. The algorithm projects the kernel onto a set of landmark points sampled by their *rid…
This paper enhances uplift modeling for multi-treatment marketing campaigns.
New method assesses individual training points' privacy risk without retraining.