Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the top- ranking problem where the goal is to recover the set of top- ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…
New method assesses individual training points' privacy risk without retraining.
New metric scores perturbations across populations, not cells, improving model comparison.
Paper develops fine-grain spatiotemporal risk scores using high-resolution mobility data.
The statistical leverage scores of a complex matrix record the degree of alignment between col and the coordinate axes in . These score are used in random sampling algorithms for solving certain numerical linear algebra problems. In this paper we present a max-plus algebr…
With the development of high-throughput technologies, principal component analysis (PCA) in the high-dimensional regime is of great interest. Most of the existing theoretical and methodological results for high-dimensional PCA are based on the spiked population model in which all the population eigenvalues are equal ex…
Credit scoring models support loan approval decisions in the financial services industry. Lenders train these models on data from previously granted credit applications, where the borrowers' repayment behavior has been observed. This approach creates sample bias. The scoring model (i.e., classifier) is trained on accep…
We derive optimal statistical and computational complexity bounds for exp-concave stochastic minimization in terms of the effective dimension. For common eigendecay patterns of the population covariance matrix, this quantity is significantly smaller than the ambient dimension. Our results reveal interesting connections…
Many in-hospital mortality risk prediction scores dichotomize predictive variables to simplify the score calculation. However, hard thresholding in these additive stepwise scores of the form "add x points if variable v is above/below threshold t" may lead to critical failures. In this paper, we seek to develop risk pre…
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…
We study convergence of a generative modeling method that first estimates the score function of the distribution using Denoising Auto-Encoders (DAE) or Denoising Score Matching (DSM) and then employs Langevin diffusion for sampling. We show that both DAE and DSM provide estimates of the score of the Gaussian smoothed p…
Unified causal inference framework using distribution adaptation.
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
Bayesian approach clusters survival data for better risk prediction.
Observational cohort studies with oversampled exposed subjects are typically implemented to understand the causal effect of a rare exposure. Because the distribution of exposed subjects in the sample differs from the source population, estimation of a propensity score function (i.e., probability of exposure given basel…
Paper tackles efficient risk estimation under dataset shift conditions.
Universal approach combines OOD detection scores for robustness.
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
ScoreFusion fuses multiple diffusion models to enhance generative modeling of a target population.
New model estimates species population trends from citizen science data.
Develops a fair classifier for deep learning models.
Efficiently approximates statistical leverage scores for faster KRR.
We consider the \mnk{classical} problem of a controller activating (or sampling) sequentially from a finite number of populations, specified by unknown distributions. Over some time horizon, at each time , the controller wishes to select a population to sample, with the goal of sampling fro…
Trimming helps in conformal prediction when it separates anomaly scores.
New models extrapolate false alarms in ASV without new data.
New algorithms estimate matrix leverage scores using rank revealing and randomization.
A robust conformal method for set estimation using non-conformity scores.
We develop a personalized real time risk scoring algorithm that provides timely and granular assessments for the clinical acuity of ward patients based on their (temporal) lab tests and vital signs. Heterogeneity of the patients population is captured via a hierarchical latent class model. The proposed algorithm aims t…
The paper tackles sampling bias in credit scoring models and proposes methods to improve their training and evaluation.
Proposes a method to infer the distributional impacts of predictive models on stakeholders.
Credit scores misclassify borrowers, especially minorities, leading to inequitable access.
The paper explores various forms of calibration scores and their implications for fairness.
Binary testing for softmax models requires many samples, similar to leverage score models.
Generative method avoids function estimation for data generation.
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion tha…
Improves generative model coverage of underrepresented modes.
The paper introduces a method for detecting principal communities and embedding vertices.
Extends importance sampling to nonlinear models using adjoint operators.
SALSA efficiently approximates leverage scores for big data, improving ARMA model fitting.
LEARNER improves low-rank matrix estimation using source population data.
We address the problem of correcting group discriminations within a score function, while minimizing the individual error. Each group is described by a probability density function on the set of profiles. We first solve the problem analytically in the case of two populations, with a uniform bonus-malus on the zones whe…
Random features provide a practical framework for large-scale kernel approximation and supervised learning. It has been shown that data-dependent sampling of random features using leverage scores can significantly reduce the number of features required to achieve optimal learning bounds. Leverage scores introduce an op…
Proposes a neural network method to combine nonprobability and probability survey samples.
A new robust PCA method uses Innovation Search and Leverage Scores.
We consider the problem of exact recovery of any matrix of rank from a small number of observed entries via the standard nuclear norm minimization framework. Such low-rank matrices have degrees of freedom . We show that any arbitrary low-rank matrices can be recovered exa…
The ROC curve is widely used to assess the quality of prediction/classification/ranking algorithms, and its properties have been extensively studied. The precision-recall (PR) curve has become the de facto replacement for the ROC curve in the presence of imbalance, namely where one class is far more likely than the oth…