Improved SRHT for linear SVM classification with higher accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified methodology for statistical inference in least squares and PCA via randomized sketching.
A well-known problem in data science and machine learning is {\em linear regression}, which is recently extended to dynamic graphs. Existing exact algorithms for updating the solution of dynamic graph regression require at least a linear time (in terms of : the size of the graph). However, this time complexity might…
A new iterative low complexity algorithm has been presented for computing the Walsh-Hadamard transform (WHT) of an dimensional signal with a -sparse WHT, where is a power of two and , scales sub-linearly in for some . Assuming a random support model for the non-zero transform domain…
Sketching, a dimensionality reduction technique, has received much attention in the statistics community. In this paper, we study sketching in the context of Newton's method for solving finite-sum optimization problems in which the number of variables and data points are both large. We study two forms of sketching that…
Linear mixed models (LMMs) are used extensively to model dependecies of observations in linear regression and are used extensively in many application areas. Parameter estimation for LMMs can be computationally prohibitive on big data. State-of-the-art learning algorithms require computational complexity which depends …
We consider a least squares regression problem where the data has been generated from a linear model, and we are interested to learn the unknown regression parameters. We consider "sketch-and-solve" methods that randomly project the data first, and do regression after. Previous works have analyzed the statistical and c…
Paper computes link determinants using Fourier-Hadamard transforms.
Uniform approximations for RHTs improve kernel approximation and distance estimation.
New algorithm reduces sketching dimension to effective problem size.
A new class of harmonic Hadamard manifolds, those spaces called of hypergeometric type, is defined in terms of Gauss hypergeometric equations. Spherical Fourier transform defined on a harmonic Hadamard manifold of hypergeometric type admits an inversion formula. A characterization of harmonic Hadamard manifold being of…
DFRot improves LLMs by reducing outlier and massive activation effects.
We prove two injectivity theorems for the geodesic ray transform on two-dimensional, complete, simply connected Riemannian manifolds with non-positive Gaussian curvature, also known as Cartan-Hadamard manifolds. The first theorem is concerned with bounded non-positive curvature and the second with decaying non-positive…
New insights into identifying mixtures of product distributions using Hadamard extensions.
In this paper, we study random subsampling of Gaussian process regression, one of the simplest approximation baselines, from a theoretical perspective. Although subsampling discards a large part of training data, we show provable guarantees on the accuracy of the predictive mean/variance and its generalization ability.…
New insights into how randomization affects greedy model selection.
Study fourth order Schrödinger equation on Cartan-Hadamard manifolds, proving existence, scattering, and blow-up results.
We study the geodesic X-ray transform on Cartan-Hadamard manifolds, and prove solenoidal injectivity of this transform acting on functions and tensor fields of any order. The functions are assumed to be exponentially decaying if the sectional curvature is bounded, and polynomially decaying if the sectional curvature de…
Enhances random forest performance with exogenous randomness.
We study the problem of subsampling in differential privacy (DP), a question that is the centerpiece behind many successful differentially private machine learning algorithms. Specifically, we provide a tight upper bound on the Rényi Differential Privacy (RDP) (Mironov, 2017) parameters for algorithms that: (1) subsamp…
Geometrically proves majorizing measure theorem on Hadamard manifolds.
Differential privacy comes equipped with multiple analytical tools for the design of private data analyses. One important tool is the so-called "privacy amplification by subsampling" principle, which ensures that a differentially private mechanism run on a random subsample of a population provides higher privacy guaran…
Early stopping is a well known approach to reduce the time complexity for performing training and model selection of large scale learning machines. On the other hand, memory/space (rather than time) complexity is the main constraint in many applications, and randomized subsampling techniques have been proposed to tackl…
A new neural subsampling method reduces data volume for deep models.
New sampling scheme improves privacy in DP-SGD without sacrificing utility.
Unified theory and debiasing framework for random oblique projections in high dimensions.
Data-driven discovery of differential equations has been an emerging research topic. We propose a novel algorithm subsampling-based threshold sparse Bayesian regression (SubTSBR) to tackle high noise and outliers. The subsampling technique is used for improving the accuracy of the Bayesian learning algorithm. It has tw…
Paper bridges statistical inference for DP-SGD, a privacy-preserving machine learning method.
A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample from the original full sample and uses it as a surrogate for subsequent computati…
A new model-free subsampling method using uniform designs is proposed.
Variational inference offers scalable and flexible tools to tackle intractable Bayesian inference of modern statistical models like Bayesian neural networks and Gaussian processes. For largely over-parameterized models, however, the over-regularization property of the variational objective makes the application of vari…
Statistical machine learning models should be evaluated and validated before putting to work. Conventional k-fold Monte Carlo Cross-Validation (MCCV) procedure uses a pseudo-random sequence to partition instances into k subsets, which usually causes subsampling bias, inflates generalization errors and jeopardizes the r…
Unified framework for subsampling mechanisms with tighter privacy guarantees.
Principal Components Regression (PCR) is a traditional tool for dimension reduction in linear regression that has been both criticized and defended. One concern about PCR is that obtaining the leading principal components tends to be computationally demanding for large data sets. While random projections do not possess…
The Weyl transform is introduced as a rich framework for data representation. Transform coefficients are connected to the Walsh-Hadamard transform of multiscale autocorrelations, and different forms of dyadic periodicity in a signal are shown to appear as different features in its Weyl coefficients. The Weyl transform …
Random forests remain among the most popular off-the-shelf supervised learning algorithms. Despite their well-documented empirical success, however, until recently, few theoretical results were available to describe their performance and behavior. In this work we push beyond recent work on consistency and asymptotic no…
We introduce extensions of stability selection, a method to stabilise variable selection methods introduced by Meinshausen and Bühlmann (J R Stat Soc 72:417-473, 2010). We propose to apply a base selection method repeatedly to random observation subsamples and covariate subsets under scrutiny, and to select covariates …
Develops an empirical likelihood framework for random forests and ensembles.
Random forests have proven to be reliable predictive algorithms in many application areas. Not much is known, however, about the statistical properties of random forests. Several authors have established conditions under which their predictions are consistent, but these results do not provide practical estimates of ran…
Extends DCP framework to Hadamard manifolds for geodesically convex functions.
This paper explains CART random forests using stochastic control theory.
Corrects bias in random sampling matrices for improved ML methods.
We study Nyström type subsampling approaches to large scale kernel methods, and prove learning bounds in the statistical learning setting, where random sampling and high probability estimates are considered. In particular, we prove that these approaches can achieve optimal learning bounds, provided the subsampling leve…
We prove, using the subspace embedding guarantee in a black box way, that one can achieve the spectral norm guarantee for approximate matrix multiplication with a dimensionality-reducing map having rows. Here is the maximum stable rank, i.e. squared ratio of Frobenius and op…
Improved DNN estimator with scalable subsampling for efficient inference.
The infinitesimal jackknife (IJ) has recently been applied to the random forest to estimate its prediction variance. These theorems were verified under a traditional random forest framework which uses classification and regression trees (CART) and bootstrap resampling. However, random forests using conditional inferenc…
We propose Subsampling MCMC, a Markov Chain Monte Carlo (MCMC) framework where the likelihood function for observations is estimated from a random subset of observations. We introduce a highly efficient unbiased estimator of the log-likelihood based on control variates, such that the computing cost is much smal…
Develops a faster model selection method using influence functions.