This note gives a simple analysis of a randomized approximation scheme for matrix multiplication proposed by Sarlos (2006) based on a random rotation followed by uniform column sampling. The result follows from a matrix version of Bernstein's inequality and a tail inequality for quadratic forms in subgaussian random ve…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper explores how random sampling and coding can speed up approximate matrix multiplication.
New method uses randomized sparse neural networks to solve time-dependent PDEs more accurately and efficiently.
Optimized sampling scheme for compressed sensing combining randomness and determinism.
New gradient coding schemes reduce decoding error in both random and adversarial straggler settings.
Improved Random Forests detect pure interactions better.
New high-order approximations for CIR process using random grids.
Low-rank structure have been profoundly studied in data mining and machine learning. In this paper, we show a dense matrix 's low-rank approximation can be rapidly built from its left and right random projections and , or bilateral random projection (BRP). We then show power scheme can further…
New schemes improve error estimates for sampling from non-log-concave distributions.
The method of random projections has become a standard tool for machine learning, data mining, and search with massive data at Web scale. The effective use of random projections requires efficient coding schemes for quantizing (real-valued) projected data into integers. In this paper, we focus on a simple 2-bit coding …
Study different masking schemes for a universal marginaliser.
Aggregates predictions from multiple regression models using random projections and kernel methods.
We consider the binary classification problem when data are large and subject to unknown but bounded uncertainties. We address the problem by formulating the nonlinear support vector machine training problem with robust optimization. To do so, we analyze and propose two bounding schemes for uncertainties associated to …
New sampling scheme improves privacy in DP-SGD without sacrificing utility.
New defense mechanism RS outperforms existing schemes in protecting against adversarial examples.
Study shows different initialization schemes for LoRA finetuning impact performance.
Two new coding schemes improve the efficient communication of noisy data.
1) We introduce random discrete Morse theory as a computational scheme to measure the complicatedness of a triangulation. The idea is to try to quantify the frequence of discrete Morse matchings with a certain number of critical cells. Our measure will depend on the topology of the space, but also on how nicely the spa…
Hash codes are efficient data representations for coping with the ever growing amounts of data. In this paper, we introduce a random forest semantic hashing scheme that embeds tiny convolutional neural networks (CNN) into shallow random forests, with near-optimal information-theoretic code aggregation among trees. We s…
We develop high-order approximations for the Heston model.
In this work, we revisit fast dimension reduction approaches, as with random projections and random sampling. Our goal is to summarize the data to decrease computational costs and memory footprint of subsequent analysis. Such dimension reduction can be very efficient when the signals of interest have a strong structure…
The nowadays massive amounts of generated and communicated data present major challenges in their processing. While capable of successfully classifying nonlinearly separable objects in various settings, subspace clustering (SC) methods incur prohibitively high computational complexity when processing large-scale data. …
Novel framework improves randomized smoothing for various norms.
This paper develops quantization algorithms for random Fourier features, simplifying the process and improving performance.
New method compresses random forest models for efficient storage.
Bayesian optimization uses triangulation candidates for better performance.
Unified theory and debiasing framework for random oblique projections in high dimensions.
Improved statistical computation through efficient matrix sampling.
This paper proposes a framework for certifying neural network defenses against data poisoning attacks.
Many modern machine learning models are trained to achieve zero or near-zero training error in order to obtain near-optimal (but non-zero) test error. This phenomenon of strong generalization performance for "overfitted" / interpolated classifiers appears to be ubiquitous in high-dimensional data, having been observed …
Hash codes are a very efficient data representation needed to be able to cope with the ever growing amounts of data. We introduce a random forest semantic hashing scheme with information-theoretic code aggregation, showing for the first time how random forest, a technique that together with deep learning have shown spe…
Boosts change-point detection power with optimal sub-sampling.
Hashing is a basic tool for dimensionality reduction employed in several aspects of machine learning. However, the perfomance analysis is often carried out under the abstract assumption that a truly random unit cost hash function is used, without concern for which concrete hash function is employed. The concrete hash f…
We present a new random sampling strategy for k-bandlimited signals defined on graphs, based on determinantal point processes (DPP). For small graphs, ie, in cases where the spectrum of the graph is accessible, we exhibit a DPP sampling scheme that enables perfect recovery of bandlimited signals. For large graphs, ie, …
The paper analyzes the Spectral Method for clustering data points on Union of Subspaces.
RGBM improves GBM efficiency with randomization.
Consider convex optimization problems subject to a large number of constraints. We focus on stochastic problems in which the objective takes the form of expected values and the feasible set is the intersection of a large number of convex sets. We propose a class of algorithms that perform both stochastic gradient desce…
We systematically investigate the problem of representing Markov chains by families of random maps, and which regularity of these maps can be achieved depending on the properties of the probability measures. Our key idea is to use techniques from optimal transport to select optimal such maps. Optimal transport theory a…
SVD-RND detects blurred images better than conventional methods.
The problem of using observed correlations to infer causal relations is relevant to a wide variety of scientific disciplines. Yet given correlations between just two classical variables, it is impossible to determine whether they arose from a causal influence of one on the other or a common cause influencing both, unle…
We present algorithms for topic modeling based on the geometry of cross-document word-frequency patterns. This perspective gains significance under the so called separability condition. This is a condition on existence of novel-words that are unique to each topic. We present a suite of highly efficient algorithms based…
Deep learning relies on good initialization schemes and hyperparameter choices prior to training a neural network. Random weight initializations induce random network ensembles, which give rise to the trainability, training speed, and sometimes also generalization ability of an instance. In addition, such ensembles pro…
A new method for approximating softmax and Gaussian kernels with reduced error.
Sharp asymptotic lower bounds of the expected quadratic variation of discretization error in stochastic integration are given. The theory relies on inequalities for the kurtosis and skewness of a general random variable which are themselves seemingly new. Asymptotically efficient schemes which attain the lower bounds a…
Novel numerical scheme for G-heat equation with uncertainty.
A new insurance and reinsurance pricing scheme based on realized loss.
We introduce an efficient message passing scheme for solving Constraint Satisfaction Problems (CSPs), which uses stochastic perturbation of Belief Propagation (BP) and Survey Propagation (SP) messages to bypass decimation and directly produce a single satisfying assignment. Our first CSP solver, called Perturbed Blief …
We propose a probabilistic numerical algorithm to solve Backward Stochastic Differential Equations (BSDEs) with nonnegative jumps, a class of BSDEs introduced in [9] for representing fully nonlinear HJB equations. In particular, this allows us to numerically solve stochastic control problems with controlled volatility,…