Newton-LESS sparsifies Gaussian sketching for faster optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm converts data into sub-gaussian designs efficiently.
A new method reduces the bias in estimating inverse covariance matrices from sketches.
Sparse OSEs achieve optimal embedding dimension of O(d).
Optimal subspace embedding with near-optimal sparsity for high-dimensional data.
EBMs become opaque in high dimensions; LASSO sparsifies them.
Improves generative model coverage of underrepresented modes.
We study the problem of large-scale network embedding, which aims to learn latent representations for network mining applications. Previous research shows that 1) popular network embedding benchmarks, such as DeepWalk, are in essence implicitly factorizing a matrix with a closed form, and 2)the explicit factorization o…
The paper examines the consistency of item embeddings in recommendation systems.
Link prediction is a popular research topic in network analysis. In the last few years, new techniques based on graph embedding have emerged as a powerful alternative to heuristics. In this article, we study the problem of systematic biases in the prediction, and show that some methods based on graph embedding offer le…
Efficiently estimates covariance for sparse functional data.
Estimates mean of distributed vectors with sparsification and spatial/temporal correlations.
Extends importance sampling to nonlinear models using adjoint operators.
We give a fast oblivious L2-embedding of to satisfying Our embedding dimension equals , a constant independent of the distortion . We use as a black-box any L2-embedding $Π…
In this paper, we provide a novel construction of the linear-sized spectral sparsifiers of Batson, Spielman and Srivastava [BSS14]. While previous constructions required running time [BSS14, Zou12], our sparsification routine can be implemented in almost-quadratic running time . The funda…
Embeds sparsity in deep neural networks, allowing exact zero parameters.
We give the first algorithm for kernel Nyström approximation that runs in *linear time in the number of training points* and is provably accurate for all kernel matrices, without dependence on regularity or incoherence conditions. The algorithm projects the kernel onto a set of landmark points sampled by their *rid…
Semantic Embeddings are a popular way to represent knowledge in the field of zero-shot learning. We observe their interpretability and discuss their potential utility in a safety-critical context. Concretely, we propose to use them to add introspection and error detection capabilities to neural network classifiers. Fir…
Regularization improves spectral embedding by focusing on the largest blocks.
Improves visualization of high-dimensional data by correcting misleading artifacts in neighbor embedding methods.
BSAC improves credit scoring models by leveraging autoencoders and addressing imbalanced datasets.
Spectral graph sparsification preserves geometry of GNN embeddings.
The statistical leverage scores of a complex matrix record the degree of alignment between col and the coordinate axes in . These score are used in random sampling algorithms for solving certain numerical linear algebra problems. In this paper we present a max-plus algebr…
CAEL-MIPS learns embeddings to improve MIPS for better OPE in contextual bandits.
Data is said to follow the transform (or analysis) sparsity model if it becomes sparse when acted on by a linear operator called a sparsifying transform. Several algorithms have been designed to learn such a transform directly from data, and data-adaptive sparsifying transforms have demonstrated excellent performance i…
In this paper, we consider a privacy preserving encoding framework for identification applications covering biometrics, physical object security and the Internet of Things (IoT). The proposed framework is based on a sparsifying transform, which consists of a trained linear map, an element-wise nonlinearity, and privacy…
LOCA learns standardized data coordinates from measurements.
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
STRATA generates code adversarial examples efficiently without gradients.
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
In order to mimic the human ability of continual acquisition and transfer of knowledge across various tasks, a learning system needs the capability for continual learning, effectively utilizing the previously acquired skills. As such, the key challenge is to transfer and generalize the knowledge learned from one task t…
DrBO uses Bayesian optimization to learn DAGs more efficiently.
ECS evaluates synthetic CXR images' distributional fidelity.
Improved bounds for sensitivity sampling reducing the sample complexity for structured matrices.
Efficiently approximates statistical leverage scores for faster KRR.
Bayesian approach sparsifies neural networks efficiently.
Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
New algorithms estimate matrix leverage scores using rank revealing and randomization.
A new model improves CT image quality from low-dose scans.
Binary testing for softmax models requires many samples, similar to leverage score models.
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion tha…
Compressing word embeddings is important for deploying NLP models in memory-constrained settings. However, understanding what makes compressed embeddings perform well on downstream tasks is challenging---existing measures of compression quality often fail to distinguish between embeddings that perform well and those th…
Many applications in signal processing benefit from the sparsity of signals in a certain transform domain or dictionary. Synthesis sparsifying dictionaries that are directly adapted to data have been popular in applications such as image denoising, inpainting, and medical image reconstruction. In this work, we focus in…
SALSA efficiently approximates leverage scores for big data, improving ARMA model fitting.
Signal models based on sparsity, low-rank and other properties have been exploited for image reconstruction from limited and corrupted data in medical imaging and other computational imaging applications. In particular, sparsifying transform models have shown promise in various applications, and offer numerous advantag…
Quizlet is the most popular online learning tool in the United States, and is used by over 2/3 of high school students, and 1/2 of college students. With more than 95% of Quizlet users reporting improved grades as a result, the platform has become the de-facto tool used in millions of classrooms. In this paper, we expl…