New algorithm reduces rank constrained optimization problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present a mechanism to compute a sketch (succinct summary) of how a complex modular deep network processes its inputs. The sketch summarizes essential information about the inputs and outputs of the network and can be used to quickly identify key components and summary statistics of the inputs. Furthermore, the sket…
New analysis proves sketching operators' RIP guarantees for mixture models without importance sampling.
DiPS learns to optimize sketching policies for better recommendation quality.
This paper speeds up kernel methods using sparsified Gaussian sketches.
With the scale of data growing every day, reducing the dimensionality (a.k.a. sketching) of high-dimensional data has emerged as a task of paramount importance. Relevant issues to address in this context include the sheer volume of data that may consist of categorical samples, the typically streaming format of acquisit…
Unified methodology for statistical inference in least squares and PCA via randomized sketching.
New algorithm optimizes positions of CountSketch non-zero entries for better data compression.
A fast sketching algorithm solves regularized least squares problems efficiently.
We introduce features for massive data streams. These stream features can be thought of as "ordered moments" and generalize stream sketches from "moments of order one" to "ordered moments of arbitrary order". In analogy to classic moments, they have theoretical guarantees such as universality that are important for lea…
Unified bounds for sketched bilinear forms in machine learning and statistics.
Detecting emergence of a low-rank signal from high-dimensional data is an important problem arising from many applications such as camera surveillance and swarm monitoring using sensors. We consider a procedure based on the largest eigenvalue of the sample covariance matrix over a sliding window to detect the change. T…
This paper compresses large datasets for efficient machine learning.
We sketch a construction of Legendrian Symplectic Field Theory (SFT) for conormal tori of knots and links. Using large duality and Witten's connection between open Gromov-Witten invariants and Chern-Simons gauge theory, we relate the SFT of a link conormal to the colored HOMFLY-PT polynomials of the link. We presen…
New sketches for weighted sampling without replacement improve accuracy and efficiency.
The immense amount of daily generated and communicated data presents unique challenges in their processing. Clustering, the grouping of data without the presence of ground-truth labels, is an important tool for drawing inferences from data. Subspace clustering (SC) is a relatively recent method that is able to successf…
COPT optimizes graph distances via simultaneous optimal transport.
Two new algorithms select matrix rows and columns to preserve distances.
In this paper, we develop a novel procedure for low-rank tensor regression, namely \emph{\underline{I}mportance \underline{S}ketching \underline{L}ow-rank \underline{E}stimation for \underline{T}ensors} (ISLET). The central idea behind ISLET is \emph{importance sketching}, i.e., carefully designed sketches based on bot…
Optimal sketching bounds for sparse linear regression under various loss functions are established.
Sketch-BERT learns vector sketches using BERT-like self-supervised learning.
A new method speeds up community detection in graphs.
New algorithm speeds up polynomial kernel approximations.
There is an especially strong need in modern large-scale data analysis to prioritize samples for manual inspection. For example, the inspection could target important mislabeled samples or key vulnerabilities exploitable by an adversarial attack. In order to solve the "needle in the haystack" problem of which samples t…
New algorithm for efficient prediction intervals in neural networks.
A new method improves convergence in low-rank approximation.
We address the statistical and optimization impacts of the classical sketch and Hessian sketch used to approximately solve the Matrix Ridge Regression (MRR) problem. Prior research has quantified the effects of classical sketch on the strictly simpler least squares regression (LSR) problem. We establish that classical …
Considering the existence of very large amount of available data repositories and reach to the very advanced system of hardware, systems meant for facial identification ave evolved enormously over the past few decades. Sketch recognition is one of the most important areas that have evolved as an integral component adop…
Localized sketching improves matrix multiplication and ridge regression complexity.
We consider the problem of selecting non-zero entries of a matrix in order to produce a sparse sketch of it, , that minimizes . For large matrices, such that (for example, representing observations over attributes) we give sampling distributions that exhibit four importa…
We introduce a new sub-linear space sketch---the Weight-Median Sketch---for learning compressed linear classifiers over data streams while supporting the efficient recovery of large-magnitude weights in the model. This enables memory-limited execution of several statistical analyses over streams, including online featu…
The paper sharpens the analysis of sketch-and-project methods using randomized singular value decomposition.
New method reduces linear regret in high-dimensional bandit problems.
Extracting information from electronic health records (EHR) is a challenging task since it requires prior knowledge of the reports and some natural language processing algorithm (NLP). With the growing number of EHR implementations, such knowledge is increasingly challenging to obtain in an efficient manner. We address…
New method for efficient maximum likelihood estimation of -generalized probit regression.
A new variable importance measure for DRFs detects broader impacts on output distributions.
The Min-Hashing approach to sketching has become an important tool in data analysis, information retrial, and classification. To apply it to real-valued datasets, the ICWS algorithm has become a seminal approach that is widely used, and provides state-of-the-art performance for this problem space. However, ICWS suffers…
New guarantees for asymmetric sketching in compressive learning.
Parameter reduction has been an important topic in deep learning due to the ever-increasing size of deep neural network models and the need to train and run them on resource limited machines. Despite many efforts in this area, there were no rigorous theoretical guarantees on why existing neural net compression methods …
A new method for estimating large-scale linear models with improved precision.
Private sketches protect linear regression data privacy.
Sketching is a randomized dimensionality-reduction method that aims to preserve relevant information in large-scale datasets. Count sketch is a simple popular sketch which uses a randomized hash function to achieve compression. In this paper, we propose a novel extension known as Higher-order Count Sketch (HCS). While …
Paper proposes Nyström sketches for better adaptive compressive learning.
In this note the long standing problem of the definition of a Poisson bracket in the framework of a multisymplectic formulation of classical field theory is solved. The new bracket operation can be applied to forms of arbitary degree. Relevant examples are discussed and important properties are stated with proofs sketc…
Sketching is more fundamental to human cognition than speech. Deep Neural Networks (DNNs) have achieved the state-of-the-art in speech-related tasks but have not made significant development in generating stroke-based sketches a.k.a sketches in vector format. Though there are Variational Auto Encoders (VAEs) for genera…
Density sketches summarize data distributions for accurate sampling and estimation.
Improved sketching for logistic and regression with near-linear dimensions.
We improve prediction risk estimation for large datasets using sketching and ridge regression.