Count-sketches reduce memory usage for deep learning models without sacrificing performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FetchSGD reduces communication in federated learning with sketching.
A new sketching method reduces tensor memory usage and enables efficient tensor operations.
Extracting information from electronic health records (EHR) is a challenging task since it requires prior knowledge of the reports and some natural language processing algorithm (NLP). With the growing number of EHR implementations, such knowledge is increasingly challenging to obtain in an efficient manner. We address…
DiffSketch combines privacy and communication efficiency in distributed learning.
Feature selection is an important challenge in machine learning. It plays a crucial role in the explainability of machine-driven decisions that are rapidly permeating throughout modern society. Unfortunately, the explosion in the size and dimensionality of real-world datasets poses a severe challenge to standard featur…
GCWSNet improves neural network training speed and accuracy with power transformation.
FedSKETCH and FedSKETCHGATE improve privacy and efficiency in federated learning.
We introduce a new sub-linear space sketch---the Weight-Median Sketch---for learning compressed linear classifiers over data streams while supporting the efficient recovery of large-magnitude weights in the model. This enables memory-limited execution of several statistical analyses over streams, including online featu…
Tensor CANDECOMP/PARAFAC (CP) decomposition has wide applications in statistical learning of latent variable models and in data mining. In this paper, we propose fast and randomized tensor CP decomposition algorithms based on sketching. We build on the idea of count sketches, but introduce many novel ideas which are un…
Structured high-cardinality data arises in many domains, and poses a major challenge for both modeling and inference. Graphical models are a popular approach to modeling structured data but they are unsuitable for high-cardinality variables. The count-min (CM) sketch is a popular approach to estimating probabilities in…
The paper develops methods to estimate frequencies in large discrete data sets with improved coverage and robustness.
The paper introduces DP algorithms using random projections and sign random projections for improved privacy in machine learning.
Improved CountSketch method reduces variance for estimating vector coordinates.
OPORP combines permutation and random projection for efficient data vector compression.