One-pass algorithm finds small subset for subspace approximation with additive error.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian online learning algorithm for one-pass data, achieving frequentist validity and uncertainty quantification.
One-pass private sketch supports various machine learning tasks.
New algorithm achieves nearly optimal regret with one-pass updates for GLB problems.
We study distribution testing with communication and memory constraints in the following computational models: (1) The {\em one-pass streaming model} where the goal is to minimize the sample complexity of the protocol subject to a memory constraint, and (2) A {\em distributed model} where the data samples reside at mul…
One-pass optimisation for high-dimensional hyperparameters.
ORFit trains models on streaming data with one pass, minimizing memory and computational costs.
In this paper, we propose a one-pass algorithm on MapReduce for penalized linear regression \[f_λ(α, β) = \|Y - α\mathbf{1} - Xβ\|_2^2 + p_λ(β)\] where is the intercept which can be omitted depending on application; is the coefficients and is the penalized function with penalizing parameter . $f_λ(α, β…
Full-batch GD outperforms one-pass SGD in learning a single-index model with quadratic activation.
We present a one-pass sparsified Gaussian mixture model (SGMM). Given data points in dimensions, , the model fits Gaussian distributions to and (softly) classifies each point to these clusters. After paying an up-front cost of to precondition the data, we subsample entries…
New method for RLHF reduces costs by integrating new data in one pass.
We develop an efficient alternating framework for learning a generalized version of Factorization Machine (gFM) on steaming data with provable guarantees. When the instances are sampled from dimensional random Gaussian vectors and the target second order coefficient matrix in gFM is of rank , our algorithm conve…
A new feature selection method using attention for neural networks.
One-pass SGD dynamics in overparameterized quadratic networks show slow escape from poor solutions.
One-pass SGD converges in overparametrized neural networks with random data.
We present a streaming model for large-scale classification (in the context of -SVM) by leveraging connections between learning and computational geometry. The streaming model imposes the constraint that only a single pass over the data is allowed. The -SVM is known to have an equivalent formulation in …
We consider the -means clustering problem in the dynamic streaming setting, where points from a discrete Euclidean space can be dynamically inserted to or deleted from the dataset. For this problem, we provide a one-pass coreset construction algorithm using space $\tilde{O}(k\cdot \mathrm{pol…
Stochastic Gradient Descent can overfit after just a few passes, contrary to initial expectations.
Mean field inference in probabilistic models is generally a highly nonconvex problem. Existing optimization methods, e.g., coordinate ascent algorithms, can only generate local optima. In this work we propose provable mean filed methods for probabilistic log-submodular models and its posterior agreement (PA) with stron…
New algorithm reduces heavy-tailed linear bandits' computational cost.
Consider a sequence of closed, orientable surfaces of fixed genus in a Riemannian manifold with uniform upper bounds on mean curvature and area. We show that on passing to a subsequence and choosing appropriate parametrisations, the inclusion maps converge in to a map from a surface of genus to . W…
A new algorithm reduces memory usage for long token attention in streaming applications.
In many large-scale machine learning applications, data are accumulated with time, and thus, an appropriate model should be able to update in an online paradigm. Moreover, as the whole data volume is unknown when constructing the model, it is desired to scan each data item only once with a storage independent with the …
A new data-oblivious sketch for logistic regression reduces data size while maintaining approximation accuracy.
Scalable methods for maximizing regularized submodular functions with improved memory and communication complexity.
We present a novel neural network algorithm, the Tensor Switching (TS) network, which generalizes the Rectified Linear Unit (ReLU) nonlinearity to tensor-valued hidden units. The TS network copies its entire input vector to different locations in an expanded representation, with the location determined by its hidden un…
This paper addresses how well we can recover a data matrix when only given a few of its elements. We present a randomized algorithm that element-wise sparsifies the data, retaining only a few its elements. Our new algorithm independently samples the data using sampling probabilities that depend on both the squares ($\e…
Real data are often with multiple modalities or from multiple heterogeneous sources, thus forming so-called multi-view data, which receives more and more attentions in machine learning. Multi-view clustering (MVC) becomes its important paradigm. In real-world applications, some views often suffer from instances missing…
Improved COD algorithm reduces streaming AMM errors and uses less space.
We study the classical problem of maximizing a monotone submodular function subject to a cardinality constraint k, with two additional twists: (i) elements arrive in a streaming fashion, and (ii) m items from the algorithm's memory are removed after the stream is finished. We develop a robust submodular algorithm STAR-…
Two new approaches for point prediction in streaming data, showing consistency and performance.
The equivariant Gromov--Hausdorff convergence of metric spaces is studied. Where all isometry groups under consideration are compact Lie, it is shown that an upper bound on the dimension of the group guarantees that the convergence is by Lie homomorphisms. Additional lower bounds on curvature and volume strengthen this…
Given a compact Riemann surface and a complex reductive Lie group equipped with real structures, we define antiholomorphic involutions on the moduli space of -Higgs bundles over . We investigate how the various components of the fixed point locus match up, as one passes from to its Langlands dual $^LG…
Polyak-Ruppert CLT for SA-Adam with momentum and non-convergent adaptive preconditioning
The probability Jaccard similarity was recently proposed as a natural generalization of the Jaccard similarity to measure the proximity of sets whose elements are associated with relative frequencies or probabilities. In combination with a hash algorithm that maps those weighted sets to compact signatures which allow f…
OPAA estimates probability densities using functional analysis.
The paper analyzes how repeating epochs affects data scaling in linear regression.
In this paper, we consider the streaming memory-limited matrix completion problem when the observed entries are noisy versions of a small random fraction of the original entries. We are interested in scenarios where the matrix size is very large so the matrix is very hard to store and manipulate. Here, columns of the o…
The stochastic gradient descent (SGD) algorithm has been widely used in statistical estimation for large-scale data due to its computational and memory efficiency. While most existing works focus on the convergence of the objective function or the error of the obtained solution, we investigate the problem of statistica…
Streaming algorithms are generally judged by the quality of their solution, memory footprint, and computational complexity. In this paper, we study the problem of maximizing a monotone submodular function in the streaming setting with a cardinality constraint . We first propose Sieve-Streaming++, which requires just…
Error bound conditions (EBC) are properties that characterize the growth of an objective function when a point is moved away from the optimal set. They have recently received increasing attention in the field of optimization for developing optimization algorithms with fast convergence. However, the studies of EBC in st…
We consider streaming, one-pass principal component analysis (PCA), in the high-dimensional regime, with limited memory. Here, -dimensional samples are presented sequentially, and the goal is to produce the -dimensional subspace that best approximates these points. Standard algorithms require memory; mea…
We describe a general framework -- compressive statistical learning -- for resource-efficient large-scale learning: the training collection is compressed in one pass into a low-dimensional sketch (a vector of random empirical generalized moments) that captures the information relevant to the considered learning task. A…
Kernel-based K-means clustering has gained popularity due to its simplicity and the power of its implicit non-linear representation of the data. A dominant concern is the memory requirement since memory scales as the square of the number of data points. We provide a new analysis of a class of approximate kernel methods…
We study -divergence contraction and its privacy implications.
A fast method combines deep mixtures of sparse GPs for flexible modeling.
New method models matrix time series using tensor CP-decomposition.
We describe many vantage points on the Baire metric and its use in clustering data, or its use in preprocessing and structuring data in order to support search and retrieval operations. In some cases, we proceed directly to clusters and do not directly determine the distances. We show how a hierarchical clustering can …