A new bootstrapping method reduces key sizes and runtime in FHE.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The main goal of this work is equipping convex and nonconvex problems with Barzilai-Borwein (BB) step size. With the adaptivity of BB step sizes granted, they can fail when the objective function is not strongly convex. To overcome this challenge, the key idea here is to bridge (non)convex problems and strongly convex …
New method ensures consistent inference across different tensor parallel sizes for large language models.
Adapts DR objectives for both sample and feature size reduction.
U-statistics improve gradient estimation in importance-weighted variational inference.
This work proposes a collaborative multi-head attention layer to reduce model size without sacrificing accuracy.
This work studies the implicit bias of mini-batch SGD in classification.
SVRN accelerates Newton methods by reducing variance and improving performance.
Improved variance reduction for Riemannian non-convex optimization with adaptive batch size.
Paper reduces recommender system model size by 90%.
Structural and topological information play a key role in modeling flow and transport through fractured rock in the subsurface. Discrete fracture network (DFN) computational suites such as dfnWorks are designed to simulate flow and transport in such porous media. Flow and transport calculations reveal that a small back…
New method for spatiotemporal data regression using Gaussian processes.
We develop a coreset for robust geometric median, reducing size dependency on outliers.
Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite, number of loss functions. In this paper, we propose a novel Riemannian extension of the Euclidean stochastic variance reduced gradient algorithm (R-SVRG) to a compact manifold search space. To this e…
In this paper we propose the macroblock scaling (MBS) algorithm, which can be applied to various CNN architectures to reduce their model size. MBS adaptively reduces each CNN macroblock depending on its information redundancy measured by our proposed effective flops. Empirical studies conducted with ImageNet and CIFAR-…
In recent years, stochastic variance reduction algorithms have attracted considerable attention for minimizing the average of a large but finite number of loss functions. This paper proposes a novel Riemannian extension of the Euclidean stochastic variance reduced gradient (R-SVRG) algorithm to a manifold search space.…
Scalable GPLVM reduces complexity in scRNA-seq data, accounting for technical and biological confounders.
To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will be able to build accurate data-driven inferences. Because sample sizes are typically orders of magnitude smaller than the dimensionality of t…
A new method uses Gram matrix for efficient multivariate functional principal components.
A new method improves active learning for large batch sizes.
SPREV simplifies visualization of complex labeled datasets.
We present a theory of reduction for Courant algebroids as well as Dirac structures, generalized complex, and generalized Kähler structures which interpolates between holomorphic reduction of complex manifolds and symplectic reduction. The enhanced symmetry group of a Courant algebroid leads us to define \emph{extended…
Value selection reduces model size while maintaining accuracy.
In this paper we study a family of variance reduction methods with randomized batch size---at each step, the algorithm first randomly chooses the batch size and then selects a batch of samples to conduct a variance-reduced stochastic update. We give the linear convergence rate for this framework for composite functions…
Similar to convolution neural networks, recurrent neural networks (RNNs) typically suffer from over-parameterization. Quantizing bit-widths of weights and activations results in runtime efficiency on hardware, yet it often comes at the cost of reduced accuracy. This paper proposes a quantization approach that increases…
New methods improve subgroup analysis in trials with limited data.
Study finds that only a fraction of data is needed for accurate patient-level prediction models.
In this paper we aim to formally explain the phenomenon of fast convergence of SGD observed in modern machine learning. The key observation is that most modern learning architectures are over-parametrized and are trained to interpolate the data by driving the empirical loss (classification and regression) close to zero…
The variance reduction class of algorithms including the representative ones, SVRG and SARAH, have well documented merits for empirical risk minimization problems. However, they require grid search to tune parameters (step size and the number of iterations per inner loop) for optimal performance. This work introduces `…
This paper develops a general framework for analyzing asymptotics of -statistics. Previous literature on limiting distribution mainly focuses on the cases when with fixed kernel size . Under some regularity conditions, we demonstrate asymptotic normality when grows with by utilizing existin…
Paper proposes a method to estimate variance reduction in DNN training using importance sampling.
Develops a new method for online conformal prediction without manual tuning.
Stochastic gradient Markov Chain Monte Carlo (SG-MCMC) has been developed as a flexible family of scalable Bayesian sampling algorithms. However, there has been little theoretical analysis of the impact of minibatch size to the algorithm's convergence rate. In this paper, we prove that under a limited computational bud…
The ability to accurately predict the fit of fashion items and recommend the correct size is key to reducing merchandise returns in e-commerce. A critical prerequisite of fit prediction is size normalization, the mapping of product sizes across brands to a common space in which sizes can be compared. At present, size n…
New theory shows how multi-head attention reduces variance and decorrelates outputs.
We demonstrate that almost all non-parametric dimensionality reduction methods can be expressed by a simple procedure: regularized loss minimization plus singular value truncation. By distinguishing the role of the loss and regularizer in such a process, we recover a factored perspective that reveals some gaps in the c…
Scalability of statistical estimators is of increasing importance in modern applications and dimension reduction is often used to extract relevant information from data. A variety of popular dimension reduction approaches can be framed as symmetric generalized eigendecomposition problems. In this paper we outline how t…
Modeling data as being sampled from a union of independent subspaces has been widely applied to a number of real world applications. However, dimensionality reduction approaches that theoretically preserve this independence assumption have not been well studied. Our key contribution is to show that projection vect…
Pruning improves model generalization in over-parameterized models, contradicting traditional theories.
This paper proposes a novel uncertainty quantification framework for computationally demanding systems characterized by a large vector of non-Gaussian uncertainties. It combines state-of-the-art techniques in advanced Monte Carlo sampling with Bayesian formulations. The key departure from existing works is the use of i…
PCA simplifies multivariate extreme data analysis.
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
FSGD uses latent factors to scale SGD for high-dimensional learning.
This paper explores how effective sample size, dimensionality, and model performance are related in covariate shift adaptation.
NCP improves deep classifier uncertainty quantification efficiency.
The Long-Short-Term-Memory Recurrent Neural Networks (LSTM RNNs) are a popular class of machine learning models for analyzing sequential data. Their training on modern GPUs, however, is limited by the GPU memory capacity. Our profiling results of the LSTM RNN-based Neural Machine Translation (NMT) model reveal that fea…
Study identifies key metrics for small and large tick assets in LOBs.
The salient properties of large empirical covariance and correlation matrices are studied for three datasets of size 54, 55 and 330. The covariance is defined as a simple cross product of the returns, with weights that decay logarithmically slowly. The key general properties of the covariance matrices are the following…