Federated learning supports exact support recovery with minimal communication.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Support vector regression (SVR) has been widely used to reduce the high computational cost of computer simulation. SVR assumes the input parameters have equal sample sizes, but unequal sample sizes are often encountered in engineering practices. To solve this issue, a new prediction approach based on SVR, namely as hig…
The paper tackles sampling from Gibbs measures with constrained support, providing a sampling guarantee.
Regularized EM algorithm improves GMM clustering in low sample settings.
A simple method for estimating PMF on large supports, preserving structure and suppressing noise.
MAGT generates data efficiently by aligning to manifold structure.
In this paper, we consider asymptotic properties of the support vector machine (SVM) in high-dimension, low-sample-size (HDLSS) settings. We show that the hard-margin linear SVM holds a consistency property in which misclassification rates tend to zero as the dimension goes to infinity under certain severe conditions. …
The study examines how much data is needed for generative and vision-language models to make reliable predictions.
Improved few-shot learning with LSSVM and transductive modules.
LMC improves sampling from complex distributions using quasi-random sequences.
Support vector machine (SVM) is a particularly powerful and flexible supervised learning model that analyzes data for both classification and regression, whose usual algorithm complexity scales polynomially with the dimension of data space and the number of data points. To tackle the big data challenge, a quantum SVM a…
Paper examines how income support affects retirement decisions for low-income individuals.
Proposes a method to compare noisy high-dimensional datasets with low-dimensional manifolds.
New algorithms reduce computational burden for principal support vector machines.
This paper improves diffusion models for low-dimensional data.
We stabilize the Kumaraswamy distribution for efficient sampling and differentiation.
We consider the problem of estimation of a low-rank matrix from a limited number of noisy rank-one projections. In particular, we propose two fast, non-convex \emph{proper} algorithms for matrix recovery and support them with rigorous theoretical analysis. We show that the proposed algorithms enjoy linear convergence a…
Introduces t-CCS for flexible tensor sampling.
Study shows how diffusion models learn on low-dimensional manifolds.
Real world data often exhibit low-dimensional geometric structures, and can be viewed as samples near a low-dimensional manifold. This paper studies nonparametric regression of Hölder functions on low-dimensional manifolds using deep ReLU networks. Suppose training data are sampled from a Hölder function in $\mathc…
We present a machine learning based method for noise classification using a low-power and inexpensive IoT unit. We use Mel-frequency cepstral coefficients for audio feature extraction and supervised classification algorithms (that is, support vector machine and k-nearest neighbors) for noise classification. We evaluate…
We consider the problem of efficiently approximating and encoding high-dimensional data sampled from a probability distribution in , that is nearly supported on a -dimensional set - for example supported on a -dimensional Riemannian manifold. Geometric Multi-Resolution Analysis (GM…
Low-rank matrix completion (LRMC) problems arise in a wide variety of applications. Previous theory mainly provides conditions for completion under missing-at-random samplings. This paper studies deterministic conditions for completion. An incomplete matrix is finitely rank- completable if there are at …
When optimizing against the mean loss over a distribution of predictions in the context of a regression task, then even if there is a distribution of targets the optimal prediction distribution is always a delta function at a single value. Methods of constructing generative models need to overcome this tendency. We con…
We simplify SSL by approximating redundant structural components with low-rank factorization.
Stochastic Gradient Descent (SGD) is a popular optimization method which has been applied to many important machine learning tasks such as Support Vector Machines and Deep Neural Networks. In order to parallelize SGD, minibatch training is often employed. The standard approach is to uniformly sample a minibatch at each…
UAPCA projects uncertain data to low dimensions using GMMs.
We study signal recovery on graphs based on two sampling strategies: random sampling and experimentally designed sampling. We propose a new class of smooth graph signals, called approximately bandlimited, which generalizes the bandlimited class and is similar to the globally smooth class. We then propose two recovery s…
Proposes a novel classification criterion for high-dimensional data with few samples.
Robust STAP with coprime arrays reduces clutter using sparse modeling.
RFSVM uses learned RF similarity for HDLSS classification.
Improved sample complexity for Gaussian process approximations.
New diffusion models learn distributions from samples with improved error bounds.
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…
Generative models learn complex data from low-dimensional manifolds.
A new method classifies color images using quaternion algebra.
New classifiers for HDLSS data classify without tuning, robustly.
Paper explores how Rectified Flow adapts to low-dimensional data.
Indoor localization is a supporting technology for a broadening range of pervasive wireless applications. One promis- ing approach is to locate users with radio frequency fingerprints. However, its wide adoption in real-world systems is challenged by the time- and manpower-consuming site survey process, which builds a …
No GANs can learn disconnected manifolds precisely.
The Nyström methods have been popular techniques for scalable kernel based learning. They approximate explicit, low-dimensional feature mappings for kernel functions from the pairwise comparisons with the training data. However, Nyström methods are generally applied without the supervision provided by the training labe…
Low-precision training reduces computational cost and produces efficient models. Recent research in developing new low-precision training algorithms often relies on simulation to empirically evaluate the statistical effects of quantization while avoiding the substantial overhead of building specific hardware. To suppor…
We study the problem of learning latent variables in Gaussian graphical models. Existing methods for this problem assume that the precision matrix of the observed variables is the superposition of a sparse and a low-rank component. In this paper, we focus on the estimation of the low-rank component, which encodes the e…
Deep neural networks can estimate Q-values efficiently on low-dimensional state-action spaces.
This paper studies binary classification problem associated with a family of loss functions called large-margin unified machines (LUM), which offers a natural bridge between distribution-based likelihood approaches and margin-based approaches. It also can overcome the so-called data piling issue of support vector machi…
Compressing data helps learn Mahalanobis metrics effectively.
We study the problem of recovering the structure underlying large Gaussian graphical models or, more generally, partial correlation graphs. In high-dimensional problems it is often too costly to store the entire sample covariance matrix. We propose a new input model in which one can query single entries of the covarian…
The paper proposes a novel tensor-based method for non-parametric density estimation.