In this paper, we explore a general Aggregated Gradient Langevin Dynamics framework (AGLD) for the Markov Chain Monte Carlo (MCMC) sampling. We investigate the nonasymptotic convergence of AGLD with a unified analysis for different data accessing (e.g. random access, cyclic access and random reshuffle) and snapshot upd…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Developing efficient and scalable algorithms for Latent Dirichlet Allocation (LDA) is of wide interest for many applications. Previous work has developed an O(1) Metropolis-Hastings sampling method for each token. However, the performance is far from being optimal due to random accesses to the parameter matrices and fr…
New methods protect privacy while providing accurate prediction sets.
Paper speeds up IoT device detection and data decoding.
We investigate the computational complexity of several basic linear algebra primitives, including largest eigenvector computation and linear regression, in the computational model that allows access to the data via a matrix-vector product oracle. We show that for polynomial accuracy, calls to the oracle are nece…
We provide tight upper and lower bounds on the complexity of minimizing the average of convex functions using gradient and prox oracles of the component functions. We show a significant gap between the complexity of deterministic vs randomized optimization. For smooth functions, we show that accelerated gradient de…
Federated Learning tackles limited user participation with a new risk-aware approach.
Query access significantly speeds up learning Multi-Index Models under Gaussian distribution.
New approach uses random matrix theory to understand tensor estimation performance.
Assume we are given a set of items from a general metric space, but we neither have access to the representation of the data nor to the distances between data points. Instead, suppose that we can actively choose a triplet of items (A,B,C) and ask an oracle whether item A is closer to item B or to item C. In this paper,…
Quantum algorithms for multi-armed bandits are explored with limited reward access.
To make efficient use of limited spectral resources, we in this work propose a deep actor-critic reinforcement learning based framework for dynamic multichannel access. We consider both a single-user case and a scenario in which multiple users attempt to access channels simultaneously. We employ the proposed framework …
Random forests is a common non-parametric regression technique which performs well for mixed-type data and irrelevant covariates, while being robust to monotonic variable transformations. Existing random forest implementations target regression or classification. We introduce the RFCDE package for fitting random forest…
In this paper we study a family of variance reduction methods with randomized batch size---at each step, the algorithm first randomly chooses the batch size and then selects a batch of samples to conduct a variance-reduced stochastic update. We give the linear convergence rate for this framework for composite functions…
Study on tensor signal estimation from incomplete data.
A study on portfolio delegation with random default times, addressing complex uncertainties.
New insights and algorithms improve prediction models with time-series privileged information.
We present a new random sampling strategy for k-bandlimited signals defined on graphs, based on determinantal point processes (DPP). For small graphs, ie, in cases where the spectrum of the graph is accessible, we exhibit a DPP sampling scheme that enables perfect recovery of bandlimited signals. For large graphs, ie, …
New methods for private statistical inference under local differential privacy.
Criterion extends identifiability for continuous mixtures of kernels.
In this paper, we investigate cost-aware joint learning and optimization for multi-channel opportunistic spectrum access in a cognitive radio system. We investigate a discrete time model where the time axis is partitioned into frames. Each frame consists of a sensing phase, followed by a transmission phase. During the …
This paper evaluates LLMs on large graph property estimation tasks.
Improved sampling efficiency with Random Reshuffling for Langevin dynamics.
Random forests have become an important tool for improving accuracy in regression and classification problems since their inception by Leo Breiman in 2001. In this paper, we revisit a historically important random forest model originally proposed by Breiman in 2004 and later studied by Gérard Biau in 2012, where a feat…
FibQuant improves KV-cache compression for long-context inference.
New method shows supervised learning can mimic unsupervised learning effectively.
Coded Federated Learning speeds up training in edge computing networks.
Paper studies federated nonparametric testing with privacy constraints, achieving optimal rates and adaptive testing.
New method attacks GNNs with limited node access, increasing misclassification rate.
Generative model uses random weighted support points for interpretable data sampling.
Performing signal processing tasks on compressive measurements of data has received great attention in recent years. In this paper, we extend previous work on compressive dictionary learning by showing that more general random projections may be used, including sparse ones. More precisely, we examine compressive K-mean…
Paper proposes a method to generate adversarial perturbations for black-box attacks without accessing inner states.
For random matrix models, the parameter estimation based on the traditional likelihood functions is not straightforward in particular when we have only one sample matrix. We introduce a new parameter optimization method for random matrix models which works even in such a case. The method is based on the spectral distri…
Develops a new deep learning framework for privacy-preserving text representations.
A new method samples from a target density without initial samples using Monte Carlo estimation of the score.
Adversarial examples pose a threat to deep neural network models in a variety of scenarios, from settings where the adversary has complete knowledge of the model and to the opposite "black box" setting. Black box attacks are particularly threatening as the adversary only needs access to the input and output of the mode…
Subspace clustering refers to the problem of clustering unlabeled high-dimensional data points into a union of low-dimensional linear subspaces, assumed unknown. In practice one may have access to dimensionality-reduced observations of the data only, resulting, e.g., from "undersampling" due to complexity and speed con…
New classifier robust to adversarial perturbations from high-accuracy models.
The paper develops efficient algorithms for sampling from random spanning trees and determinantal point processes.
Proposes MCLLO for assessing and recalibrating multiclass probability predictions.
Domain randomization (DR) is a successful technique for learning robust policies for robot systems, when the dynamics of the target robot system are unknown. The success of policies trained with domain randomization however, is highly dependent on the correct selection of the randomization distribution. The majority of…
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number…
We give algorithms for estimating the expectation of a given real-valued function on a sample drawn randomly from some unknown distribution over domain , namely . Our algorithms work in two well-studied models of restricted access to data samples. The first o…
Algorithm solves job acceptance problem with random arrivals and values.
Classical clients can verify quantum learning tasks efficiently.
Motivated by a sampling problem basic to computational statistical inference, we develop a nearly optimal algorithm for a fundamental problem in spectral graph theory and numerical analysis. Given an SDDM matrix , and a constant , our algorithm gives efficient access to a…
Recent progress has shown that few-shot learning can be improved with access to unlabelled data, known as semi-supervised few-shot learning(SS-FSL). We introduce an SS-FSL approach, dubbed as Prototypical Random Walk Networks(PRWN), built on top of Prototypical Networks (PN). We develop a random walk semi-supervised lo…
This paper compresses large datasets for efficient machine learning.