It was proved in 1998 by Ben-David and Litman that a concept space has a sample compression scheme of size d if and only if every finite subspace has a sample compression scheme of size d. In the compactness theorem, measurability of the hypotheses of the created sample compression scheme is not guaranteed; at the same…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Enhanced Sampling Scheme improves masked generative modeling.
We establish a tight characterization of the worst-case rates for the excess risk of agnostic learning with sample compression schemes and for uniform convergence for agnostic sample compression schemes. In particular, we find that the optimal rates of convergence for size- agnostic sample compression schemes are of…
Reduces multiclass and regression compression schemes to binary ones.
Learnable multiclass hypothesis classes don't always have a sample compression scheme of fixed size.
Two new coding schemes improve the efficient communication of noisy data.
In this paper we propose a generalized numerical scheme for backward stochastic differential equations(BSDEs). The scheme is based on approximation of derivatives via Lagrange interpolation. By changing the distribution of sample points used for interpolation, one can get various numerical schemes with different stabil…
This paper studies the sample complexity of searching over multiple populations. We consider a large number of populations, each corresponding to either distribution P0 or P1. The goal of the search problem studied here is to find one population corresponding to distribution P1 with as few samples as possible. The main…
Optimized sampling scheme for compressed sensing combining randomness and determinism.
We obtain the first positive results for bounded sample compression in the agnostic regression setting with the loss, where . We construct a generic approximate sample compression scheme for real-valued function classes exhibiting exponential size in the fat-shattering dimension but independen…
Boosts change-point detection power with optimal sub-sampling.
The paper presents two schemes for sampling matrices from specific distributions on a manifold.
Extends JKO scheme for iterative algorithms with unknown parameters.
We study the performance of the adaptive construction scheme for a Bayesian inference on the Quadratic GARCH model which introduces the asymmetry in time series dynamics. In the adaptive construction scheme a proposal density in the Metropolis-Hastings algorithm is constructed adaptively by changing the parameters of t…
Paper connects sampling and labeling biases in large-output spaces.
Efficiently simulates SABR model with novel sampling methods.
Several approximate policy iteration schemes without value functions, which focus on policy representation using classifiers and address policy learning as a supervised learning problem, have been proposed recently. Finding good policies with such methods requires not only an appropriate classifier, but also reliable e…
New sampling scheme improves ML accuracy in physics simulations.
We perform Markov chain Monte Carlo simulations for a Bayesian inference of the GJR-GARCH model which is one of asymmetric GARCH models. The adaptive construction scheme is used for the construction of the proposal density in the Metropolis-Hastings algorithm and the parameters of the proposal density are determined ad…
New algorithms improve sampling from constrained distributions.
Suppose that one particular block in a stochastic block model is of interest, but block labels are only observed for a few of the vertices in the network. Utilizing a graph realized from the model and the observed block labels, the vertex nomination task is to order the vertices with unobserved block labels into a rank…
We propose a Monte Carlo algorithm to sample from high dimensional probability distributions that combines Markov chain Monte Carlo and importance sampling. We provide a careful theoretical analysis, including guarantees on robustness to high dimensionality, explicit comparison with standard Markov chain Monte Carlo me…
A Bayesian estimation of a GARCH model is performed for US Dollar/Japanese Yen exchange rate by the Metropolis-Hastings algorithm with a proposal density given by the adaptive construction scheme. In the adaptive construction scheme the proposal density is assumed to take a form of a multivariate Student's t-distributi…
New schemes improve error estimates for sampling from non-log-concave distributions.
Improved sampling for network community detection.
Quantum Proof-of-Work uses boson sampling to secure blockchain consensus.
New method approximates CVaR with less data for heavy-tailed risks.
New method for LLMs to learn reasoning by optimizing latent variables.
Shrunk sample covariance matrix is a factor model of a special form combining some (typically, style) risk factor(s) and principal components with a (block-)diagonal factor covariance matrix. As such, shrinkage, which essentially inherits out-of-sample instabilities of the sample covariance matrix, is not an alternativ…
The paper provides theoretical guarantees for optimized sampling in compressed sensing, showing error vanishes with more measurements.
New algorithms sample from log concave distributions without gradient Lipschitz continuity.
The paper explores when to prioritize easy or hard samples in learning tasks.
This work deals with the simulation of Wishart processes and affine diffusions on positive semidefinite matrices. To do so, we focus on the splitting of the infinitesimal generator, in order to use composition techniques as Ninomiya and Victoir or Alfonsi. Doing so, we have found a remarkable splitting for Wishart proc…
Improved Gumbel watermark detection method.
Recurrent neural networks are nowadays successfully used in an abundance of applications, going from text, speech and image processing to recommender systems. Backpropagation through time is the algorithm that is commonly used to train these networks on specific tasks. Many deep learning frameworks have their own imple…
Meta two-sample testing uses auxiliary data to quickly find powerful tests from limited samples.
A new method simplifies sampling from complex distributions without using diffusions.
We investigate the class of -stable Poisson-Kingman random probability measures (RPMs) in the context of Bayesian nonparametric mixture modeling. This is a large class of discrete RPMs which encompasses most of the the popular discrete RPMs used in Bayesian nonparametrics, such as the Dirichlet process, Pitman-Yor p…
The Epps effect varies under different sampling schemes, affecting correlation emergence rates.
Matrix multiplication is a fundamental building block for large scale computations arising in various applications, including machine learning. There has been significant recent interest in using coding to speed up distributed matrix multiplication, that are robust to stragglers (i.e., machines that may perform slower …
Adaptive importance sampling for estimating point process statistics.
The paper optimizes RV estimation by efficient sampling in time-changed diffusion models.
A new term weighting scheme TF-IDFC-RF outperforms others in sentiment analysis.
Optimizes network sampling for efficient community detection.
Novel data acquisition schemes have been an emerging need for scanning microscopy based imaging techniques to reduce the time in data acquisition and to minimize probing radiation in sample exposure. Varies sparse sampling schemes have been studied and are ideally suited for such applications where the images can be re…
Improved statistical computation through efficient matrix sampling.
We consider the problem of sampling from posterior distributions for Bayesian models where some parameters are restricted to be orthogonal matrices. Such matrices are sometimes used in neural networks models for reasons of regularization and stabilization of training procedures, and also can parameterize matrices of bo…
Estimates expected information gain using density approximations and dimension reduction.