The paper cleans label noise in supervised classification using Bernoulli sampling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Introduces t-CCS for flexible tensor sampling.
A new sampling method balances multi-label datasets by preserving category frequency order.
Improved training of GRBMs for image generation.
A new method for efficient nonlinear process monitoring using random Bernoulli features.
The question of the optimality of Thompson Sampling for solving the stochastic multi-armed bandit problem had been open since 1933. In this paper we answer it positively for the case of Bernoulli rewards by providing the first finite-time analysis that matches the asymptotic rate given in the Lai and Robbins lower boun…
New acquisition functions improve Bernoulli LSE.
This paper tackles open problem of tight bounds for KBs with Bernoulli rewards.
Spectral method speeds fitting of binary time series models.
Variational autoencoders (VAE) have quickly become a central tool in machine learning, applicable to a broad range of data types and latent variable models. By far the most common first step, taken by seminal papers and by core software libraries alike, is to model MNIST data using a deep network parameterizing a Berno…
This work extends score-based methods to binary data on the Boolean hypercube.
We introduce a novel multivariate random process producing Bernoulli outputs per dimension, that can possibly formalize binary interactions in various graphical structures and can be used to model opinion dynamics, epidemics, financial and biological time series data, etc. We call this a Bernoulli Autoregressive Proces…
A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity analysis in social networks. In this paper, we analyze the clusterability of BMMs from…
A new method trains discrete EBMs without sampling.
Develops Thompson Sampling algorithms for mean-variance bandits.
This paper addresses the mapping problem. Using a conjugate prior form, we derive the exact theoretical batch multi-object posterior density of the map given a set of measurements. The landmarks in the map are modeled as extended objects, and the measurements are described as a Poisson process, conditioned on the map. …
Characterizes symmetric Bernoulli distributions with minimal convex sums.
In this paper, we are concerned with obtaining distribution-free concentration inequalities for mixture of independent Bernoulli variables that incorporate a notion of variance. Missing mass is the total probability mass associated to the outcomes that have not been seen in a given sample which is an important quantity…
Study analyzes symmetric two-armed Bernoulli bandit problem with zero mean gap.
Upper bound on expected supremum of Bernoulli process.
Exact simulation of correlated binary outcomes using PMF constraints and linear programming.
KL-MS improves regret bounds for multi-armed bandits with bounded rewards.
Improved BAI under DP reduces gap to constant.
In this paper, we consider the multivariate Bernoulli distribution as a model to estimate the structure of graphs with binary nodes. This distribution is discussed in the framework of the exponential family, and its statistical properties regarding independence of the nodes are demonstrated. Importantly the model can e…
Finite index solutions to Bernoulli problem are always axially symmetric.
New Gibbs sampling reduces GLMB filtering complexity to linear time.
Proves a principle for one-phase Bernoulli problem minimizers.
The paper develops sampling methods for ocean phenomena based on temperature and salinity measurements.
A very simple event frequency approximation algorithm that is sensitive to event timeliness is suggested. The algorithm iteratively updates categorical click-distribution, producing (path of) a random walk on a standard -dimensional simplex. Under certain conditions, this random walk is self-similar and corresponds …
The key idea of Bayesian optimization is replacing an expensive target function with a cheap surrogate model. By selection of an acquisition function for Bayesian optimization, we trade off between exploration and exploitation. The acquisition function typically depends on the mean and the variance of the surrogate mod…
No algorithm outperforms uniform sampling in A/B testing.
Adaptive learning method identifies and corrects corrupted data.
Data processing inequalities link Fisher information to local differential privacy constraints.
Bayesian autoencoders improve OOD detection by addressing Bernoulli likelihood issues.
This paper proposed a new regression model called -regularized outlier isolation and regression (LOIRE) and a fast algorithm based on block coordinate descent to solve this model. Besides, assuming outliers are gross errors following a Bernoulli process, this paper also presented a Bernoulli estimate model which, …
We solve Euler equations on graph manifolds, classifying steady flows with Morse-Bott Bernoulli functions.
Thompson sampling, a Bayesian method for balancing exploration and exploitation in bandit problems, has theoretical guarantees and exhibits strong empirical performance in many domains. Traditional Thompson sampling, however, assumes perfect compliance, where an agent's chosen action is treated as the implemented actio…
The paper extends consistency results for sequential design strategies to vector-valued Gaussian processes.
Motivated by the ever-increasing demands for limited communication bandwidth and low-power consumption, we propose a new methodology, named joint Variational Autoencoders with Bernoulli mixture models (VAB), for performing clustering in the compressed data domain. The idea is to reduce the data dimension by Variational…
Berry et al. (1997) initiated the development of the infinite arms bandit problem. They derived a regret lower bound of all allocation strategies for Bernoulli rewards with uniform priors, and proposed strategies based on success runs. Bonald and Proutière (2013) proposed a two-target algorithm that achieves the regret…
Fictitious play is a simple and widely studied adaptive heuristic for playing repeated games. It is well known that fictitious play fails to be Hannan consistent. Several variants of fictitious play including regret matching, generalized regret matching and smooth fictitious play, are known to be Hannan consistent. In …
Dasgupta and Shulman showed that a two-round variant of the EM algorithm can learn mixture of Gaussian distributions with near optimal precision with high probability if the Gaussian distributions are well separated and if the dimension is sufficiently high. In this paper, we generalize their theory to learning mixture…
Optimization-based pruning eliminates backpropagation for large language models.
New TS algorithms improve performance in non-stationary multi-armed bandit problems.
We consider the problem of estimating the parameters of a multivariate Bernoulli process with auto-regressive feedback in the high-dimensional setting where the number of samples available is much less than the number of parameters. This problem arises in learning interconnections of networks of dynamical systems with …
We study covariance matrix estimation for the case of partially observed random vectors, where different samples contain different subsets of vector coordinates. Each observation is the product of the variable of interest with a Bernoulli random variable. We analyze an unbiased covariance estimator under this mod…
We study the fundamental problem of learning an unknown, smooth probability function via pointwise Bernoulli tests. We provide a scalable algorithm for efficiently solving this problem with rigorous guarantees. In particular, we prove the convergence rate of our posterior update rule to the true probability function in…
A new method uses Mean Field Games to optimize mixture models of Bernoulli and categorical distributions.