Generation of pseudorandom numbers from different probability distributions has been studied extensively in the Monte Carlo simulation literature. Two standard generation techniques are the acceptance-rejection and inverse transformation methods. An alternative approach to Monte Carlo simulation is the quasi-Monte Carl…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Variational inference using the reparameterization trick has enabled large-scale approximate Bayesian inference in complex probabilistic models, leveraging stochastic optimization to sidestep intractable expectations. The reparameterization trick is applicable when we can simulate a random variable by applying a differ…
New MC simulation methods use classifiers to estimate pdf ratios without explicit pdfs.
Polynomial-time algorithm for near-optimal community detection in graphs.
Score calibration enables automatic speaker recognizers to make cost-effective accept / reject decisions. Traditional calibration requires supervised data, which is an expensive resource. We propose a 2-component GMM for unsupervised calibration and demonstrate good performance relative to a supervised baseline on NIST…
ABC method uses machine learning for likelihood-free inference.
We consider the problem of sampling from a strongly log-concave density in , and prove a non-asymptotic upper bound on the mixing time of the Metropolis-adjusted Langevin algorithm (MALA). The method draws samples by simulating a Markov chain obtained from the discretization of an appropriate Langevin dif…
New method improves sampling from score-based models by correcting bias.
In this work, we present an application of Locally Interpretable Machine-Agnostic Explanations to 2-D chemical structures. Using this framework we are able to provide a structural interpretation for an existing black-box model for classifying biologically produced fuel compounds with regard to Research Octane Number. T…
DART optimizes subset selection in non-linear bandit problems.
Markov chain (MC) algorithms are ubiquitous in machine learning and statistics and many other disciplines. Typically, these algorithms can be formulated as acceptance rejection methods. In this work we present a novel estimator applicable to these methods, dubbed Markov chain importance sampling (MCIS), which efficient…
Recent developments in differentially private (DP) machine learning and DP Bayesian learning have enabled learning under strong privacy guarantees for the training data subjects. In this paper, we further extend the applicability of DP Bayesian learning by presenting the first general DP Markov chain Monte Carlo (MCMC)…
We propose Learned Accept/Reject Sampling (LARS), a method for constructing richer priors using rejection sampling with a learned acceptance function. This work is motivated by recent analyses of the VAE objective, which pointed out that commonly used simple priors can lead to underfitting. As the distribution induced …
Develops a new bivariate process for energy markets with improved simulation methods.
Study reveals bias in machine learning conference reviews.
New algorithm speeds up SVAR inference for large datasets.
NUTS mixing time scales as d^(1/4) for Gaussian distributions.
The paper connects ABC to GBI, suggesting ABC as a robustification strategy.
Energy-based models (EBMs) are powerful probabilistic models, but suffer from intractable sampling and density evaluation due to the partition function. As a result, inference in EBMs relies on approximate sampling algorithms, leading to a mismatch between the model and inference. Motivated by this, we consider the sam…
A new method estimates expectations from subtractive mixture models without sampling.
The transition probability of a Cox-Ingersoll-Ross process can be represented by a non-central chi-square density. First we prove a new representation for the central chi-square density based on sums of powers of generalized Gaussian random variables. Second we prove Marsaglia's polar method extends to this distributio…
Accelerating Speculative Diffusions via Block Verification
Enhanced SMC uses gradients from CRN-PF in Langevin proposals for improved state and parameter estimation.
Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally expensive. Both the calculation of the acceptance probability and the creation of informed proposals usually require an iteration through the whole data set. The recently proposed stochastic gradient Langevin dynamics (SG…
Adaptive source selection for positive transfer in linear models improves target dataset performance.
Proposes a new undersampling method for imbalanced data classification.
New algorithm improves sampling from constrained spaces.
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
Framework for Bayesian inference using GP emulated MH sampler for noisy likelihoods.
Interacting particle methods are increasingly used to sample from complex and high-dimensional distributions. These stochastic particle integration techniques can be interpreted as an universal acceptance-rejection sequential particle sampler equipped with adaptive and interacting recycling mechanisms. Practically, the…
New MCMC methods use auxiliary variables to sample from intractable distributions.
Detects corruption in agentic models during execution.
Piecewise normalizing flows improve multi-modal distribution modeling.
Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally infeasible. The recently proposed stochastic gradient Langevin dynamics (SGLD) method circumvents this problem in three ways: it generates proposed moves using only a subset of the data, it skips the Metropolis-Hastings a…
For classification problems with significant class imbalance, subsampling can reduce computational costs at the price of inflated variance in estimating model parameters. We propose a method for subsampling efficiently for logistic regression by adjusting the class balance locally in feature space via an accept-reject …
A new method estimates Bayes error for deep networks, suggesting they may have reached the limit.
Method samples triangulations of manifolds using biased random walks.
The presence of noisy instances in mobile phone data is a fundamental issue for classifying user phone call behavior (i.e., accept, reject, missed and outgoing), with many potential negative consequences. The classification accuracy may decrease and the complexity of the classifiers may increase due to the number of re…
We propose a new metaheuristic training scheme that combines Stochastic Gradient Descent (SGD) and Discrete Optimization in an unconventional way. Our idea is to define a discrete neighborhood of the current SGD point containing a number of "potentially good moves" that exploit gradient information, and to search this …
This paper improves parameter estimation in cardiac models using Gaussian process-based MH sampling.
A dealer manages quotes and rejection rules to control slippage risk in FX markets.
MESSY estimation recovers symbolic density functions from samples using maximum entropy.
Paper proposes a human-algorithm approach to reduce medical device recall risk and workload.