Paper explores variable skipping to speed up range density estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper develops estimators for unbounded density ratios with applications in error control.
Roundtrip uses deep generative models for flexible density estimation.
Optimizes kernel density ratios for better predictions and information measures.
Kernel Density Machines learn probability densities without structural assumptions.
Autoregressive models are among the best performing neural density estimators. We describe an approach for increasing the flexibility of an autoregressive model, based on modelling the random numbers that the model uses internally when generating data. By constructing a stack of autoregressive models, each modelling th…
Adaptive kernel density estimation improves accuracy in high dimensions.
Given observations from an unknown absolute continuous distribution defined on some domain , we propose a nonparametric method to learn a piecewise constant function to approximate the underlying probability density function. Our density estimate is a piecewise constant function defined on a binary partition o…
Study minimax rates for density estimation under Huber contamination and Besov IPM losses.
Machine learning models, especially based on deep architectures are used in everyday applications ranging from self driving cars to medical diagnostics. It has been shown that such models are dangerously susceptible to adversarial samples, indistinguishable from real samples to human eye, adversarial samples lead to in…
M-flows learn data manifolds and densities, improving manifold learning and inference.
We present a first procedure that can estimate -- with statistical consistency guarantees -- any local-maxima of a density, under benign distributional conditions. The procedure estimates all such local maxima, or , of any bounded shape or dimension, including usual point-modes. In practice, modal-…
This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.
We study in this paper the rate of convergence for learning densities under the Generative Adversarial Networks (GAN) framework, borrowing insights from nonparametric statistics. We introduce an improved GAN estimator that achieves a faster rate, through simultaneously leveraging the level of smoothness in the target d…
Generative model for tabular data density regression.
We introduce a new family of estimators for unnormalized statistical models. Our family of estimators is parameterized by two nonlinear functions and uses a single sample from an auxiliary distribution, generalizing Maximum Likelihood Monte Carlo estimation of Geyer and Thompson (1992). The family is such that we can e…
How can one perform Bayesian inference on stochastic simulators with intractable likelihoods? A recent approach is to learn the posterior from adaptively proposed simulations using neural network-based conditional density estimators. However, existing methods are limited to a narrow range of proposal distributions or r…
Density estimation is a versatile technique underlying many data mining tasks and techniques,ranging from exploration and presentation of static data, to probabilistic classification, or identifying changes or irregularities in streaming data. With the pervasiveness of embedded systems and digitisation, this latter typ…
Bayes Error Rate estimators are evaluated for accuracy and sample requirements.
A new model simulates non-linear adsorption using Gaussian KDEs.
Autoregressive generative models consistently achieve the best results in density estimation tasks involving high dimensional data, such as images or audio. They pose density estimation as a sequence modeling task, where a recurrent neural network (RNN) models the conditional distribution over the next element conditio…
MBORE optimizes multi-objective problems using density-ratio estimation.
Mixture models are powerful statistical models used in many applications ranging from density estimation to clustering and classification. When dealing with mixture models, there are many issues that the experimenter should be aware of and needs to solve. The MixEst toolbox is a powerful and user-friendly package for M…
Kernel density matrices simplify probabilistic deep learning.
Probability Density Estimation (PDE) is a multivariate discrimination technique based on sampling signal and background densities defined by event samples from data or Monte-Carlo (MC) simulations in a multi-dimensional phase space. In this paper, we present a modification of the PDE method that uses a self-adapting bi…
Kernel density estimation (KDE) is a popular statistical technique for estimating the underlying density distribution with minimal assumptions. Although they can be shown to achieve asymptotic estimation optimality for any input distribution, cross-validating for an optimal parameter requires significant computation do…
A probability density function (pdf) encodes the entire stochastic knowledge about data distribution, where data may represent stochastic observations in robotics, transition state pairs in reinforcement learning or any other empirically acquired modality. Inferring data pdf is of prime importance, allowing to analyze …
Discrimination between non-stationarity and long-range dependency is a difficult and long-standing issue in modelling financial time series. This paper uses an adaptive spectral technique which jointly models the non-stationarity and dependency of financial time series in a non-parametric fashion assuming that the time…
Pathfinder uses quasi-Newton optimization for variational inference.
Fully augmented links have dense volume densities but discrete in certain ranges.
The paper studies knot densities under various constraints and degenerations.
Improved flow-based models capture dependencies better with multi-scale autoregressive priors.
This research improves demand forecasting by predicting complete probability density functions using machine learning.
The COS method for European options pricing is improved with a new bound for the number of terms.
This paper compares log-likelihood and BLEU scores for sequence generation tasks.
A new distance metric derived from information theory and estimation theory.
We consider the Cauchy problem for doubly non-linear degenerate parabolic equations on Riemannian manifolds of infinite volume, or in . The equation contains a weight function as a capacitary coefficient which we assume to decay at infinity. We connect the behavior of non-negative solutions to the interplay betwe…
Neural networks help create summary statistics for complex models.
Driving styles have a great influence on vehicle fuel economy, active safety, and drivability. To recognize driving styles of path-tracking behaviors for different divers, a statistical pattern-recognition method is developed to deal with the uncertainty of driving styles or characteristics based on probability density…
WS-KDE provides robust confidence bounds for stochastic functions.
Upper bound on Jones polynomials density modulo primes.
A new metric evaluates generative models by comparing real and generated samples.
A new method for learning conditional distributions using ODEs and neural networks.
A new density model using Fourier basis achieves better approximations and compression.
The paper proposes a method to estimate latent structures in multivariate data without assuming their existence.
New geometric analysis of PWSPDs balances density and geometry in high-dimensional data.
Modes and ridges of the probability density function behind observed data are useful geometric features. Mode-seeking clustering assigns cluster labels by associating data samples with the nearest modes, and estimation of density ridges enables us to find lower-dimensional structures hidden in data. A key technical cha…
We study the task of unsupervised domain adaptation, where no labeled data from the target domain is provided during training time. To deal with the potential discrepancy between the source and target distributions, both in features and labels, we exploit a copula-based regression framework. The benefits of this approa…