We study the problem of learning a mixture model of non-parametric product distributions. The problem of learning a mixture model is that of finding the component distributions along with the mixing weights using observed samples generated from the mixture. The problem is well-studied in the parametric setting, i.e., w…
Non-parametric time series forecasting without assuming a specific distribution.
problem Time series forecasting with numerical stability issues in classical models.
method Generates predictions by sampling from the empirical distribution of time series data.
result The proposed method produces reasonable forecasts without numerical stability issues.
A novel MCMC method clusters data faster and more accurately.
problem Efficiently clustering large datasets with unknown number of clusters.
method Master/Worker architecture for distributed MCMC inference.
result Significant improvement in clustering accuracy and speed.
The paper reinterprets Bayesian priors and posteriors using Riemannian manifolds.
problem The dependence of maximum a posteriori estimates on parametrization.
method Assuming a Riemannian manifold with Fisher metric, the paper reinterprets priors and posteriors as distributions over probability distributions, making estimates independent of parametrization.
result A maximum a posteriori estimate independent of parametrization is defined.
We solve the mean parametrization of von Mises-Fisher distribution.
problem No closed-form normalization function for mean parameters exists.
method Derived a second-order ODE for mean normalizer and provided approximations.
result Rapid evaluation of densities and natural parameters in terms of mean parameters.
We develop quantile regression models in order to derive risk margin and to evaluate capital in non-life insurance applications. By utilizing the entire range of conditional quantile functions, especially higher quantile levels, we detail how quantile regression is capable of providing an accurate estimation of risk ma…
Improved spatial distribution learning with Bayesian transport maps and parametric shrinkage.
problem Learning non-Gaussian spatial distributions with limited training data.
method Proposed ShrinkTM approach using Bayesian transport maps with parametric shrinkage.
result ShrinkTM outperforms existing BTM, especially with few training samples.
Develops coresets for scalable multivariate distribution estimation.
problem Handling large-scale data in non-parametric or semi-parametric regression and density estimation.
method Novel coreset construction for multivariate conditional transformation models (MCTMs).
result Substantial data reduction with high log-likelihood accuracy.
Optimal learning for parametric prophet inequalities with exponential-type distributions
problem Learning in prophet inequalities with unknown parameters
method Confidence-based dynamic-programming policy
result Achieves optimal asymptotic competitive ratio using online observations
We consider an investor, whose portfolio consists of a single risky asset and a risk free asset, who wants to maximize his expected utility of the portfolio subject to the Value at Risk assuming a heavy tail distribution of the stock prices return. We use Markov Decision Process and dynamic programming principle to get…
The paper develops a theory for identifying the best arm in non-parametric multi-armed bandits with a fixed budget.
problem Identifying the best arm in non-parametric multi-armed bandits with a limited number of trials.
method The paper proposes upper and lower bounds on the average log-probability of misidentification using information-theoretic quantities and a refined analysis of the successive-rejects strategy.
result The paper provides new upper and lower bounds on the average log-probability of misidentification, which generalize existing bounds.
DPPS uses DP priors for Bayesian non-parametric multi-arm bandits.
problem Optimizing multi-arm bandit environments with prior beliefs.
method Bayesian non-parametric algorithm based on Dirichlet Process priors.
result DPPS provides principled incorporation of prior beliefs and is optimal in Bayesian regret setup.
This paper finds a unique partition of a sample space for estimating continuous distributions.
problem Estimating continuous probability distributions from finite samples.
method Equal-probability partition of the sample space using order statistics.
result The partition yields an entropy of log2(N+1) bits, providing a discrete entropy estimate.
Dirichlet Process(DP) is a Bayesian non-parametric prior for infinite mixture modeling, where the number of mixture components grows with the number of data items. The Hierarchical Dirichlet Process (HDP), is an extension of DP for grouped data, often used for non-parametric topic modeling, where each group is a mixtur…
Although various distributed machine learning schemes have been proposed recently for pure linear models and fully nonparametric models, little attention has been paid on distributed optimization for semi-paramemetric models with multiple-level structures (e.g. sparsity, linearity and nonlinearity). To address these is…
We use variational Gaussian approximations to analyze parametric models with unknown data-generating distributions.
problem Analyzing inference and learning in parametric models with unknown or intractable data-generating distributions.
method Replica method with variational Gaussian approximation in grand canonical formalism.
result Stationarity conditions adaptively determine parameters of the trial Hamiltonian for each dataset.
Flexible copula model using implicit generative neural networks.
problem Limited flexibility of parametric copulas and curse of dimensionality in non-parametric methods.
method Implicit generative neural networks to model high-dimensional copula distributions with unspecified marginals.
result Demonstrated flexibility and performance on various datasets.
Extends DeTEcT framework for token economies with dynamic and probabilistic parameters.
problem Modeling wealth distribution in token economies with dynamic and probabilistic parameters.
method Introduces four parametrization techniques: dynamic vs static, probabilistic vs non-probabilistic.
result Derives existing wealth distribution models from DeTEcT framework with added restrictions.
Parametric adversarial divergences, which are a generalization of the losses used to train generative adversarial networks (GANs), have often been described as being approximations of their nonparametric counterparts, such as the Jensen-Shannon divergence, which can be derived under the so-called optimal discriminator …
New neural network models extreme value distributions with preserved shape constraints.
problem Modeling multivariate extreme value distributions with preserved shape constraints.
method d-max-decreasing neural network architecture for non-parametric calibration and generation of MEVs.
result The proposed architecture approximates the dependence structure of MEVs at parametric rate and preserves essential shape constraints.
A parametrization of hypergraphs based on the geometry of points in Rd is developed. Informative prior distributions on hypergraphs are induced through this parametrization by priors on point configurations via spatial processes. This prior specification is used to infer conditional independence models or M…
Proposes a parametric t-SNE without perplexity tuning.
problem Non-parametric t-SNE's perplexity parameter limits DR quality.
method Multi-scale parametric t-SNE with deep neural network.
result Produces reliable embeddings with competitive neighborhood preservation.
A tractable pseudo-metric for non-parametric distributions via SPD geometry.
problem Computing distances between non-parametric probability distributions is intractable.
method Two-stage framework: projection onto parametric family, embedding into SPD matrices.
result Closed-form pseudo-metric for two-sample hypothesis testing.
Paper simplifies data carving inference with a parametric distribution.
problem Valid inference after selection with data carving.
method Developed a parametric distribution for data carving inference.
result Exact inference for data carving can be computed trivially.
A nonparametric two-sample test using a parametric integral probability metric
problem Detecting distributional differences between two independent samples
method Propose a new two-sample test statistic based on a newly introduced integral probability metric (IPM)
result Establish theoretical guarantees for the associated two-sample testing procedure
This paper studies the rates of convergence for learning distributions implicitly with the adversarial framework and Generative Adversarial Networks (GANs), which subsume Wasserstein, Sobolev, MMD GAN, and Generalized/Simulated Method of Moments (GMM/SMM) as special cases. We study a wide range of parametric and nonpar…
Efficient methods estimate bid and value distributions in auctions.
problem Estimating bid and value distributions in auctions with limited information.
method Non-parametric estimation algorithms for first- and second-price auctions.
result Uniform estimation bounds for bid and value distributions, independent of distributions being estimated.
Study finds a method to discover causal relationships that are invariant to marginal distributions.
problem Current causal discovery methods are sensitive to marginal distributions, leading to unreliable results.
method Proposes a non-parametric estimator that marginalizes the marginals to find intrinsic causal relationships.
result The proposed method yields causal estimators competitive with current methodologies and emphasizes uncertainty.
We introduce a non-parametric method to recover physical probability distributions of asset returns based on their European option prices and some other sparse parametric information. Thus the main problem is similar to the one considered foir instance in the Recovery Theorem by Ross (2015), except that here we conside…
Motivated by the application of real-time pricing in e-commerce platforms, we consider the problem of revenue-maximization in a setting where the seller can leverage contextual information describing the customer's history and the product's type to predict her valuation of the product. However, her true valuation is un…
A non-parametric method for evaluation of the aggregate loss distribution (ALD) by combining and numerically inverting the empirical characteristic functions (CFs) is presented and illustrated. This approach to evaluate ALD is based on purely non-parametric considerations, i.e., based on the empirical CFs of frequency …
Proposes method for eliciting non-parametric joint priors using normalizing flows.
problem Learning complex non-parametric joint priors for model parameters.
method Expert elicitation combined with normalizing flows for generative modeling.
result Framework supports elicitation of both parametric and non-parametric priors.
This work introduces the concept of parametric Gaussian processes (PGPs), which is built upon the seemingly self-contradictory idea of making Gaussian processes parametric. Parametric Gaussian processes, by construction, are designed to operate in "big data" regimes where one is interested in quantifying the uncertaint…
Parametric UMAP learns a mapping from data to embeddings.
problem Representing and learning from structured data.
method Parametric optimization over neural network weights for UMAP.
result Parametric UMAP performs comparably to non-parametric UMAP with faster online embeddings.
Motivated by the need for parametric families of rich and yet tractable distributions in financial mathematics, both in pricing and risk management settings, but also considering wider statistical applications, we investigate a novel technique for introducing skewness or kurtosis into a symmetric or other distribution.…
Improved inference for models with continuous latent variables.
problem Inference accuracy with traditional variational methods is limited.
method Reparameterized Variational Rejection Sampling (RVRS) using a proposal distribution with a reparameterized gradient estimator.
result RVRS offers a better trade-off between computational cost and inference fidelity.
Learning algorithms for implicit generative models can optimize a variety of criteria that measure how the data distribution differs from the implicit model distribution, including the Wasserstein distance, the Energy distance, and the Maximum Mean Discrepancy criterion. A careful look at the geometries induced by thes…
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
problem Online monitoring of multivariate data streams for detecting changes.
method Combines Kernel-QuantTree histogram and EWMA statistic for non-parametric monitoring.
result Controls Average Run Length (ARL0) while achieving comparable detection delays.
The paper improves prediction intervals for non-parametric regression using histograms.
problem Computing accurate prediction intervals for non-parametric regression models.
method Uses conditional histograms to estimate conditional distributions and compute shortest prediction intervals.
result The method provides prediction intervals with provable marginal coverage and asymptotic conditional coverage.
IQ-BART models conditional quantiles using a non-parametric Bayesian approach.
problem Capturing multimodal predictive distributions in time series forecasting.
method Implicit Quantile BART (IQ-BART) augments data with quantile values for non-parametric quantile function estimation.
result IQ-BART provides flexible distribution-free regression with theoretical guarantees.
We propose a Bayesian non-parametric approach for modeling the distribution of multiple returns. In particular, we use an asymmetric dynamic conditional correlation (ADCC) model to estimate the time-varying correlations of financial returns where the individual volatilities are driven by GJR-GARCH models. The ADCC-GJR-…
We argue that a stochastic model of economic exchange, whose steady-state distribution is a Generalized Beta Prime (also known as GB2), and some unique properties of the latter, are the reason for GB2's success in describing wealth/income distributions. We use housing sale prices as a proxy to wealth/income distributio…
Process capability index (PCI) is a commonly used statistic to measure ability of a process to operate within the given specifications or to produce products which meet the required quality specifications. PCI can be univariate or multivariate depending upon the number of process specifications or quality characteristi…
Develops non-parametric tests for group symmetry in data.
problem Lack of statistical tests for group symmetry in data.
method Formulates and implements non-parametric tests for distributional symmetry under specified groups.
result Develops tests for conditional invariance/equivariance and applies them to real-world data.
New model for time series classification from single example.
problem Classifying time series patterns from limited data.
method Developed a Hidden semi-Markov Model with variable state duration.
result Different representations of state duration have distinct strengths and weaknesses.
Method learns conditional distributions using neural entropic optimal transport.
problem Challenges in learning multiple conditional distributions.
method Neural entropic optimal transport method with two networks and regularization.
result Effective learning of conditional distributions with limited samples.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
Study shows improper learning can outperform proper learning in misspecified models.
problem Misspecification in probabilistic prediction models.
method Investigates the performance of proper and improper learning strategies in misspecified models.
result Improper learning can achieve lower regret compared to proper learning, especially in high-dimensional settings.