Study shows SW distance estimators are consistent and asymptotically valid for generative models.
problem Theoretical guarantees for SW distance in generative models.
method Investigation of asymptotic properties of SW distance estimators.
result Asymptotic consistency and central limit theorem for SW distance estimators.
Private minimum Hellinger distance estimators maintain robustness and efficiency while ensuring privacy.
problem Ensuring privacy in robust statistical estimation.
method Derive private minimum Hellinger distance estimators satisfying Hellinger differential privacy.
result Private minimum Hellinger distance estimators retain robustness and efficiency under privacy constraints.
FKEE estimates expectations without samples, using diffusion bridges and PINNs.
problem Estimating expectations without large sample sizes.
method Diffusion bridge models and Feynman-Kac operator approximation using PINNs.
result Significantly reduces variance and improves efficiency.
We propose a minimum distance estimation method for robust regression in sparse high-dimensional settings. The traditional likelihood-based estimators lack resilience against outliers, a critical issue when dealing with high-dimensional noisy data. Our method, Minimum Distance Lasso (MD-Lasso), combines minimum distanc…
We investigate a robust penalized logistic regression algorithm based on a minimum distance criterion. Influential outliers are often associated with the explosion of parameter vector estimates, but in the context of standard logistic regression, the bias due to outliers always causes the parameter vector to implode, t…
Study robust distribution estimation with Wasserstein distance, achieving optimal risk.
problem Robust distribution estimation under adversarial corruption.
method Combining partial OT and minimum distance estimation, proving structural properties and deriving a novel dual form.
result Achieves minimax-optimal robust estimation risk in many settings.
Paper proposes MWDE for estimating finite location-scale mixtures.
problem Estimating finite location-scale mixtures using MLE is problematic.
method Investigates minimum Wasserstein distance estimators (MWDE).
result MWDE is consistent and provides a numerical solution.
Improved MMD estimator for likelihood-free inference.
problem Computational challenges in estimating MMD for likelihood-free inference.
method Optimally-weighted MMD estimator with improved sample complexity.
result Significantly improved sample complexity for accurate MMD estimation.
Generalizes robust statistics to various perturbations under Wasserstein distance.
problem Robust statistics for datasets corrupted by various perturbations.
method Generalizes robust statistics to any Wasserstein distance, showing robust estimation under certain perturbations.
result Generalized resilience property holds under moment or hypercontractive conditions, simplifying and improving known results.
New estimator handles covariate shift with closed-form solution and super-efficiency.
problem Handling covariate shift in missing data and causal inference problems.
method Minimum Wasserstein distance estimation framework.
result Closed-form expression and super-efficiency relative to semiparametric efficient estimator.
Estimates population profile from small random samples.
problem Learning population composition from limited data.
method Minimum distance estimator based on linear programming.
result Consistent estimation of profile in sublinear sample size.
MC-MCL improves MCL for nonlinear clustering.
problem Nonlinear clustering in data science.
method MC-MCL combines MCL with Minimum Curvilinearity for nonlinear distances.
result MC-MCL outperforms classical MCL and baseline clustering algorithms in nonlinear datasets.
This work improves understanding of projection robust optimal transport distances.
problem Understanding the behavior of minimum Wasserstein estimators in high-dimensional and misspecified models.
method Adopting projection robust (PR) optimal transport, establishing statistical properties, proposing IPRW distance, and providing asymptotic guarantees.
result Established fundamental statistical properties and proposed new distances that outperform Wasserstein distances empirically.
Optimal transport framework for density estimation with constraints.
problem Density estimation under expectation constraints.
method Minimizes Wasserstein distance subject to expected value constraints and regularization.
result Framework effectively addresses non-smooth constraints through annealing-like algorithm.
In this work, a novel solution to the speaker identification problem is proposed through minimization of statistical divergences between the probability distribution (g). of feature vectors from the test utterance and the probability distributions of the feature vector corresponding to the speaker classes. This approac…
Defines MER for Bayesian learning, a gap between achievable and optimal performance.
problem Analyzing the best performance of Bayesian learning under generative models.
method Two methods for deriving upper bounds for MER: conditional mutual information and minimum estimation error.
result Quantifies the rate at which MER decays to zero with more data and relates it to model richness.
Algorithm reduces historical expected shortfall computation by focusing on worst-case scenarios.
problem Computing the historical expected shortfall efficiently and accurately.
method Multi-step algorithm using Monte Carlo simulations to identify and reduce the number of worst-case scenarios.
result Non-asymptotic bounds for the L p-error of the expected shortfall estimator are derived.
Approximations of loopy belief propagation, including expectation propagation and approximate message passing, have attracted considerable attention for probabilistic inference problems. This paper proposes and analyzes a generalization of Opper and Winther's expectation consistent (EC) approximate inference method. Th…
Proposes a new model to maximize out-of-sample Sharpe ratios by forecasting tangency portfolios.
problem Maximizing Sharpe ratios when returns and covariances are not stationary.
method Forecast the tangency portfolio using vector autoregressions and invest in the minimum Euclidean distance portfolio.
result Empirically validated superior out-of-sample Sharpe ratios.
Study shows the corrected Akaike criterion is inadmissible for estimating Kullback-Leibler discrepancy.
problem Inadmissibility of the corrected Akaike information criterion for estimating Kullback-Leibler discrepancy.
method Loss estimation framework to demonstrate inadmissibility and provide improved estimators.
result Improved estimators of Kullback-Leibler discrepancy are provided and perform well in reduced-rank situations.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
Efficient global optimization is the problem of minimizing an unknown function f, using as few evaluations f(x) as possible. It can be considered as a continuum-armed bandit problem, with noiseless data and simple regret. Expected improvement is perhaps the most popular method for solving this problem; the algorithm pe…
We extend techniques due to Pardon to show that there is a lower bound on the distortion of a knot in R3 proportional to the minimum of the bridge distance and the bridge number of the knot. We also exhibit an infinite family of knots for which the minimum of the bridge distance and the bridge number is unb…
Paper proposes deep neural networks for nonparametric regression from dependent data.
problem Nonparametric regression from strongly mixing observations.
method Minimum error entropy principle applied to deep neural networks.
result Deep neural networks achieve minimax optimal convergence rates for Gaussian errors.
Proposes variance reduction techniques for sliced Wasserstein distance estimation.
problem Intractability of estimating sliced Wasserstein distances.
method Uses control variates based on Gaussian approximations of projected measures.
result Significant reduction in variance of SW distance estimators.
Study on Bayesian reinforcement learning performance bounds.
problem Achieving optimal performance in model-based Bayesian reinforcement learning.
method Defining minimum Bayesian regret, deriving upper bounds using relative entropy and Wasserstein distance, and applying these to specific MDP cases.
result Upper bounds on minimum Bayesian regret for MDPs, including specific cases like MAB and online optimization with partial feedback.
Inference for normal and Monte Carlo distributions using minimum relative entropy.
problem Inference from partial information on expectations and covariances.
method Minimum relative entropy sub-manifolds, analytical formulas, Monte Carlo simulations.
result Improved numerical implementation for inference from partial information.
New insights into correntropy-based regression reveal robustness and unified approaches.
problem Learning robust regression functions under additive noise.
method Minimum distance estimation and conditional mean, mode, median functions.
result Unified approach to conditional mean, mode, and median functions.
The paper develops tests for comparing means in high dimensions with unknown covariance.
problem Testing if the mean of a high-dimensional distribution is close to zero or different from another.
method Develops nonasymptotic tests using concentration inequalities and operator norms.
result Obtains bounds on the minimal separation distance for controlling Type I and Type II errors.
New examples show flip distance and polyhedron triangulation numbers differ, with ratio close to 3/2.
problem Understanding the relationship between flip distance and polyhedron triangulation numbers.
method Provided examples to demonstrate the difference between flip distance and polyhedron triangulation numbers.
result Ratio of flip distance to polyhedron triangulation numbers can be arbitrarily close to 3/2.
Asynchronous Gibbs sampling has been recently shown to be fast-mixing and an accurate method for estimating probabilities of events on a small number of variables of a graphical model satisfying Dobrushin's condition~\cite{DeSaOR16}. We investigate whether it can be used to accurately estimate expectations of functions…
Plug-in robust NPE method adapts summaries independently of pretrained NPE.
problem Misspecification of neural posterior estimators under test data distribution.
method Minimum-distance summaries using maximum mean discrepancy (MMD).
result Substantial robustness gains with minimal additional overhead.
The paper converts metric bounds to distance function Hölder bounds and proves compactness theorems.
problem Proving geometric stability results with scalar curvature bounds.
method Transforming Lp bounds to Hölder bounds for distance functions. result Compactness theorems and convergence guarantees for Riemannian manifolds.
A non-Euclidean generalization of conditional expectation is introduced and characterized as the minimizer of expected intrinsic squared-distance from a manifold-valued target. The computational tractable formulation expresses the non-convex optimization problem as transformations of Euclidean conditional expectation. …
We consider the problem of subspace estimation in a Bayesian setting. Since we are operating in the Grassmann manifold, the usual approach which consists of minimizing the mean square error (MSE) between the true subspace U and its estimate U^ may not be adequate as the MSE is not the natural metric in the Gra…
In this article we study point configurations minimizing the discrete energy on a compact Riemannian manifold, where the energy kernel is taken to be the Green's function for the Laplacian. We show that every point in a minimizing configuration lies inside an open set called harmonic ball where no other point can enter…
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
Large graphs abound in machine learning, data mining, and several related areas. A useful step towards analyzing such graphs is that of obtaining certain summary statistics - e.g., or the expected length of a shortest path between two nodes, or the expected weight of a minimum spanning tree of the graph, etc. These sta…
Algorithm identifies nearest mode in noisy data.
problem Identifying the point with the minimum k-th nearest neighbor distance in unknown multivariate probability density.
method Sequential learning algorithm using noisy oracle queries to adaptively decide which points to query.
result Upper bounds on query complexity show significant improvement over baselines.
Understanding separation effects on parameter estimation in finite Gaussian mixtures
problem Minimum component separation impact on convergence rates in finite Gaussian mixtures
method Developing a unified geometric framework using Hellinger lower bounds and specialized moment-extraction test functions
result Separation complexity driven by spatial configuration of mixture components
Unified proof of knot unknotting bounds using Ma-Qiu index.
problem Finding bounds on the number of moves to unknot knots.
method Using the Ma-Qiu index to bound presentation distances and Gordian distances.
result Unified proof of various unknotting number bounds.
Distance-based hierarchical clustering (HC) methods are widely used in unsupervised data analysis but few authors take account of uncertainty in the distance data. We incorporate a statistical model of the uncertainty through corruption or noise in the pairwise distances and investigate the problem of estimating the HC…
Estimating entropy and mutual information consistently is important for many machine learning applications. The Kozachenko-Leonenko (KL) estimator (Kozachenko & Leonenko, 1987) is a widely used nonparametric estimator for the entropy of multivariate continuous random variables, as well as the basis of the mutual inform…
New smoothing technique improves Wasserstein distance estimation in high dimensions.
problem Estimating statistical distances between high-dimensional distributions.
method Gaussian smoothing of p-Wasserstein distance and analysis of its asymptotic behavior. result Gaussian-smoothed p-Wasserstein distance converges at rate n−1/2, improving over n−1/d for unsmoothed distances. Diffusion models achieve nearly optimal distribution estimation in various spaces.
problem Theoretical limitations of diffusion modeling for distribution estimation.
method Analysis of approximation and generalization abilities of diffusion models in Besov spaces.
result Diffusion models achieve nearly minimax optimal estimation rates in total variation and Wasserstein distances.
Estimates generalization error for two-layer ReLU NNs through minimum norm solutions.
problem Estimating generalization error for two-layer ReLU NNs trained by mean squared error.
method Uses minimum norm solutions and Neural Tangent Kernel (NTK) regime to derive generalization error bounds.
result Derives an a priori generalization error bound for two-layer ReLU NNs without requiring exponentially large number of neurons.
In the past decade many researchers have proposed new optimal portfolio selection strategies to show that sophisticated diversification can outperform the naïve 1/N strategy in out-of-sample benchmarks. Providing an updated review of these models since DeMiguel et al. (2009b), I test sixteen strategies across six empir…
This thesis improves kernel-based distances for statistical inference and integration.
problem Efficiently measuring distances between probability distributions for robust and smooth modeling.
method Kernel-based distances, focusing on maximum mean discrepancy (MMD) and novel kernel quantile discrepancies.
result Improved MMD estimators for simulation-based inference and conditional expectations.