Private minimum Hellinger distance estimators maintain robustness and efficiency while ensuring privacy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a minimum distance estimation method for robust regression in sparse high-dimensional settings. The traditional likelihood-based estimators lack resilience against outliers, a critical issue when dealing with high-dimensional noisy data. Our method, Minimum Distance Lasso (MD-Lasso), combines minimum distanc…
We investigate a robust penalized logistic regression algorithm based on a minimum distance criterion. Influential outliers are often associated with the explosion of parameter vector estimates, but in the context of standard logistic regression, the bias due to outliers always causes the parameter vector to implode, t…
Study robust distribution estimation with Wasserstein distance, achieving optimal risk.
Paper proposes MWDE for estimating finite location-scale mixtures.
Improved MMD estimator for likelihood-free inference.
New estimator handles covariate shift with closed-form solution and super-efficiency.
The development of algorithms for unsupervised pattern recognition by nonlinear clustering is a notable problem in data science. Markov clustering (MCL) is a renowned algorithm that simulates stochastic flows on a network of sample similarities to detect the structural organization of clusters in the data, but it has n…
Estimates population profile from small random samples.
This work improves understanding of projection robust optimal transport distances.
In this work, a novel solution to the speaker identification problem is proposed through minimization of statistical divergences between the probability distribution (g). of feature vectors from the test utterance and the probability distributions of the feature vector corresponding to the speaker classes. This approac…
Proposes a new model to maximize out-of-sample Sharpe ratios by forecasting tangency portfolios.
We extend techniques due to Pardon to show that there is a lower bound on the distortion of a knot in proportional to the minimum of the bridge distance and the bridge number of the knot. We also exhibit an infinite family of knots for which the minimum of the bridge distance and the bridge number is unb…
New insights into correntropy-based regression reveal robustness and unified approaches.
New examples show flip distance and polyhedron triangulation numbers differ, with ratio close to 3/2.
Plug-in robust NPE method adapts summaries independently of pretrained NPE.
In this article we study point configurations minimizing the discrete energy on a compact Riemannian manifold, where the energy kernel is taken to be the Green's function for the Laplacian. We show that every point in a minimizing configuration lies inside an open set called harmonic ball where no other point can enter…
We consider the problem of subspace estimation in a Bayesian setting. Since we are operating in the Grassmann manifold, the usual approach which consists of minimizing the mean square error (MSE) between the true subspace and its estimate may not be adequate as the MSE is not the natural metric in the Gra…
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
Algorithm identifies nearest mode in noisy data.
Unified proof of knot unknotting bounds using Ma-Qiu index.
Understanding separation effects on parameter estimation in finite Gaussian mixtures
New smoothing technique improves Wasserstein distance estimation in high dimensions.
Minimum expected distance estimation (MEDE) algorithms have been widely used for probabilistic models with intractable likelihood functions and they have become increasingly popular due to their use in implicit generative modeling (e.g. Wasserstein generative adversarial networks, Wasserstein autoencoders). Emerging fr…
Robust statistics traditionally focuses on outliers, or perturbations in total variation distance. However, a dataset could be corrupted in many other ways, such as systematic measurement errors and missing covariates. We generalize the robust statistics approach to consider perturbations under any Wasserstein distance…
Recently, a framework for application-oriented optimal experiment design has been introduced. In this context, the distance of the estimated system from the true one is measured in terms of a particular end-performance metric. This treatment leads to superior unknown system estimates to classical experiment designs bas…
New algorithms improve robust estimation in contaminated Gaussian models.
We consider the problem of estimating undirected triangle-free graphs of high dimensional distributions. Triangle-free graphs form a rich graph family which allows arbitrary loopy structures but 3-cliques. For inferential tractability, we propose a graphical Fermat's principle to regularize the distribution family. Suc…
Region-based classification of PolSAR data can be effectively performed by seeking for the assignment that minimizes a distance between prototypes and segments. Silva et al (2013) used stochastic distances between complex multivariate Wishart models which, differently from other measures, are computationally tractable.…
We consider the problem of high-dimensional Ising (graphical) model selection. We propose a simple algorithm for structure estimation based on the thresholding of the empirical conditional variation distances. We introduce a novel criterion for tractable graph families, where this method is efficient, based on the pres…
Constructs portfolios based on Hellinger distance to normal, finding market invariance.
Alternative method improves SVM for data classification.
In this paper, we use the distance comparison principle, first been developed by G. Huisken, to study the spatial curve shortening flow. We have got the result that if the initial curve is the helix, then the local minimum of the ratio of the extrinsic and intrinsic distance is non-decreasing. And we have proved a Gray…
A reliable, accurate, and affordable positioning service is highly required in wireless networks. In this paper, the novel Message Passing Hybrid Localization (MPHL) algorithm is proposed to solve the problem of cooperative distributed localization using distance and direction estimates. This hybrid approach combines t…
This work develops a generic framework, called the bag-of-paths (BoP), for link and network data analysis. The central idea is to assign a probability distribution on the set of all paths in a network. More precisely, a Gibbs-Boltzmann distribution is defined over a bag of paths in a network, that is, on a representati…
While likelihood-based inference and its variants provide a statistically efficient and widely applicable approach to parametric inference, their application to models involving intractable likelihoods poses challenges. In this work, we study a class of minimum distance estimators for intractable generative models, tha…
A well-known Lemma in Riemannian geometry by Klingenberg says that if is a minimum point of the distance function to in the cut locus of , then either there is a minimal geodesic from to along which they are conjugate, or there is a geodesic loop at that smoothly goes throu…
Paper introduces a new robust method for estimating Pareto tail index from grouped data.
This paper establishes information-theoretic limits in estimating a finite field low-rank matrix given random linear measurements of it. These linear measurements are obtained by taking inner products of the low-rank matrix with random sensing matrices. Necessary and sufficient conditions on the number of measurements …
Study entropic regularization of Gaussian measures and processes on Hilbert space.
Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We study two variants th…
Graph regularized autoencoder improves anomaly detection performance.
The theory of tunnel number 1 knots detailed in our previous paper, The tree of knot tunnels, provides a non-negative integer invariant called the depth of the tunnel. We give various results related to the depth invariant. Noting that it equals the minimum number of Goda-Scharlemann-Thompson tunnel moves needed to con…
In this work, we present a method to compute the Kantorovich-Wasserstein distance of order one between a pair of two-dimensional histograms. Recent works in Computer Vision and Machine Learning have shown the benefits of measuring Wasserstein distances of order one between histograms with bins, by solving a classic…
In this letter, we derive the optimal discriminant functions for modulation classification based on the sampled distribution distance. The proposed method classifies various candidate constellations using a low complexity approach based on the distribution distance at specific testpoints along the cumulative distributi…
Optimal estimator derived for partially observable LTI systems.
A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested …
We generalize to tree graphs obtained by connecting path graphs an oracle result obtained for the Fused Lasso over the path graph. Moreover we show that it is possible to substitute in the oracle inequality the minimum of the distances between jumps by their harmonic mean. In doing so we prove a lower bound on the comp…