Private minimum Hellinger distance estimators maintain robustness and efficiency while ensuring privacy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We extend techniques due to Pardon to show that there is a lower bound on the distortion of a knot in proportional to the minimum of the bridge distance and the bridge number of the knot. We also exhibit an infinite family of knots for which the minimum of the bridge distance and the bridge number is unb…
New examples show flip distance and polyhedron triangulation numbers differ, with ratio close to 3/2.
The development of algorithms for unsupervised pattern recognition by nonlinear clustering is a notable problem in data science. Markov clustering (MCL) is a renowned algorithm that simulates stochastic flows on a network of sample similarities to detect the structural organization of clusters in the data, but it has n…
We investigate a robust penalized logistic regression algorithm based on a minimum distance criterion. Influential outliers are often associated with the explosion of parameter vector estimates, but in the context of standard logistic regression, the bias due to outliers always causes the parameter vector to implode, t…
In this article we study point configurations minimizing the discrete energy on a compact Riemannian manifold, where the energy kernel is taken to be the Green's function for the Laplacian. We show that every point in a minimizing configuration lies inside an open set called harmonic ball where no other point can enter…
Unified proof of knot unknotting bounds using Ma-Qiu index.
We propose a minimum distance estimation method for robust regression in sparse high-dimensional settings. The traditional likelihood-based estimators lack resilience against outliers, a critical issue when dealing with high-dimensional noisy data. Our method, Minimum Distance Lasso (MD-Lasso), combines minimum distanc…
Paper proposes MWDE for estimating finite location-scale mixtures.
Region-based classification of PolSAR data can be effectively performed by seeking for the assignment that minimizes a distance between prototypes and segments. Silva et al (2013) used stochastic distances between complex multivariate Wishart models which, differently from other measures, are computationally tractable.…
Study robust distribution estimation with Wasserstein distance, achieving optimal risk.
In this work, a novel solution to the speaker identification problem is proposed through minimization of statistical divergences between the probability distribution (g). of feature vectors from the test utterance and the probability distributions of the feature vector corresponding to the speaker classes. This approac…
This work improves understanding of projection robust optimal transport distances.
Constructs portfolios based on Hellinger distance to normal, finding market invariance.
Alternative method improves SVM for data classification.
In this paper, we use the distance comparison principle, first been developed by G. Huisken, to study the spatial curve shortening flow. We have got the result that if the initial curve is the helix, then the local minimum of the ratio of the extrinsic and intrinsic distance is non-decreasing. And we have proved a Gray…
This work develops a generic framework, called the bag-of-paths (BoP), for link and network data analysis. The central idea is to assign a probability distribution on the set of all paths in a network. More precisely, a Gibbs-Boltzmann distribution is defined over a bag of paths in a network, that is, on a representati…
Estimates population profile from small random samples.
Study entropic regularization of Gaussian measures and processes on Hilbert space.
Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We study two variants th…
Graph regularized autoencoder improves anomaly detection performance.
Improved MMD estimator for likelihood-free inference.
New estimator handles covariate shift with closed-form solution and super-efficiency.
In this letter, we derive the optimal discriminant functions for modulation classification based on the sampled distribution distance. The proposed method classifies various candidate constellations using a low complexity approach based on the distribution distance at specific testpoints along the cumulative distributi…
Proposes a new model to maximize out-of-sample Sharpe ratios by forecasting tangency portfolios.
Global optimization problems whose objective function is expensive to evaluate can be solved effectively by recursively fitting a surrogate function to function samples and minimizing an acquisition function to generate new samples. The acquisition step trades off between seeking for a new optimization vector where the…
Though deep learning has been applied successfully in many scenarios, malicious inputs with human-imperceptible perturbations can make it vulnerable in real applications. This paper proposes an error-correcting neural network (ECNN) that combines a set of binary classifiers to combat adversarial examples in the multi-c…
We propose a geometric method for quantifying the difference between parametrized curves in Euclidean space by introducing a distance function on the space of parametrized curves up to rigid transformations (rotations and translations). Given two curves, the distance between them is defined as the infimum of an energy …
New bounds on curve distances on surfaces of arbitrary genus.
MultiDendrograms is a Java-written application that computes agglomerative hierarchical clusterings of data. Starting from a distances (or weights) matrix, MultiDendrograms is able to calculate its dendrograms using the most common agglomerative hierarchical clustering methods. The application implements a variable-gro…
In case of a standard form vN-algebra, the Bures distance is the natural distance between the fibres of implementing vectors at normal positive linear forms. Thereby, it is well-known that to each two normal positive linear forms implementing vectors exist such that the Bures distance is attained by the metric distance…
A knot K in 1-bridge position with respect to a genus-g Heegaard surface in a 3-manifold can be moved by isotopy through knots in 1-bridge position until it lies in a union of n parallel genus-g surfaces tubed together by n-1 straight tubes, with K intersecting each tube in two arcs connecting the ends. We prove that t…
This paper presents a novel method to compute the exact Kantorovich-Wasserstein distance between a pair of -dimensional histograms having bins each. We prove that this problem is equivalent to an uncapacitated minimum cost flow problem on a -partite graph with nodes and arcs,…
We show that the volume of any Riemannian metric on a three sphere is bounded below by the length of the shortest closed curve that links its antipodal image. In particular, the volume is bounded below by the minimum of the length of the shortest closed geodesic and the minimal distance between antipodal points.
New formula and algorithm for computing distances on complex Riemann surfaces.
New insights into correntropy-based regression reveal robustness and unified approaches.
Plug-in robust NPE method adapts summaries independently of pretrained NPE.
Lumbermark clusters data robustly, slicing limbs of mutual reachability trees.
We study colorings of the hyperbolic plane, analogously to the Hadwiger-Nelson problem for the Euclidean plane. The idea is to color points using the minimum number of colors such that no two points at distance exactly are of the same color. The problem depends on and, following a strategy of Kloeckner, we show…
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from \emph{unknown} continuous distributions, which form clusters with each cluster containing a composite set of closely located distributions (based on a certain distance metric between di…
Study on Bayesian reinforcement learning performance bounds.
A method for quickly determining deployment schedules that meet a given fuel cycle demand is presented here. This algorithm is fast enough to perform in situ within low-fidelity fuel cycle simulators. It uses Gaussian process regression models to predict the production curve as a function of time and the number of depl…
We investigate the use of Minimax distances to extract in a nonparametric way the features that capture the unknown underlying patterns and structures in the data. We develop a general-purpose and computationally efficient framework to employ Minimax distances with many machine learning methods that perform on numerica…
We prove that S^2 x S^2 satisfies an intermediate condition between having metrics with positive Ricci and positive sectional curvature. Namely, there exist metrics for which the average of the sectional curvatures of any two planes tangent at the same point, but separated by a minimum distance in the 2-Grassmannian, i…
We study a new class of codes for lossy compression with the squared-error distortion criterion, designed using the statistical framework of high-dimensional linear regression. Codewords are linear combinations of subsets of columns of a design matrix. Called a Sparse Superposition or Sparse Regression codebook, this s…
The volume distance from a point p to a convex hypersurface M of the (N+1)-dimensional space is defined as the minimum (N+1)-volume of a region bounded by M and a hyperplane H through the point. This function is differentiable in a neighborhood of M and if we restrict its hessian to the minimizing hyperplane H(p) we ob…
Proves minimum number of normals to curves in 3D space.
This paper establishes information-theoretic limits in estimating a finite field low-rank matrix given random linear measurements of it. These linear measurements are obtained by taking inner products of the low-rank matrix with random sensing matrices. Necessary and sufficient conditions on the number of measurements …