Extends L2-norm LDA to 2D inputs using Bhattacharyya bound.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on error probability for classification of heavy-tailed renewal processes.
Paper connects rejection learning to Bhattacharyya divergence.
Geometrically refines Cramér-Rao bound using extrinsic manifold curvature.
New CRB derived for curved models using extrinsic geometry.
In this paper, we propose a novel linear discriminant analysis criterion via the Bhattacharyya error bound estimation based on a novel L1-norm (L1BLDA) and L2-norm (L2BLDA). Both L1BLDA and L2BLDA maximize the between-class scatters which are measured by the weighted pairwise distances of class means and meanwhile mini…
Efficiently learns polytrees with known skeleton in polynomial time and sample complexity.
Study reconstructs hidden perfect matchings in random graphs with specific edge weights.
A framework for disentangling class-related and class-independent factors in data.
Efficient unsupervised training and inference in deep generative models remains a challenging problem. One basic approach, called Helmholtz machine, involves training a top-down directed generative model together with a bottom-up auxiliary model used for approximate inference. Recent results indicate that better genera…
Mixture distributions arise in many parametric and non-parametric settings -- for example, in Gaussian mixture models and in non-parametric estimation. It is often necessary to compute the entropy of a mixture, but, in most cases, this quantity has no closed-form expression, making some form of approximation necessary.…
Article provides Bernstein gradient estimates for heat equations with potential terms.
In many machine learning problems, labeled training data is limited but unlabeled data is ample. Some of these problems have instances that can be factored into multiple views, each of which is nearly sufficent in determining the correct labels. In this paper we present a new algorithm for probabilistic multi-view lear…
We propose a representation of graph as a functional object derived from the power iteration of the underlying adjacency matrix. The proposed functional representation is a graph invariant, i.e., the functional remains unchanged under any reordering of the vertices. This property eliminates the difficulty of handling e…
We study strictly proper scoring rules in the Reproducing Kernel Hilbert Space. We propose a general Kernel Scoring rule and associated Kernel Divergence. We consider conditions under which the Kernel Score is strictly proper. We then demonstrate that the Kernel Score includes the Maximum Mean Discrepancy as a special …
Algorithm learns latent simplex from perturbed points in input-sparsity time.
Unified geometric interpretation of statistical estimation inequalities.
The clustering algorithms that view each object data as a single sample drawn from a certain distribution, Gaussian distribution, for example, has been a hot topic for decades. Many clustering algorithms: such as k-means and spectral clustering are proposed based on the single sample assumption. However, in real life, …
Efficiently learns Gaussian tree models with near-optimal sample complexity.
We develop a novel methodology based on the marriage between the Bhattacharyya distance, a measure of similarity across distributions of random variables, and the Johnson-Lindenstrauss Lemma, a technique for dimension reduction. The resulting technique is a simple yet powerful tool that allows comparisons between data-…
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
The scaled complex Wishart distribution is a widely used model for multilook full polarimetric SAR data whose adequacy has been attested in the literature. Classification, segmentation, and image analysis techniques which depend on this model have been devised, and many of them employ some type of dissimilarity measure…
Develops geometric framework for uncertainty-aware multi-class classification.
There has been a growing interest in mutual information measures due to their wide range of applications in Machine Learning and Computer Vision. In this paper, we present a generalized structured regression framework based on Shama-Mittal divergence, a relative entropy measure, which is introduced to the Machine Learn…
We quantify forgetting in post-training models, distinguishing mass and drift.
Hybrid clustering combines partitional and hierarchical clustering for computational effectiveness and versatility in cluster shape. In such clustering, a dissimilarity measure plays a crucial role in the hierarchical merging. The dissimilarity measure has great impact on the final clustering, and data-independent prop…
Images obtained with coherent illumination, as is the case of sonar, ultrasound-B, laser and Synthetic Aperture Radar -- SAR, are affected by speckle noise which reduces the ability to extract information from the data. Specialized techniques are required to deal with such imagery, which has been modeled by the G0 dist…
Deep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a deep random feature (DRF) expansion model, which makes inference tractable by app…
In the field of statistics, many kind of divergence functions have been studied as an amount which measures the discrepancy between two probability distributions. In the differential geometrical approach in statistics (information geometry), dually flat spaces play a key role. In a dually flat space, there exist dual a…
The space of probability distributions on a given sample space possesses natural geometric properties. For example, in the case of a smooth parametric family of probability distributions on the real line, the parameter space has a Riemannian structure induced by the embedding of the family into the Hilbert space of squ…
New algorithms SVCA and SSPA improve robustness to noise in nonnegative matrix factorization.
The paper introduces a new risk measure for financial models with jumps.
Wireless sensor networks usually comprise a large number of sensors monitoring changes in variables. These changes in variables represent changes in physical quantities. The changes can occur for various reasons; these reasons are highlighted in this work. Outliers are unusual measurements. Outliers are important; they…
Market Microstructure is the investigation of the process and protocols that govern the exchange of assets with the objective of reducing frictions that can impede the transfer. In financial markets, where there is an abundance of recorded information, this translates to the study of the dynamic relationships between o…
We explore the connection between two problems that have arisen independently in the signal processing and related fields: the estimation of the geometric mean of a set of symmetric positive definite (SPD) matrices and their approximate joint diagonalization (AJD). Today there is a considerable interest in estimating t…
The study optimizes bounds for comparing training and population loss.
Introduces bounded scale measure and generalizes property A.
Paper improves PAC-Bayes bounds for various loss types.
Improved bounds for Monte Carlo Rademacher Averages using self-bounding functions.
Study bounds on self-shrinkers with bounded HA for applications.
Investigates tight PAC-Bayes bounds for small datasets.
Extends Fatou theorem to bounded harmonic maps.
New bound relaxes uniform gradient norm assumptions for PAC-Bayesian bounds.
Jiang et al. (2020) found no uniformly tight generalization bounds for neural networks in the overparameterized setting.
Willmore-type inequalities for bounded domains in manifolds with curvature bounds.
Lower bounds on curvature integral for manifolds with curvature constraints.
Paper improves SLCB regret bound for bounded noise.
Study on CMC hypersurfaces with bounded index and area, proving multiplicity one convergence and bounds on genus.