New method estimates density ratio for well-separated distributions using multi-class logistic regression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New findings on robust learning with well-separated data.
Consistent estimator for mixtures of nonparametric elliptical distributions helps cluster analysis.
The moduli space metric and its Kahler potential for well-separated non-Abelian vortices are obtained in U(N) gauge theories with N Higgs fields in the fundamental representation.
EM algorithm achieves optimal sample complexity for well-separated Gaussian mixtures.
New algorithm learns POMDPs without computational oracles.
Stochastic Neighbor Embedding and its variants are widely used dimensionality reduction techniques -- despite their popularity, no theoretical results are known. We prove that the optimal SNE embedding of well-separated clusters from high dimensions to any Euclidean space R^d manages to successfully separate the cluste…
This study examines when non-parametric methods are robust to adversarial examples.
Algorithm distinguishes Gaussian mixtures from pure Gaussians in quasi-polynomial time.
The paper analyzes the risk of CV-tuned regularized estimators and connects it to SURE.
New algorithm for planning in observable POMDPs in quasi-polynomial time.
We introduce a convex approach for mixed linear regression over features. This approach is a second-order cone program, based on L1 minimization, which assigns an estimate regression coefficient in for each data point. These estimates can then be clustered using, for example, -means. For problem…
At critical coupling, the interactions of Ginzburg-Landau vortices are determined by the metric on the moduli space of static solutions. The asymptotic form of the metric for two well separated vortices is shown here to be expressible in terms of a Bessel function. A straightforward extension gives the metric for N vor…
Recent progress has shown that few-shot learning can be improved with access to unlabelled data, known as semi-supervised few-shot learning(SS-FSL). We introduce an SS-FSL approach, dubbed as Prototypical Random Walk Networks(PRWN), built on top of Prototypical Networks (PN). We develop a random walk semi-supervised lo…
Improved sample efficiency with normalized RBF kernels in neural networks.
Improved KSD test for better detection of differences in distributions.
DDSME outperforms SME in estimating multimodal distributions.
Neural networks have many successful applications, while much less theoretical understanding has been gained. Towards bridging this gap, we study the problem of learning a two-layer overparameterized ReLU neural network for multi-class classification via stochastic gradient descent (SGD) from random initialization. In …
Sampling from posterior distributions using Markov chain Monte Carlo (MCMC) methods can require an exhaustive number of iterations, particularly when the posterior is multi-modal as the MCMC sampler can become trapped in a local mode for a large number of iterations. In this paper, we introduce the pseudo-extended MCMC…
The paper cleans label noise in supervised classification using Bernoulli sampling.
We analyze the spectral clustering procedure for identifying coarse structure in a data set , and in particular study the geometry of graph Laplacian embeddings which form the basis for spectral clustering algorithms. More precisely, we assume that the data is sampled from a mixture model supported on …
The study bounds the stability of Gaussian mixtures under small perturbations.
The main contribution of the paper is to show that Gaussian sketching of a kernel-Gram matrix yields an operator whose counterpart in an RKHS , is a \emph{random projection} operator---in the spirit of Johnson-Lindenstrauss (J-L) lemma. To be precise, given a random matrix with i.i.d. Ga…
We present a simple noise-robust margin-based active learning algorithm to find homogeneous (passing the origin) linear separators and analyze its error convergence when labels are corrupted by noise. We show that when the imposed noise satisfies the Tsybakov low noise condition (Mammen, Tsybakov, and others 1999; Tsyb…
Hybrid model for multimodal distributions using diffusion and classification.
Sparse subspace clustering (SSC) is an elegant approach for unsupervised segmentation if the data points of each cluster are located in linear subspaces. This model applies, for instance, in motion segmentation if some restrictions on the camera model hold. SSC requires that problems based on the -norm are solved …
A key task in Bayesian statistics is sampling from distributions that are only specified up to a partition function (i.e., constant of proportionality). However, without any assumptions, sampling (even approximately) can be #P-hard, and few works have provided "beyond worst-case" guarantees for such settings. For log-c…
KPCA improves OoD detection by separating InD and OoD data.
A new Wasserstein -means method for clustering probability distributions.
Neural networks can interpolate noisy data and still generalize well.
Relative moduli spaces of periodic monopoles provide novel examples of Asymptotically Locally Flat hyperkahler manifolds. By considering the interactions between well-separated periodic monopoles, we infer the asymptotic behavior of their metrics. When the monopole moduli space is four-dimensional, this construction yi…
t-distributed Stochastic Neighborhood Embedding (t-SNE), a clustering and visualization method proposed by van der Maaten & Hinton in 2008, has rapidly become a standard tool in a number of natural sciences. Despite its overwhelming success, there is a distinct lack of mathematical foundations and the inner workings of…
While Bayesian neural networks (BNNs) hold the promise of being flexible, well-calibrated statistical models, inference often requires approximations whose consequences are poorly understood. We study the quality of common variational methods in approximating the Bayesian predictive distribution. For single-hidden laye…
Dasgupta and Shulman showed that a two-round variant of the EM algorithm can learn mixture of Gaussian distributions with near optimal precision with high probability if the Gaussian distributions are well separated and if the dimension is sufficiently high. In this paper, we generalize their theory to learning mixture…
GAT-GMM improves GANs' performance in learning Gaussian mixture models.
Researchers create initial data for multiple collapsing boson stars.
Suppose M is a compact orientable irreducible 3-manifold with Heegaard splitting surfaces P and Q. Then either Q is isotopic to a possibly stabilized copy of P or the Hempel distance of the splitting P is no greater than twice the genus of Q. More generally, if P and Q are bicompressible but weakly incompressible conne…
New indices for determining cluster compactness and separability.
EDD uses entropy of distance distributions to cluster unlabeled data.
A new method improves Bayesian inference for multimodal posteriors.
We investigate the problem of nodes clustering under privacy constraints when representing a dataset as a graph. Our contribution is threefold. First we formally define the concept of differential privacy for structured databases such as graphs, and give an alternative definition based on a new neighborhood notion betw…
Deep-embedding methods aim to discover representations of a domain that make explicit the domain's class structure and thereby support few-shot learning. Disentangling methods aim to make explicit compositional or factorial structure. We combine these two active but independent lines of research and propose a new parad…
This work introduces a novel method to evaluate generative model novelty.
Irregular features disrupt the desired classification. In this paper, we consider aggressively modifying scales of features in the original space according to the label information to form well-separated clusters in low-dimensional space. The proposed method exploits spectral clustering to derive scaling factors that a…
Constructs classifiers for neural networks with specific data configurations.
On a complete manifold, such as Euclidean 3-space or hyperbolic 3-space, the limit at infinity of the norm of the Higgs field is called the mass of the monopole. We show the existence, on hypebolic 3-space, of monopoles with given magnetic charge and arbitrary mass. Previously, aside from charge one monopoles, existenc…
Graph convolutional networks (GCNs) are a widely used method for graph representation learning. To elucidate the capabilities and limitations of GCNs, we investigate their power, as a function of their number of layers, to distinguish between different random graph models (corresponding to different class-conditional d…
Constructs initial data for multiple black holes with specified ADM parameters.