Separating mixed distributions is a long standing challenge for machine learning and signal processing. Most current methods either rely on making strong assumptions on the source distributions or rely on having training samples of each source in the mixture. In this work, we introduce a new method---Neural Egg Separat…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Guiding the design of neural networks is of great importance to save enormous resources consumed on empirical decisions of architectural parameters. This paper constructs shallow sigmoid-type neural networks that achieve 100% accuracy in classification for datasets following a linear separability condition. The separab…
The paper solves optimal bounds for separating data points in high dimensions.
New conditions ensure MMDs separate and converge to target distributions.
We quantify the separation between the numbers of labeled examples required to learn in two settings: Settings with and without the knowledge of the distribution of the unlabeled data. More specifically, we prove a separation by multiplicative factor for the class of projections over the Boolean hypercube o…
Despite substantial progress in signal source separation, results for richly structured data continue to contain perceptible artifacts. In contrast, recent deep generative models can produce authentic samples in a variety of domains that are indistinguishable from samples of the data distribution. This paper introduces…
Neural networks use their hidden layers to transform input data into linearly separable data clusters, with a linear or a perceptron type output layer making the final projection on the line perpendicular to the discriminating hyperplane. For complex data with multimodal distributions this transformation is difficult t…
DSI measures dataset separability for neural networks.
The paper extends optimal transport for linear separability of sheared distributions in supervised learning.
We propose a new blind source separation algorithm based on mixtures of alpha-stable distributions. Complex symmetric alpha-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, infer…
Algorithm clusters mixtures with bounded covariances under specific separation conditions.
Enhanced FastMNMF for better speech separation.
Consistent estimator for mixtures of nonparametric elliptical distributions helps cluster analysis.
We provide new results concerning label efficient, polynomial time, passive and active learning of linear separators. We prove that active learning provides an exponential improvement over PAC (passive) learning of homogeneous linear separators under nearly log-concave distributions. Building on this, we provide a comp…
Single-microphone, speaker-independent speech separation is normally performed through two steps: (i) separating the specific speech sources, and (ii) determining the best output-label assignment to find the separation error. The second step is the main obstacle in training neural networks for speech separation. Recent…
Develops a new framework for causal models on cyclic graphs, solving unique solvability issues.
Mixup reduces the sample complexity of finding optimal decision boundaries for more separable data.
Graph convolution improves linear separability and generalizes to out-of-distribution data.
We introduce a new approach for designing computationally efficient learning algorithms that are tolerant to noise, and demonstrate its effectiveness by designing algorithms with improved noise tolerance guarantees for learning linear separators. We consider both the malicious noise model and the adversarial label nois…
Generalizes underlap coefficient for multivariate group separation.
Kernel embeddings separate distinct probability distributions, simplifying testing.
Separable losses are inconsistent for structured prediction models.
New method estimates density ratio for well-separated distributions using multi-class logistic regression.
In this paper, we presented a novel semi-supervised one-class classification algorithm which assumes that class is linearly separable from other elements. We proved theoretically that class is linearly separable if and only if it is maximal by probability within the sets with the same mean. Furthermore, we presented an…
KPCA improves OoD detection by separating InD and OoD data.
Paper develops robust methods for panel data with latent groups, improving inference under group separation violations.
Adversarial noises are linearly separable for random neural networks.
We consider the closeness testing problem for discrete distributions. The goal is to distinguish whether two samples are drawn from the same unspecified distribution, or whether their respective distributions are separated in -norm. In this paper, we focus on adapting the rate to the shape of the underlying distri…
New indices for determining cluster compactness and separability.
Generative source separation methods such as non-negative matrix factorization (NMF) or auto-encoders, rely on the assumption of an output probability density. Generative Adversarial Networks (GANs) can learn data distributions without needing a parametric assumption on the output density. We show on a speech source se…
Categorical d-separation criterion simplifies probability graph analysis.
Develops large-sample theory for non-stationary source separation.
New methods test discrete distributions faster with local privacy constraints.
SAL framework uses unlabeled data to improve OOD detection.
Separates estimation and control in risk-sensitive investment problems with partial observation.
The separability assumption (Donoho & Stodden, 2003; Arora et al., 2012) turns non-negative matrix factorization (NMF) into a tractable problem. Recently, a new class of provably-correct NMF algorithms have emerged under this assumption. In this paper, we reformulate the separable NMF problem as that of finding the ext…
Federated learning studies separate client data and distribution gaps.
This paper proposes a multichannel source separation technique called the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class l…
We study exact recovery conditions for convex relaxations of point cloud clustering problems, focusing on two of the most common optimization problems for unsupervised clustering: -means and -median clustering. Motivations for focusing on convex relaxations are: (a) they come with a certificate of optimality, and…
Polynomial-time algorithm for clustering mixtures with separation Δ=Ω(√(log k)).
The simplicial condition and other stronger conditions that imply it have recently played a central role in developing polynomial time algorithms with provable asymptotic consistency and sample complexity guarantees for topic estimation in separable topic models. Of these algorithms, those that rely solely on the simpl…
While on some natural distributions, neural-networks are trained efficiently using gradient-based algorithms, it is known that learning them is computationally hard in the worst-case. To separate hard from easy to learn distributions, we observe the property of local correlation: correlation between local patterns of t…
Estimates statistical power for cluster analysis in biomedical research.
High temporal resolution measurements of human brain activity can be performed by recording the electric potentials on the scalp surface (electroencephalography, EEG), or by recording the magnetic fields near the surface of the head (magnetoencephalography, MEG). The analysis of the data is problematic due to the fact …
Measures mode separation in high-dimensional densities via a reversible diffusion process.
Study high-dimensional Bayesian linear regression using variational inference.
Standard methods in deep learning for natural language processing fail to capture the compositional structure of human language that allows for systematic generalization outside of the training distribution. However, human learners readily generalize in this way, e.g. by applying known grammatical rules to novel words.…
This paper strengthens the computational separation between multimodal and unimodal learning, showing unimodal learning is hard on typical instances.