Research examines the distribution of curve components in random multicurves.
problem Distribution of curve components in random multicurves.
method Action of the mapping class group on random multicurves.
result Distribution of curve components analyzed.
In the present work, eigenvalue distributions defined by a random rectangular matrix whose components are neither independently nor identically distributed are analyzed using replica analysis and belief propagation. In particular, we consider the case in which the components are independently but not identically distri…
Training on mixed distributions improves test performance even when components are unrelated.
problem Improving test performance with mismatched training and test distributions.
method Analyzing mixture distributions with different training and test proportions.
result Distribution shift can be beneficial, improving test performance even when components are unrelated.
Two derivations of PCA for distributional data.
problem PCA for datasets of distributions.
method Two derivations: variance maximization and reconstruction error minimization.
result Closed-form solution for distributional PCA.
Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.
New supervised and unsupervised NFLTs for elliptical distributions.
problem Understanding unsupervised No Free Lunch Theorems for elliptical distributions.
method Proved two equally optimal strategies for elliptical distributions, inspired PRIM-based bump-hunting algorithms.
result Optimal strategies for selecting principal components based on variance or volume.
Recently, Generative Adversarial Networks (GANs) have emerged as a popular alternative for modeling complex high dimensional distributions. Most of the existing works implicitly assume that the clean samples from the target distribution are easily available. However, in many applications, this assumption is violated. I…
We consider the estimation of Dirichlet Process Mixture Models (DPMMs) in distributed environments, where data are distributed across multiple computing nodes. A key advantage of Bayesian nonparametric models such as DPMMs is that they allow new components to be introduced on the fly as needed. This, however, posts an …
Autoencoder estimates parameters of noisy, multi-component damped signals.
problem Parameter estimation of damped sinusoidal signals under rapid decay and noise.
method Autoencoder-based approach using latent space for frequency, phase, decay, and amplitude estimation.
result High accuracy in parameter estimation, robustness to subdominant components and phase differences.
We study the problem of learning a mixture model of non-parametric product distributions. The problem of learning a mixture model is that of finding the component distributions along with the mixing weights using observed samples generated from the mixture. The problem is well-studied in the parametric setting, i.e., w…
Discover causal structure from mixtures of DAGs using latent variable algorithms.
problem Discover causal structure from distributions arising from mixtures of DAGs.
method Causal structure discovery algorithms such as FCI for latent variables.
result Recover a 'union' of the component DAGs and identify varying conditional distributions.
Quaternion self-attention reduces computational cost and improves performance.
problem Existing quaternion self-attention increases computational cost and diverges attention distributions.
method Proposes a shared-score quaternion self-attention mechanism.
result Reduces score-computation multiplications by 75% and softmax operations from four to one.
Random surfaces with boundary have predictable properties.
problem Understanding the statistical properties of random surfaces.
method Generating surfaces by gluing polygons and analyzing their genus and boundary components.
result Genus and boundary components of random surfaces follow a bivariate normal distribution.
New method identifies shared components from unpaired multimodal mixtures.
problem Identify shared components from unpaired multimodal mixtures.
method Distribution divergence minimization-based loss with sufficient conditions for identifiability.
result Sufficient conditions for shared component identifiability from unaligned multimodal mixtures.
We propose to formulate multi-label learning as a estimation of class distribution in a non-linear embedding space, where for each label, its positive data embeddings and negative data embeddings distribute compactly to form a positive component and negative component respectively, while the positive component and nega…
Domain adaptation framework identifies latent variables for target distribution identifiability.
problem Unsupervised domain adaptation without identifiable joint distribution of features and labels.
method Formulated latent variable model with invariant and changing components, constrained domain shift to influence only changing components.
result Joint distribution of data and labels in target domain is identifiable under mild conditions.
Score-based methods fail with isolated components and incorrect mixing proportions.
problem Score-based methods struggle with distributions having isolated components and incorrect mixing proportions.
method Score-based methods, including score matching, are used but fail in the presence of isolated components and incorrect mixing proportions.
result Score-based methods cannot discover isolated components or identify correct mixing proportions.
New algorithm selects GMM components robustly and efficiently.
problem Estimating the number of components in GMMs when not known in advance.
method Robust model selection for GMMs with poly(k/ε) samples.
result Constructs a GMM with O(k) components approximating the distribution within ε.
New method uses birth-death process and exploration component to accelerate sampling from multimodal distributions.
problem Sampling from multimodal probability distributions efficiently.
method Combines birth-death process and exploration component to accelerate sampling.
result Proves exponential asymptotic convergence under mild assumptions.
This paper tackles distributed estimation of the top-L eigenspace in PCA for large data sets.
problem Challenges in estimating the top-L eigenspace in principal component analysis for large data sets.
method Proposes a novel multi-round algorithm using shift-and-invert preconditioning and convex optimization.
result Achieves a fast convergence rate and covers the targeted top-L eigenspace without explicit eigengap assumption.
Study properties of pointwise k-slant submanifolds in Kähler manifolds.
problem Characterize the integrability of component distributions in Kähler manifolds.
method Characterization through integrability and totally geodesic cases.
result Characterize the integrability of component distributions in Kähler manifolds.
DPA autoencoders learn data distribution and intrinsic dimensionality with guarantees.
problem Learning data distribution and intrinsic dimensionality in unsupervised learning.
method Combines distributionally correct reconstruction with principal-component-like interpretability.
result Exact theoretical guarantees on disentangling factors of variation and intrinsic dimensionality.
Many methods for machine learning rely on approximate inference from intractable probability distributions. Variational inference approximates such distributions by tractable models that can be subsequently used for approximate inference. Learning sufficiently accurate approximations requires a rich model family and ca…
This paper improves PPCA robustness using t-distributions.
problem Improving robustness of probabilistic PCA.
method Using multivariate t-distributions and a hierarchical model. result Clarified the correct correspondence between the multivariate t-PPCA framework and the hierarchical model. We propose a kernel method to identify finite mixtures of nonparametric product distributions. It is based on a Hilbert space embedding of the joint distribution. The rank of the constructed tensor is equal to the number of mixture components. We present an algorithm to recover the components by partitioning the data p…
Many financial variables are found to exhibit multifractal nature, which is usually attributed to the influence of temporal correlations and fat-tailedness in the probability distribution (PDF). Based on the partition function approach of multifractal analysis, we show that there is a marked finite-size effect in the d…
Score matching errors are not sufficient for measuring diffusion model quality.
problem The L2 score matching error is not a reliable measure of diffusion model performance. method Decomposed score errors into gradient and solenoidal components and analyzed their geometric properties.
result Only the gradient component of the score error affects the marginal distributional quality.
Analyzes geodesic lengths in sparse networks, deriving a distribution.
problem Understanding connectivity and robustness in networked systems.
method Analytic derivation of geodesic length distribution in sparse networks.
result Simple closed-form expression for geodesic length distribution.
New geometric analysis shows L2 score error is flawed for diffusion models.
problem Score matching errors in diffusion models do not fully capture distributional quality.
method Decomposed score errors into gradient and solenoidal components, focusing on gradient's role in Fokker-Planck dynamics.
result Only gradient component affects marginal distributional quality; solenoidal component is structurally invisible.
Paper solves NGCA for discrete distributions using LLL method.
problem Learning hidden non-Gaussian components in discrete distributions.
method Utilizes LLL lattice basis reduction method.
result Sample and computationally efficient algorithm for NGCA in discrete distributions.
Simplifies denoising score matching for manifold learning.
problem Learning distributions on manifolds is computationally intensive.
method Modifies denoising score matching to implicitly account for the manifold.
result Reduces computational burden while maintaining efficiency.
Proposes BATer for improved adversarial example detection.
problem Detecting adversarial examples in neural networks.
method Introduces a Bayesian adversarial example detector (BATer) using random components in a Bayesian neural network.
result BATer outperforms state-of-the-art detectors in adversarial example detection.
This paper develops GPCA for probability distributions using Otto-Wasserstein geometry.
problem Analyzing modes of variation in datasets of probability measures.
method Geodesic Principal Component Analysis (GPCA) on Wasserstein space with neural networks.
result Identification of geodesic curves that capture modes of variation in probability distributions.
System learns to combine multiple model components for personalized text generation.
problem Adapting and biasing language models for personal preferences.
method Combines model-defined components, learns activation and probability combination from unlabeled text.
result Directly generates text with personalized components from unlabeled data.
Langevin Dynamics fails to sample from mixture distributions efficiently.
problem Analyzing Langevin Dynamics for sampling from mixture distributions.
method Theoretical analysis of Langevin Dynamics and proposing Chained-Langevin Dynamics.
result Langevin Dynamics fails to sample from mixture distributions efficiently.
Model predicts credit portfolio losses with contagion effects.
problem Predicting credit portfolio losses with contagion effects.
method Introduced a model with a recursive algorithm and flexible distributions.
result Good fit for synthetic CDO tranches of the iTraxx index.
The paper develops methods to accurately locate change points in high-dimensional mean shift models.
problem Locating change points in high-dimensional mean shift models.
method Locally refitted least squares estimator, component-wise and simultaneous rates of estimation.
result Asymptotic validity of component-wise and simultaneous confidence intervals for change point parameters.
This paper presents a methodology for creating streaming, distributed inference algorithms for Bayesian nonparametric (BNP) models. In the proposed framework, processing nodes receive a sequence of data minibatches, compute a variational posterior for each, and make asynchronous streaming updates to a central model. In…
Generalized principal component analysis (GLM-PCA) facilitates dimension reduction of non-normally distributed data. We provide a detailed derivation of GLM-PCA with a focus on optimization. We also demonstrate how to incorporate covariates, and suggest post-processing transformations to improve interpretability of lat…
Principal component analysis (PCA) is very popular to perform dimension reduction. The selection of the number of significant components is essential but often based on some practical heuristics depending on the application. Only few works have proposed a probabilistic approach able to infer the number of significant c…
OT-ICA uses optimal transport to find independent components, outperforming traditional methods.
problem Finding independent components from linear mixtures of signals.
method OT-ICA uses the squared Wasserstein distance to maximize non-Gaussianity, optimizing projections via gradient descent.
result OT-ICA outperforms traditional proxy-based methods in various applications.
Fourier PCA is Principal Component Analysis of a matrix obtained from higher order derivatives of the logarithm of the Fourier transform of a distribution.We make this method algorithmic by developing a tensor decomposition method for a pair of tensors sharing the same vectors in rank-1 decompositions. Our main appli…
Study identifies components of unknown interventions in a mixture.
problem Identify components of a mixture of unknown interventions on a causal Bayesian Network.
method Construct example showing components not identifiable. Prove identifiability under mild conditions. Develop efficient algorithm for recovery. Analyze performance in simulation.
result Components of a mixture of unknown interventions can be uniquely identified under certain conditions.
Fuses posterior distributions from different datasets using KL divergence.
problem Combining information from multiple datasets with uncertainty.
method Mean field assumption, KL divergence, assign-and-average approach.
result Efficient non-parametric algorithm for fused model computation.
Novel prior for orthogonal functions improves functional component estimation.
problem Improving orthogonality in functional principal component analysis.
method Sequential adaptive priors for orthogonal functions using hierarchical conditionally normal distributions.
result Proposed prior leads to nearly orthogonal posterior estimates.
A new method detects changes in mixture models quickly and accurately.
problem Detecting changes in mixture models with heavy-tailed components.
method Change-point methods based on robust and quick approach.
result The method is up to 500 times faster and more accurate than existing methods.
Optimal transport is #P-hard when components are independent, even with approximate solutions.
problem Computational complexity of optimal transport with independent marginals.
method Proved #P-hardness and developed a pseudo-polynomial time approximation algorithm.
result Optimal transport is #P-hard even with independent components and approximate solutions.
New algorithm speeds Bayesian nonparametric model inference.
problem Slow inference in Bayesian nonparametric models.
method Decompose random measures into finite and infinite sub-measures; use different algorithms for each.
result Hybrid algorithm improves scalability and mixing.