Unified Bayesian methods for sparse signal recovery using scale mixtures.
problem Sparse signal recovery using various priors and methods.
method Unified MAP and Type II Bayesian approaches using Power Exponential Scale Mixture (PESM) family.
result Type II methods show better support recovery than Type I methods.
Higher granularity in MoE models boosts expressivity exponentially.
problem Expressivity of Mixture-of-Experts models with varying granularity.
method Comparing models with different numbers of active experts (granularity).
result Exponential separation in network expressivity based on granularity.
In this paper we propose a novel framework for the construction of sparsity-inducing priors. In particular, we define such priors as a mixture of exponential power distributions with a generalized inverse Gaussian density (EP-GIG). EP-GIG is a variant of generalized hyperbolic distributions, and the special cases inclu…
Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.
problem Achieving optimal error rates in clustering sub-exponential mixture models.
method Establishes universal lower bounds and demonstrates iterative algorithms' optimality in sub-exponential mixture models.
result Iterative algorithms achieve the universal lower bound in sub-exponential mixture models.
Extends Stein's lemma to exponential-family mixtures for gradient computation.
problem Computing gradients for complex distributions with weak assumptions.
method Generalizes Stein's lemma to exponential-family mixtures and applies it to reparameterization trick.
result Derives new gradient identities for various distributions.
ELU algorithm improves on EM for over-specified Gaussian mixtures.
problem Slow convergence of EM in over-specified Gaussian mixtures.
method Developed ELU algorithm for two-component mixtures, combining exponential location update and gradient descent.
result ELU converges to final statistical radius after logarithmic iterations, resolving open question.
Unified framework for sparse non-negative least squares and matrix factorization.
problem Recovering non-negative quantities from linear measurements.
method Unified rectified power exponential scale mixture prior and multiplicative update rules.
result Proposed algorithms converge to stationary points of the objective function.
We prove a conjecture about approximating Gaussian Processes on one dimension.
problem Computational scaling issues with Gaussian Processes on one dimension.
method Developed a new family of state-space models (LEG) to approximate any stationary GP on one dimension.
result Proved that any stationary GP on one dimension can be approximated using the LEG family.
In recent years, a rich variety of shrinkage priors have been proposed that have great promise in addressing massive regression problems. In general, these new priors can be expressed as scale mixtures of normals, but have more complex forms and better properties than traditional Cauchy and double exponential priors. W…
New pruning method breaks power law scaling, potentially reducing error to exponential.
problem Improving neural network performance through scaling alone is costly.
method Developed a new data pruning metric to break power law scaling.
result Pruned datasets show better than power law scaling on various image datasets.
New spectral mixture representation for isotropic kernels simplifies random Fourier features.
problem Applying Random Fourier Features to complex kernels.
method Decompose isotropic kernels into scale mixtures of α-stable random vectors.
result Constructive spectral sampling formula for various kernels.
We derive relations between theoretical properties of restricted Boltzmann machines (RBMs), popular machine learning models which form the building blocks of deep learning models, and several natural notions from discrete mathematics and convex geometry. We give implications and equivalences relating RBM-representable …
New algorithm speeds up sampling from complex Bayesian mixture models.
problem Sampling from non-log-concave, multi-modal posterior distributions in Bayesian Gaussian mixtures.
method Introduced Reflected Metropolis-Hastings Random Walk (RMRW) algorithm.
result Proved mixing time bound for RMRW in symmetric two-component Gaussian mixtures.
Extends likelihood ratio exponential families to analyze various optimization methods.
problem Analyzing optimization methods like rate-distortion and information bottleneck.
method Linking geometric mixture paths to exponential families and using hypothesis testing.
result Provides a common mathematical framework for understanding these methods.
Mixture models are a fundamental tool in applied statistics and machine learning for treating data taken from multiple subpopulations. The current practice for estimating the parameters of such models relies on local search heuristics (e.g., the EM algorithm) which are prone to failure, and existing consistent methods …
MoEs can efficiently model complex tasks with low-dimensionality and sparsity.
problem Understanding the theoretical foundations of MoEs for complex tasks.
method Systematic study of MoEs with two structural priors: low-dimensionality and sparsity.
result MoEs can approximate functions on low-dimensional manifolds and exhibit exponential structured tasks.
We develop coresets for various clustering problems using Bregman divergences.
problem Efficiently clustering large datasets with provable accuracy.
method Proposed a general algorithm to construct strong coresets for a wide range of clustering problems.
result Demonstrated practicality and efficiency of the algorithm through empirical evaluation.
A new method normalizes flow mixtures for better inference across different data types.
problem Inference failure across diverse posterior geometries in normalizing flows.
method Introduces a two-stage framework with a stable global weighting mechanism based on sEMA.
result Achieves consistent NLL improvements and stable weight trajectories over baselines.
AutoGMM automates Gaussian mixture modeling in Python.
problem Automatic clustering of complex data with uncertainty-aware grouping.
method Strategic initialization using an agglomerative Mahalanobis heuristic, parallelized model selection by information criteria.
result Strong out-of-the-box performance on classic benchmarks and real datasets.
Modeling financial returns as conditionally independent random variables explains power-law tails.
problem Understanding the distribution of financial returns and their relation to volatility.
method Assuming returns are conditionally independent given volatility, which varies randomly over time.
result Returns distribution can be described by the sum of conditionally independent random variables, showing scaling and power-law tails.
We propose the Bayesian bridge estimator for regularized regression and classification. Two key mixture representations for the Bayesian bridge model are developed: (1) a scale mixture of normals with respect to an alpha-stable random variable; and (2) a mixture of Bartlett--Fejer kernels (or triangle densities) with r…
Study reveals a universal formula for knotting in random equilateral polygons.
problem Probability of knotting in equilateral random polygons.
method Extensive Monte Carlo simulations with improved algorithms and knot invariants.
result A universal scaling formula for knotting probability with number of edges, involving exponential and power law factors.
Develops a semi-supervised learning method using exponential tilt mixture models.
problem Improves classification accuracy with labeled and unlabeled data.
method Extends logistic regression to exponential tilt modeling, derives maximum likelihood estimation, and proposes regularized estimation.
result Demonstrates improved prediction accuracy compared to existing methods.
We provide guarantees for learning latent variable models emphasizing on the overcomplete regime, where the dimensionality of the latent space can exceed the observed dimensionality. In particular, we consider multiview mixtures, spherical Gaussian mixtures, ICA, and sparse coding models. We provide tight concentration…
We study the problem of finding the smallest m such that every element of an exponential family can be written as a mixture of m elements of another exponential family. We propose an approach based on coverings and packings of the face lattice of the corresponding convex support polytopes and results from coding th…
Algorithm learns mixtures of linear regressions in subexponential time.
problem Learning mixtures of linear regressions with high accuracy.
method Fourier moment descent method using univariate density estimation and low-degree moments of Fourier transforms.
result First algorithm for learning MLRs in subexponential time.
The paper analyzes EM for Mixtures of Experts and shows its equivalence to projected Mirror Descent.
problem Training Mixtures of Experts (MoE) models.
method Rigorously analyzes Expectation Maximization (EM) for MoE models using a Mirror Descent perspective.
result Derives new convergence results and identifies conditions for local linear convergence.
Study calculates tail risk for various mixture distributions.
problem Estimating tail risk for complex distribution mixtures.
method Analyzes tail conditional expectation for location-scale mixtures of elliptical distributions.
result Developed methods for calculating tail risk in various distributions.
We introduce a stochastic model to explain a double power-law distribution which exhibits two different Paretian behaviors in the upper and the lower tail and widely exists in social and economic systems. The model incorporates fitness consideration and noise fluctuation. We find that if the number of variables (e.g. t…
SCI-PI solves scale invariant problems efficiently.
problem Solving scale invariant problems in optimization.
method Introduces SCI-PI and proves its convergence.
result SCI-PI achieves local linear convergence.
We construct an infinite-dimensional information manifold based on exponential Orlicz spaces without using the notion of exponential convergence. We then show that convex mixtures of probability densities lie on the same connected component of this manifold, and characterize the class of densities for which this mixtur…
New theorem improves spectral gap for sampling from mixture distributions.
problem Sampling from multimodal distributions with simulated tempering.
method Introduced a decomposition theorem for the restricted spectral gap of simulated tempering.
result Lower bound on the restricted spectral gap for mixture distributions.
We study by theoretical analysis and by direct numerical simulation the dynamics of a wide class of asynchronous stochastic systems composed of many autocatalytic degrees of freedom. We describe the generic emergence of truncated power laws in the size distribution of their individual elements. The exponents α of the…
The paper introduces a new class of multivariate mixtures for actuarial applications.
problem Developing a new class of multivariate mixtures for actuarial calculations.
method Proposed a class of multivariate matrix-exponential affine mixtures with matrix-exponential marginals.
result Explicit calculations of actuarial quantities are possible due to the proposed class's properties.
Appropriately designing the proposal kernel of particle filters is an issue of significant importance, since a bad choice may lead to deterioration of the particle sample and, consequently, waste of computational power. In this paper we introduce a novel algorithm adaptively approximating the so-called optimal proposal…
New model explains volatility after extreme stock market events.
problem Understanding volatility dynamics after extreme stock market events.
method Proposed a new dynamical model using high frequency minute data.
result Volatility after extreme events follows a stretched exponential decay initially and a power law decay later.
This paper studies stability of the exponential utility maximization when there are small variations on agent's utility function. Two settings are considered. First, in a general semimartingale model where random endowments are present, a sequence of utilities defined on R converges to the exponential utility. Under a …
We extend natural-gradient methods to mixtures of exponential-family distributions, improving inference speed.
problem Complex, multimodal posterior distributions are difficult to approximate with simple exponential-family distributions.
method We use minimal conditional-EF representations and derive simple natural-gradient updates.
result Our natural-gradient method converges faster than black-box methods with reparameterization gradients.
Sampling can be faster than optimization in nonconvex settings.
problem Limited theoretical understanding of optimization vs sampling efficiency.
method Examined nonconvex objective functions in mixture modeling and multi-stable systems.
result Sampling algorithms are linearly scalable in model dimension, while optimization algorithms are exponentially scalable.
Optimizes portfolios with GM returns using convex optimization.
problem Maximizing expected exponential utility with GM asset returns.
method Formulated as a convex optimization problem.
result Optimal solutions found without sampling or scenarios.
We present the data on wealth and income distributions in the United Kingdom, as well as on the income distributions in the individual states of the USA. In all of these data, we find that the great majority of population is described by an exponential distribution, whereas the high-end tail follows a power law. The di…
Proposes a differentiable LSE-ICNN for modeling multi-well potentials.
problem Modeling multi-well potentials in various scientific domains.
method Log-sum-exponential (LSE) mixture of input convex neural network (ICNN) modes.
result Smooth surrogate that retains convexity within basins and allows gradient-based learning.
VI approximates complex densities in Bayesian stats.
problem Approximating complex posterior densities in Bayesian stats.
method Optimizing a family of densities close to the target via Kullback-Leibler divergence.
result VI can be faster than classical methods like MCMC.
Estimates statistical power for cluster analysis in biomedical research.
problem Lack of established methods to compute a priori statistical power for cluster analysis.
method Simulation studies varying subgroup size, number, separation, and covariance structure.
result Sufficient statistical power achieved with small samples (N=20-30) for large effect sizes.
Develops a new class of forward performance processes for investment pools.
problem Investment performance in market models with continuous semimartingale stock prices.
method Constructs a broad class of forward performance processes with power mixture initial conditions.
result Characterizes and derives properties of two-power mixture forward performance processes.
A new model for Gaussian process experts tackles scalability and uncertainty issues.
problem Scalability and excessive number of experts degrade predictive performance and increase uncertainty.
method Nested partitioning scheme infers the number of components, a generalised GP framework accommodates multiple response types, and a factorised exponential family structure handles multiple input types.
result Effectiveness demonstrated on synthetic data and an Alzheimer's challenge dataset.
Proposes EMG mixture model for spectroscopy data.
problem Modeling residuals in spectroscopy data with positive support.
method Exponentially-modified Gaussian mixture (EMG) model with expectation-maximization algorithm.
result EMG mixture outperforms existing models in spectroscopy applications.
This work explains scaling laws as redundancy laws in deep learning.
problem The mathematical origins of scaling laws in deep learning models remain unclear.
method Kernel regression and analysis of data covariance spectra.
result Scaling laws can be explained as redundancy laws, revealing the learning curve's slope depends on data redundancy.