A new MLSS model approximates high-order Markov chains efficiently.
problem Efficiently approximating high-order Markov chains.
method Decaying mixture over past states, simple sampling algorithm.
result Approximates high-order Markov chains with fixed time and memory costs.
We consider unsupervised estimation of mixtures of discrete graphical models, where the class variable corresponding to the mixture components is hidden and each mixture component over the observed variables can have a potentially different Markov graph structure and parameters. We propose a novel approach for estimati…
We informally call a stochastic process learnable if it admits a generalization error approaching zero in probability for any concept class with finite VC-dimension (IID processes are the simplest example). A mixture of learnable processes need not be learnable itself, and certainly its generalization error need not de…
SDP relaxation for Sub-Gaussian Mixture Model achieves optimal error bound and robustness.
problem Estimating discrete clustering structures in Sub-Gaussian Mixture Model.
method Hidden integrality property of SDP relaxation and semi-random robustness analysis.
result SDP relaxation achieves optimal error bound and robustness in semi-random setting.
New model explains volatility after extreme stock market events.
problem Understanding volatility dynamics after extreme stock market events.
method Proposed a new dynamical model using high frequency minute data.
result Volatility after extreme events follows a stretched exponential decay initially and a power law decay later.
Nuclear magnetic resonance (NMR) spectroscopy exploits the magnetic properties of atomic nuclei to discover the structure, reaction state and chemical environment of molecules. We propose a probabilistic generative model and inference procedures for NMR spectroscopy. Specifically, we use a weighted sum of trigonometric…
A new law limits kurtosis contrast in balanced mixtures.
problem Kurtosis-based ICA fails in wide, balanced mixtures.
method Proved a redundancy law and showed purification restores contrast.
result Kurtosis contrast obeys O(κmax/Reff) in balanced mixtures. Study improves flow-based model training from few samples.
problem Training flow-based models from limited data.
method Sharp analysis of two-layer autoencoder with finite sample complexity.
result Generative flow approximates target density with rate Θ_n(1/n).
Deep neural networks classify unbounded Gaussian mixture data without dimensionality issues.
problem Binary classification of unbounded Gaussian mixture data.
method Deep ReLU neural networks with non-asymptotic upper bounds and convergence rates.
result Deep ReLU networks can classify unbounded Gaussian mixture data without dimensionality constraints.
MSFA clusters high-dimensional spatial data using spline-based covariance structures.
problem Clustering high-dimensional spatial data with flexible covariance structures.
method Mixture of spatial factor analyzers with spline-based covariance and matrix variate factor analyzers for dimensionality reduction.
result Proposed models accurately infer and differentiate distinct spatial patterns in tensor-variate data.
Study the geometric structure of graph Laplacian embeddings for manifold data.
problem Identifying coarse structure in manifold data sampled from a mixture model.
method Analyze spectral clustering procedure for data sampled from a manifold, focusing on graph Laplacian embeddings.
result Embedded data concentrates on cones centered around orthogonal vectors when the mixture model is well-separated.
One-bit clustering method for two-component sub-Gaussian mixture models
problem Clustering in sub-Gaussian mixture models
method One-bit clustering using dithered quantization
result Decaying misclassification rate with exponential signal-to-noise ratio
Universal model for soft tissue mechanics under shock waves.
problem Modeling shock wave mechanics in soft biological tissues.
method Continuum mixture theory with phase-field mechanics.
result Universal thermodynamically consistent formulation for soft porous tissues.
Paper analyzes Annealed Langevin Dynamics for multimodal sampling stability.
problem Ensuring stability of Annealed Langevin Dynamics across dimensions.
method Uniform-in-dimension analysis of ALD for Gaussian-mixture targets.
result ALD achieves prescribed accuracy in KL divergence with spectral conditions.
In this paper we propose a copula contagion mixture model for correlated default times. The model includes the well known factor, copula, and contagion models as its special cases. The key advantage of such a model is that we can study the interaction of different models and their pricing impact. Specifically, we model…
New method models fat-tailed distributions with anisotropic tail-adaptive flows.
problem Gaussian-based variational inference fails to accurately capture tail decay in fat-tailed distributions.
method Improved theory on tails of flows, developed anisotropic tail-adaptive flows (ATAF).
result ATAF models tail-anisotropy, outperforming prior work on synthetic and real-world targets.
The paper shows how training with synthetic data can lead to model improvement, not degradation, under certain conditions.
problem Model collapse in iterative training on contaminated sources.
method Statistical analysis of iterative training on a mixture of true and synthetic data.
result Training with synthetic data can lead to model improvement, not degradation, under specific conditions.
A neural network method estimates densities from characteristic functions.
problem Estimating fixed-horizon probability densities from empirical characteristic functions.
method Data-driven Fourier-mixture neural-network method trained in Fourier space.
result Competitive performance and clear gains on heavy-tailed targets.
GradPower speeds up language model training with minimal code changes.
problem Slower training of large language models.
method Elementwise sign-power transformation applied to gradients.
result Consistently lower terminal loss across various models and datasets.
We analyze two communication-efficient algorithms for distributed statistical optimization on large-scale data sets. The first algorithm is a standard averaging method that distributes the N data samples evenly to $\nummac$ machines, performs separate minimization on each subset, and then averages the estimates. We p…
Robust clustering algorithm for datasets with outliers.
problem Clustering with arbitrary outliers.
method Spectral clustering with a rounding scheme on a Gaussian kernel matrix.
result Misclassification error decays exponentially with signal-to-noise ratio.
The study uses DPGMM to analyze pulsar families in parameter space.
problem Classifying and understanding different types of pulsars.
method Unsupervised machine learning with Dirichlet process Gaussian mixture model (DPGMM).
result DPGMM provides insights into pulsar families and their relations.
Paper analyzes Langevin dynamics for multimodal Gaussian mixtures, controlling errors across dimensions.
problem Challenges in obtaining stable diffusion-based samplers in high- and infinite-dimensional settings.
method Study of preconditioned Annealed Langevin Dynamics (ALD) for Gaussian mixtures, focusing on Euler-Maruyama (EM) and exponential-integrator schemes.
result Proves dimension-uniform KL bounds for the exponential-integrator scheme, allowing arbitrarily small divergence with dimension.
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
problem Deconvolving mixed DNA profiles from crime samples.
method Multiple population evolutionary algorithm (MEA) with guided mutation.
result The MEA successfully deconvoluted DNA profiles from crime samples.
New stability theory for Sinkhorn semigroups with explicit decay rates.
problem Stability and convergence of Sinkhorn iterations for various divergences.
method Operator-theoretic framework based on Lyapunov techniques.
result Explicit exponential decay rates for Sinkhorn iterates.
This paper introduces a deep learning ensemble forecasting model using Dirichlet process.
problem Forecasting with deep learning ensemble models.
method Infinite mixture model based on Dirichlet process, with decaying learning rate strategy.
result The ensemble model outperforms single benchmark models in prediction accuracy and stability.
Study examines robust regression in high dimensions with heavy-tailed data.
problem Analyzing robust regression in high-dimensional settings with heavy-tailed data.
method Sharp asymptotic characterisation of M-estimators and ridge regression in elliptical distributions.
result Ridge regression is optimal and universal for finite second moments but can decay faster without them.
Suppose k centers are fit to m points by heuristically minimizing the k-means cost; what is the corresponding fit over the source distribution? This question is resolved here for distributions with p≥4 bounded moments; in particular, the difference between the sample cost and distribution cost decays with $…
This study reveals the critical role of scale vectors in large language models, improving optimization and expressivity.
problem Understanding and optimizing the scale vectors in large language models.
method Systematic study of scale vectors from expressivity, optimization, and architectural perspectives; theoretical and empirical analysis of weight decay; proposing and evaluating improvements.
result Scale vectors improve optimization through a self-amplifying preconditioning effect and are beneficial for expressivity in certain architectures.
Extends Minkowski stability proof to minimal decay assumptions.
problem Global stability of Minkowski spacetime with minimal decay.
method Extends Christodoulou-Klainerman's proof to minimal decay assumptions.
result Exterior stability of Minkowski holds with borderline decay.
New method creates vacuum data at minimal and borderline decay thresholds.
problem Creating vacuum initial data at specific decay thresholds.
method Conical solution-operator method applied to vacuum asymptotically flat initial data.
result Demonstrates global and exterior stability of Minkowski spacetime.
We quantify forgetting in post-training models, distinguishing mass and drift.
problem Understanding and preventing forgetting in post-training generative models.
method Developed theoretical results under a two-mode mixture abstraction, formalizing mass and drift forgetting.
result Forgetting can be precisely quantified based on divergence direction, geometric overlap, and training regime.
Unique solutions found for wave-like decaying null infinity equations.
problem Wave-like decaying null infinity equations with spherically symmetric Einstein-scalar-field.
method Local and global unique solutions for small initial data.
result Sharp decaying condition for unique solutions.
New findings on flatness of certain metrics with fast decay.
problem Rigidity of positive mass theorem under fast metric decay.
method Considered metrics with nonnegative scalar curvature and rapid decay at infinity.
result Any such metric is necessarily flat in dimensions 4 and higher if decay rate exceeds Schwarzschild metric.
Study on curvature decay in steady Ricci solitons, proving dichotomy.
problem Curvature decay in steady Ricci solitons.
method Established a dichotomy for curvature decay in specific types of solitons.
result Proved a dichotomy on curvature decay for certain steady Ricci solitons.
Study on decay rates of higher derivatives for nonlinear Dirac equations.
problem Estimating decay rates of higher derivatives of solutions to nonlinear Dirac equations.
method Similar to Li and Zang's method, focusing on 'good' spin null form.
result Obtained decay rates of higher derivatives of solutions.
Cautious Weight Decay modifies weight decay for better optimization.
problem Improving optimization in deep learning models.
method Applies weight decay selectively based on parameter sign alignment.
result Consistently improves model performance across various tasks and scales.
New proof shows certain solitons must be symmetric if curvature decays linearly.
problem Understanding noncompact steady Ricci solitons with specific curvature properties.
method Proved that nonnegative curvature operator and linear curvature decay imply rotational symmetry.
result Noncompact κ-noncollapsed steady Ricci solitons with nonnegative curvature operator and linear curvature decay must be rotationally symmetric.
Study on scalar curvature decay on non-compact manifolds linked at infinity.
problem Understanding scalar curvature decay on non-compact manifolds with topological linking at infinity.
method Analyzing polynomial decay, developing obstruction theory, using μ--bubble exhaustions, and index theory. result Topological linking at infinity forces polynomial decay of scalar curvature on manifolds of weakly bounded geometry.
Study on massless Vlasov equation on Reissner-Nordström spacetimes, showing decay rates and non-decay phenomena.
problem Analyzing decay and non-decay rates of solutions to the massless Vlasov equation on Reissner-Nordström spacetimes.
method Quantitative analysis of geodesic flow and comparison to wave equation instability results.
result Exponential decay rates in subextremal cases and polynomial rates in extremal cases, with non-decay of transversal derivatives in extremal cases.
Introduces gradient decay in Softmax for better generalization.
problem Improving generalization performance in neural networks.
method Gradient decay hyperparameter in Softmax for varying gradient rates based on probability.
result Gradient decay rate affects generalization performance and can be tuned for better optimization.
Study examines wave equation decay and Strichartz estimates on conic manifolds.
problem Analyzing wave equation behavior on conic spaces with critical electromagnetic potentials.
method Established decay and Strichartz estimates through localized spectral measure construction.
result Extended and improved previous results on wave equation behavior with critical potentials.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
Three mechanisms of weight decay found for different optimizers and architectures.
problem Understanding the regularization effect of weight decay in neural networks.
method Empirical investigation of weight decay for SGD, Adam, and K-FAC with various network architectures.
result Identified three distinct mechanisms of weight decay effect.
Random walks on hyperbolic spaces show linear progress with exponential decay.
problem Understanding progress and decay in random walks on hyperbolic spaces.
method Analyzing random walks on separable, geodesic hyperbolic metric spaces with specific step distributions.
result Exponential decay in progress is extended to non-acylindrical actions.
Study shows uniform decay rate for singular mean curvature flows.
problem Understanding singularities in mean curvature flows.
method Rescaled flow analysis near compact singularities.
result Uniform decay order bound for the rescaled flow.
Polynomial decay of correlations shown for curved surfaces.
problem Analyzing geodesic flows on curved surfaces.
method Proving polynomial decay of correlations for geodesic flows on nonpositively curved surfaces.
result Polynomial decay of correlations for geodesic flows on nonpositively curved surfaces.
Introduces gradient normalization and decay for deep learning.
problem Improving convergence time in deep neural networks.
method Gradient normalization and decay with respect to depth.
result Improvements in convergence time on image classification and natural language processing tasks.