Proposes a method to sample from flat basins of posterior distributions in Bayesian deep learning.
problem Sampling from multi-modal posterior distributions leads to overfitting due to trapping in bad modes.
method Introduces an auxiliary guiding variable to bias MCMC sampling towards flat basins of the energy landscape.
result The method converges faster and outperforms existing methods in sampling from flat basins of the posterior.
This research explains why SGD generalizes better than ADAM in deep learning.
problem Understanding the generalization gap between SGD and ADAM in deep learning.
method Analyzing local convergence behaviors through Levy-driven stochastic differential equations (SDEs).
result SGD is more locally unstable and better escapes from sharp minima to flatter ones, leading to better generalization.
Neural networks' optimization dynamics are confined to a single basin despite connected basins in the loss landscape.
problem Neural networks' optimization dynamics are confined to a single basin despite connected basins in the loss landscape.
method Identifying entropic barriers arising from the interplay between curvature variations along low-loss paths and noise in optimization dynamics.
result Curvature-induced entropic forces bias noisy dynamics back toward the endpoints, explaining the confinement and connectivity of solutions.
Quantization-aware training can recover accuracy lost by post-training quantization.
problem Post-training quantization (PTQ) can fail sharply at aggressive bitwidths.
method A unified geometric framework that explains PTQ failure and QAT recovery.
result QAT has a useful bias that steers iterates back into the basin.
EDLP samples flat modes in discrete spaces using entropy.
problem Sampling flat modes in discrete spaces is challenging.
method EDLP uses a continuous auxiliary variable and local entropy to guide sampling.
result EDLP consistently outperforms traditional methods in various tasks.
New method detects metastable basins in high dimensions using trajectory sampling.
problem Identifying distinct basins in high-dimensional Markov processes.
method Discriminative approach based on marginal trajectory distribution comparison.
result Bayes-optimal classifier achieves high accuracy distinguishing between basins.
fSGLD optimizes deep learning by favoring flat regions in the loss landscape.
problem Understanding and improving the behavior and generalization of deep learning algorithms.
method Flatness-Aware Stochastic Gradient Langevin Dynamics (fSGLD) that biases learning towards flat basins.
result fSGLD targets a flatness-biased Gibbs distribution with explicit excess risk guarantees.
New method improves ensemble quality by exploring the pre-train basin more effectively.
problem Limited diversity in ensembles trained from a single pre-trained checkpoint.
method Proposed StarSSE modification of Snapshot Ensembles for transfer learning.
result Stronger ensembles and uniform model soups achieved.
SLT reveals how grokking occurs via basin selection in training.
problem Understanding grokking in machine learning models.
method Singular Learning Theory (SLT) to analyze the loss landscape and local learning coefficient (LLC).
result LLC ranks basins by statistical preference, leading to grokking.
Proves properties of neural network basins of attraction and their expressiveness.
problem Characterize the properties of basins of attraction in neural networks.
method Analyzes width-bounded neural networks, proving properties of basins of attraction.
result Boundedness and path-connectedness of basins of attraction under certain conditions.
Wide networks are often believed to have a nice optimization landscape, but what rigorous results can we prove? To understand the benefit of width, it is important to identify the difference between wide and narrow networks. In this work, we prove that from narrow to wide networks, there is a phase transition from havi…
Straight lines are a basin of attraction for the elastic flow at least to level 1.9615π.
problem Understanding the basin of attraction for the free boundary free elastic flow.
method Steepest descent gradient flow for elastic energy, numerical evidence.
result Straight lines have a basin of attraction at least to level 1.9615π.
Study characterizes spike deconvolution basin for noisy data.
problem Recover spike locations from noisy convolution with PSF across multiple snapshots.
method Variable-projection formulation, explicit basin of convexity characterization, local convergence guarantees.
result Consistent estimator within basin of convexity under stochastic noise, complementary error bound under adversarial noise.
Proposes a probabilistic model to improve hydrology predictions and trust.
problem Noisy or missing basin characteristics impact streamflow prediction.
method Probabilistic inverse modeling framework to reconstruct basin characteristics.
result 6% improvement in R2 for streamflow prediction, 17% reduction in uncertainty. Study of separatrix configurations in holomorphic flows with real time.
problem Characterize separatrices in holomorphic flows with real-valued time.
method Establish continuity of transit times, classify path components, prove blow-up scenarios.
result Separatrices of different types of equilibria exhibit blow-up in finite time.
Flooding is a destructive and dangerous hazard and climate change appears to be increasing the frequency of catastrophic flooding events around the world. Physics-based flood models are costly to calibrate and are rarely generalizable across different river basins, as model outputs are sensitive to site-specific parame…
Quantification of the stationary points and the associated basins of attraction of neural network loss surfaces is an important step towards a better understanding of neural network loss surfaces at large. This work proposes a novel method to visualise basins of attraction together with the associated stationary points…
Study global geometry of dynamical systems with entire vector fields.
problem Understanding the global structure of equilibria and their basins.
method Step-by-step analysis of basins of centers, nodes, and foci; introduction of global elliptic sectors.
result Characterization of heteroclinic regions connecting equilibria.
Regional rainfall-runoff modeling is an old but still mostly out-standing problem in Hydrological Sciences. The problem currently is that traditional hydrological models degrade significantly in performance when calibrated for multiple basins together instead of for a single basin alone. In this paper, we propose a nov…
We study codimension one foliations with singularities defined locally by Bott-Morse functions on closed oriented manifolds. We carry to this setting the classical concepts of holonomy of invariant sets and stability, and prove a stability theorem in the spirit of the local stability theorem of Reeb. This yields, among…
New method finds basins of attraction without needing system models.
problem Determining basins of attraction (BoA) for nonlinear systems without prior knowledge.
method Hybrid Active Learning (HAL) method combining AST, AL, and DBS.
result Efficiently finds and labels boundary of BoA without model knowledge.
Researchers use Gaussian processes with non-stationary kernels to model precipitation patterns in the Upper Indus Basin.
problem Uncertainty in precipitation patterns in the Upper Indus Basin, Himalayas.
method Proposes Gaussian processes with structured non-stationary kernels to model precipitation patterns, accounting for spatial variation with a latent Gaussian process.
result The proposed model adapts to varying precipitation patterns across distinct topography and outperforms stationary models in ablation experiments.
In several experimental reports on nonconvex optimization problems in machine learning, stochastic gradient descent (SGD) was observed to prefer minimizers with flat basins in comparison to more deterministic methods, yet there is very little rigorous understanding of this phenomenon. In fact, the lack of such work has…
SGD transitions between maxima and minima with varying time scales.
problem Understanding SGD's behavior near critical points in noisy landscapes.
method Analyzing SGD convergence and escape dynamics in 1D landscapes with infinite- and finite-variance noise.
result SGD reliably moves to the basin's minimum unless close to a local maximum, where it can linger.
Study SRB measures for Anosov actions on manifolds.
problem Characterize SRB measures for Anosov actions.
method Use Ruelle-Taylor resonances and properties of Sinai-Ruelle-Bowen measures.
result SRB measures have properties like smooth disintegrations, positive basins, and are unique under certain conditions.
Multifractality in time series arises from temporal correlations, not just fat tails.
problem Understanding the origin of multifractality in time series data.
method Mathematical arguments and numerical simulations using MFDFA approach.
result Genuine multifractality requires temporal correlations, not just fat tails.
Biological neurons learn tensor decompositions of higher-order correlations using nonlinear Hebbian plasticity.
problem Learning higher-order correlations in biological neurons.
method Introduce and study generalized nonlinear Hebbian learning rules.
result Neurons can learn tensor eigenvectors of higher-order input correlation tensors.
WASH trains ensembles with shuffled weights to improve accuracy and reduce communication.
problem Training ensembles for weight averaging leads to models converging to different loss basins.
method WASH randomly shuffles a small percentage of weights during training to keep models within the same basin.
result WASH achieves state-of-the-art image classification accuracy with lower communication costs.
Recently, a general method for analyzing the statistical accuracy of the EM algorithm has been developed and applied to some simple latent variable models [Balakrishnan et al. 2016]. In that method, the basin of attraction for valid initialization is required to be a ball around the truth. Using Stein's Lemma, we exten…
Paper proposes a new method for population-wise matching of sulcal graphs.
problem Challenges in matching cortical fold variations across individuals.
method Population-wise multi-graph matching of sulcal graphs.
result Effectiveness of multi-graph matching in obtaining consistent labeling of sulcal basins.
Framework analyzes neural network dynamics for better understanding and optimization.
problem Understanding the fundamental mechanisms of deep neural networks.
method Dynamical systems theory, transformation units, attraction basins.
result Different transformation modes lead to distinct learning phases and network performance.
DeepWeightFlow generates diverse neural network weights efficiently.
problem Generating complete neural network weights efficiently and accurately.
method Flow Matching in weight space with Git Re-Basin and TransFusion.
result DeepWeightFlow generates high-accuracy neural networks without fine-tuning.
We introduce a mathematical model on the dynamics of demand and supply incorporating collectability and saturation factors. Our analysis shows that when the fluctuation of the determinants of demand and supply is strong enough, there is chaos in the demand-supply dynamics. Our numerical simulation shows that such a cha…
Study of holomorphic correspondences combining entire maps and Fuchsian groups.
problem Understanding dynamics of entire maps and their interactions with Fuchsian groups.
method Systematic study of (∞:∞) holomorphic correspondences arising from conformal combinations of transcendental entire maps and Fuchsian groups. result The resulting correspondence is the composition of a Möbius involution and the deleted covering correspondence of a meromorphic function with a simple pole.
The basin of infinity of a polynomial map $f : {\bf C} \arrow {\bf C}$ carries a natural foliation and a flat metric with singularities, making it into a metrized Riemann surface X(f). As f diverges in the moduli space of polynomials, the surface X(f) collapses along its foliation to yield a metrized simplicial t…
We consider compressed sensing formulated as a minimization problem of nonconvex sparse penalties, Smoothly Clipped Absolute deviation (SCAD) and Minimax Concave Penalty (MCP). The nonconvexity of these penalties is controlled by nonconvexity parameters, and L1 penalty is contained as a limit with respect to these para…
An image pattern can be represented by a probability distribution whose density is concentrated on different low-dimensional subspaces in the high-dimensional image space. Such probability densities have an astronomical number of local modes corresponding to typical pattern appearances. Related groups of modes can join…
HydroNets use river structure to improve hydrologic predictions.
problem Scalable and accurate hydrologic models are needed for climate change impacts.
method HydroNets are deep neural networks that incorporate river network structure.
result HydroNets improve predictions with fewer data, especially at longer horizons.
The paper investigates what enables successful transfer learning and separates feature reuse from data statistics.
problem Understanding what enables successful transfer learning and identifying the responsible parts of the network.
method Analyzes transfer learning on block-shuffled images to distinguish feature reuse from data statistics.
result Some benefit of transfer learning comes from learning low-level statistics of data, not just feature reuse.
Deep learning, in the form of artificial neural networks, has achieved remarkable practical success in recent years, for a variety of difficult machine learning applications. However, a theoretical explanation for this remains a major open problem, since training neural networks involves optimizing a highly non-convex …
RAMBO optimizes multi-regime problems by discovering and modeling distinct energy basins.
problem Multi-regime problems in molecular conformation and drug discovery.
method Dirichlet Process Mixture of Gaussian Processes with adaptive hyperparameters and concentration parameters.
result Consistent improvements over state-of-the-art on multi-regime objectives.
A new associative memory uses Sinkhorn divergence for efficient pattern retrieval.
problem Efficiently retrieving patterns from large datasets of weighted point clouds.
method Derived retrieval dynamics as a SHK gradient flow, discretized for a deterministic algorithm.
result Proved basin invariance, geometric convergence, and robust recovery from perturbations.
LSTM models with DI enhance streamflow forecasts across diverse regions.
problem Challenges in integrating varied discharge measurements for accurate streamflow forecasts.
method Flexible data integration (DI) using LSTM models with CNN units for lagged inputs.
result DI significantly improved streamflow forecast performance, reaching record efficiency coefficients.
The paper explains how continuous language models can produce discrete, interpretable meanings.
problem Semantic collapse in continuous systems of large language models.
method Formalizing large language models as Continuous State Machines (CSMs) and analyzing the associated transfer operator.
result The leading eigenfunctions of the transfer operator induce a finite number of invariant meaning basins, explaining how continuous computation can produce discrete, interpretable semantics.
Recent work has demonstrated the effectiveness of gradient descent for directly recovering the factors of low-rank matrices from random linear measurements in a globally convergent manner when initialized properly. However, the performance of existing algorithms is highly sensitive in the presence of outliers that may …
The study revisits Hopfield's associative memory model and calculates its capacity for two specific pattern basins.
problem Determining the capacity of a Hebbian-Hopfield network for storing binary patterns.
method Using fully lifted random duality theory and numerical analysis, the study calculates the capacity for two specific pattern basins.
result Explicit characterizations of the capacity for the AGS and NLT pattern basins, with remarkable fast lifting convergence.
New optimization method helps models generalize better after achieving near-perfect training performance.
problem Models can achieve near-perfect training performance but fail to generalize well to unseen examples.
method GROKtimizer combines rapid convergence to interpolation with post-interpolation norm minimization using Critically Damped Momentum.
result GROKtimizer provides a quadratic speedup over classical gradient descent, offering a natural solution for selecting low-norm interpolating solutions.
Density mode clustering is a nonparametric clustering method. The clusters are the basins of attraction of the modes of a density estimator. We study the risk of mode-based clustering. We show that the clustering risk over the cluster cores --- the regions where the density is high --- is very small even in high dimens…