Smoothed fitness landscape improves protein optimization.
problem Infeasibility of combinatorially large protein sequence space.
method Formulate protein fitness as a graph signal, smooth using Tikunov regularization, and optimize with Gibbs sampling.
result 2.5 fold fitness improvement over training set.
Riemannian geometry improves protein dynamics analysis.
problem Efficient analysis of protein dynamics data in non-linear spaces.
method Developed a local approximation technique for geodesics and a smooth manifold of protein conformations.
result Geodesics approximate molecular dynamics trajectories and provide realistic summary statistics.
EP approximates Bayesian posterior distributions using gradient descent.
problem Bayesian inference's uncomputable posterior distributions.
method Relates EP to gradient descent on a smoothed energy landscape.
result EP is equivalent to gradient descent on a smoothed energy landscape.
New sampler tackles complex discrete energy landscapes efficiently.
problem Stagnation in gradient-based discrete samplers for non-convex settings.
method DREXEL sampler with Replica Exchange and Adjusted Metropolis.
result Proves samplers satisfy detailed balance and converge to target distribution.
This paper explores machine learning landscapes using molecular energy analogy.
problem Understanding the solution space and nature of predictions in machine learning.
method Analogy with molecular potential energy landscapes to explore machine learning landscapes.
result Emergent properties of machine learning landscapes can be related to molecular structure, thermodynamics, and kinetics.
Neural networks' energy landscape is surprisingly flat, suggesting minimal structural changes between minima.
problem Understanding the structure of neural network energy landscapes.
method Constructing continuous paths between minima of recent neural network architectures on CIFAR10 and CIFAR100.
result Paths between minima are essentially flat in both training and test landscapes, implying minimal structural changes.
Study shows neural networks have less complexity than expected.
problem Understanding the complexity of neural networks.
method Explored the energy landscape of a simple neural network.
result Neural networks have less complexity than expected, leading to better generalization.
Study energy landscapes in glass models, focusing on Gaussian and spiked-tensor functions.
problem Characterize statistical properties and phase transitions of high-dimensional energy landscapes.
method Developed a Kac-Rice method framework to compute landscape complexity and analyze phase transitions rigorously.
result Characterized the ruggedness and arrangements of local minima in energy landscapes.
Unified framework for sampling and approximating high-dimensional energy landscapes.
problem Sampling and approximating complex energy landscapes in physical systems with constraints and energy barriers.
method Formulates a minimax optimization problem that jointly adapts surrogate approximation and adaptive sampling.
result Demonstrates effectiveness in biomolecular systems with up to 30 collective variables.
New theory shows predictive coding makes learning landscape easier to navigate.
problem Understanding the impact of predictive coding's inference procedure on learning efficiency.
method Analyzed the geometry of the energy landscape of deep linear networks, proving many non-strict saddles become strict in the equilibrated energy.
result All highly degenerate (non-strict) saddles of the loss become strict in the equilibrated energy, suggesting a more robust learning landscape.
Entropy-SGD improves deep learning by favoring flat minima in the energy landscape.
problem Training deep neural networks to avoid poorly-generalizing sharp minima.
method Constructs a local-entropy-based objective function to bias SGD towards flat regions.
result Entropy-SGD leads to improved generalization over SGD, as shown by experiments and theoretical underpinnings.
In many statistical learning problems, the target functions to be optimized are highly non-convex in various model spaces and thus are difficult to analyze. In this paper, we compute \emph{Energy Landscape Maps} (ELMs) which characterize and visualize an energy function with a tree structure, in which each leaf node re…
Study on the optimization landscape of half-rectified networks without simplifying assumptions.
problem Understanding the optimization landscape of deep neural networks, focusing on half-rectified networks.
method Theoretical analysis and empirical study of gradient descent on half-rectified networks.
result Proves that half-rectified single layer networks are asymptotically connected and provides bounds on the interplay between data distribution and model over-parametrization.
SGD with machine learning noise converges to global minimum exponentially fast.
problem Optimizing machine learning models with stochastic gradient descent.
method Analysis of SGD with machine learning noise, focusing on energy landscapes and gradient noise.
result SGD converges to the global minimum exponentially fast under certain conditions.
Autoencoders discover and accelerate molecular dynamics simulations.
problem Efficient sampling of macromolecular folding landscapes with high free energy barriers.
method Employing auto-associative artificial neural networks to learn nonlinear collective variables (CVs) that are explicit and differentiable functions of atomic coordinates.
result Substantial speedups in exploration of configurational space and discovery of data-driven CVs.
We reveal connections between RBMs and Bosons, explaining symmetry breaking in their energy landscapes.
problem Understanding the relationships among different deep generative models and their learning mechanisms.
method Introducing a reciprocal space formulation to RBMs, revealing connections to diffusion processes and Bosons.
result Symmetry breaking in RBM energy landscapes is characterized by singular values and weight matrix eigenvectors.
Experimental fractal landscape dynamics observed in emulsions.
problem Understanding anomalous motions in soft glassy materials.
method Quantitative analysis of oil droplet trajectories in dense emulsions.
result Experimental fractal geometry matches computational model of soft glassy dynamics.
LSAM optimizes deep learning training with improved efficiency.
problem Inefficiency in distributed large-batch training with Sharpness-Aware Minimization (SAM).
method Integrates SAM's adversarial steps with an asynchronous distributed sampling strategy.
result Higher final accuracy compared to data-parallel SAM.
SGD vs quasi-Newton optimization in neural networks: different landscapes, different generalizability.
problem Understanding neural network optimization and generalizability.
method Comparison of stochastic gradient descent (SGD) and quasi-Newton optimization methods using computational tools.
result SGD solutions are separated by lower barriers than quasi-Newton solutions, but quasi-Newton solutions are deeper and more isolated.
Improved Langevin Monte Carlo reduces energy barriers for faster optimization.
problem Optimizing functions with high energy barriers.
method Proposes a modified landscape for Langevin Monte Carlo.
result Polynomial dependence on energy barrier in Log-Sobolev constant.
Sparse neural networks training is difficult due to optimization failures and energy landscape issues.
problem Training sparse neural networks leads to suboptimal solutions and optimization failures.
method Investigated optimization dynamics and energy landscape in sparse neural networks.
result Sparse neural networks have a linear path with a monotonically decreasing objective from initialization to a good solution, but not from a bad solution.
Deep neural networks are optimizable due to their multilayered structure.
problem Understanding why deep neural networks are easily optimizable despite their non-convex loss functions.
method Analysis of a spin glass model of deep neural networks using random matrix theory and algebraic geometry.
result The multilayered structure of deep neural networks leads to fewer stationary points, more clustered minima, and less severe tradeoffs between depth and width of minima.
IDM learns clustered representations for complex FELs.
problem Learning meaningful representations for complex free energy landscapes.
method Information Distilling of Metastability (IDM) method.
result IDM achieves physically meaningful representations of FELs.
Model explains surprising properties of neural network training landscapes.
problem Understanding the geometry of neural network loss landscapes.
method Developed a simple theoretical model of gradients and Hessians.
result Unified model accounts for 4 surprising properties of neural loss landscapes.
EWFM trains continuous flows with only energy evaluations, improving sample quality with fewer computations.
problem Efficiently sampling from complex, high-dimensional Boltzmann distributions using only energy evaluations.
method Energy-Weighted Flow Matching (EWFM) using importance sampling and iterative/annealed training.
result Improved sample quality with up to 3 orders of magnitude fewer energy evaluations compared to existing methods.
iDEM generates samples from Boltzmann densities without data.
problem Generating statistically independent samples from unnormalized distributions.
method Iterative algorithm using energy and gradient for diffusion-based sampler training.
result iDEM achieves state-of-the-art performance and trains faster than existing methods.
Proposes a method to sample from flat basins of posterior distributions in Bayesian deep learning.
problem Sampling from multi-modal posterior distributions leads to overfitting due to trapping in bad modes.
method Introduces an auxiliary guiding variable to bias MCMC sampling towards flat basins of the energy landscape.
result The method converges faster and outperforms existing methods in sampling from flat basins of the posterior.
Black holes offer insights into machine learning's loss landscapes.
problem Understanding the loss landscape in machine learning.
method Comparing machine learning loss landscapes to black hole entropy.
result Black holes provide an infinite family of potential landscapes with known minima.
Trained SPENs with efficient search in reward function for structured prediction.
problem Expensive ground-truth labeling in structured output prediction.
method Efficient truncated randomized search in reward function for training SPENs.
result Local improvements and effective supervision for SPENs without labeled data.
Deep neural networks undergo hierarchical free-energy landscape transitions with increasing data size.
problem Understanding the design space and dynamics of deep neural networks.
method Statistical mechanical approach based on replica method.
result Hierarchical free-energy landscape transitions with ultrametricity, leading to simpler configurations in deeper layers.
Study of phase separation and geometry on a closed elastic curve, including dynamics and free energy minimization.
problem Free energy and dynamics of a closed elastic filament coupled to a scalar concentration field.
method Analytical and numerical simulations of coupled Willmore flow and Cahn--Hilliard gradient flow on differential geometry.
result Qualitative changes in free energy landscape due to closure constraint, leading to metastable and stable multi-domain morphologies.
SmoothDARTS stabilizes DARTS-based architecture search by smoothing loss landscapes.
problem DARTS-based NAS methods suffer from instability, leading to deteriorating architectures.
method SmoothDARTS (SDARTS) uses perturbation-based regularization to smooth the loss landscape.
result SmoothDARTS improves the generalizability and performance of DARTS-based methods.
Quantum annealing outperforms classical in solving non-convex optimization problems crucial for machine learning.
problem Solving non-convex optimization problems efficiently.
method Designing a classical energy function and adding a quantum transverse field to facilitate tunneling.
result Quantum annealing converges efficiently to optimal solutions in a wide class of non-convex problems, unlike classical thermal annealing.
DMs emerge from DenseAMs, transitioning from memorization to generalization.
problem Hindered memory retrieval in DenseAMs due to spurious states.
method Examined diffusion models through the lens of DenseAMs, focusing on their generative process.
result Identified a critical phase in DMs transitioning from memorization to generalization.
ELS framework improves safety alignment by dynamically steering LLMs towards helpful responses.
problem Over-Refusal in Aligned Large Language Models
method Fine-tuning free framework using an Energy-Based Model (EBM) to dynamically steer LLMs during inference.
result Extensive experiments show a significant reduction in false refusals (from 57.3% to 82.6%) while maintaining safety performance.
GMVAE improves clustering in molecular simulations data.
problem Clustering metastable states in multi-basin free-energy landscapes.
method Gaussian mixture variational autoencoder (GMVAE) for dimensionality reduction and clustering.
result Enhanced clustering of metastable states compared to standard VAEs.
Current state-of-the-art discrete optimization methods struggle behind when it comes to challenging contrast-enhancing discrete energies (i.e., favoring different labels for neighboring variables). This work suggests a multiscale approach for these challenging problems. Deriving an algebraic representation allows us to…
We study the pull-back of the 2-parameter family of quotient elastic metrics introduced in Mio-Srivastava-Joshi on the space of arc-length parameterized loops. This point of view has the advantage of concentrating on the manifold of arc-length parameterized curves, which is a very natural manifold when the analysis of …
Exploring loss landscapes of XOR networks reveals complex structures.
problem Understanding the optimization landscape of XOR networks.
method Using molecular science optimization tools, analyzing the number and types of stationary points.
result The landscape of XOR networks becomes more convex with increased regularisation, embedding smaller networks in larger ones.
Study of neural network training using mean-field Langevin dynamics and energy landscapes.
problem Understanding convergence of stochastic gradient algorithms for non-convex learning tasks.
method Leverage infinite-dimensional convexification, Mean-Field Langevin Dynamics, and energy functional.
result Gradient flow converges to a unique minimizer in 2-Wasserstein metric, with exponential convergence under regularisation.
Study identifies new stable climate states in climate model.
problem Understanding multistability and transitions in climate models.
method Combination of quasipotential theory and manifold learning.
result Discovery of a third stable climate state not previously known.
Discovering quasipotential equations from data using machine learning.
problem Understanding escape mechanisms from metastable states in nonlinear systems.
method Combining neural networks and sparse regression to symbolically reconstruct quasipotential equations.
result Model-unbiased analytical forms of quasipotential discovered directly from data.
Quantified limits of nuclear stability beyond drip lines.
problem Predicting nuclear stability beyond known isotopes.
method Microscopic nuclear mass models, Bayesian methodology, Gaussian processes.
result Quantified predictions of one- and two-nucleon separation energies.
Deep neural networks and glassy systems share dynamics but differ in landscape properties.
problem Comparing training dynamics of DNNs and glassy systems.
method Statistical physics methods applied to DNN training.
result DNN dynamics slow down due to many flat directions, diffusing at the loss minimum.
Bayesian inference learns free energy landscapes from experimental data.
problem Characterize the free energy landscape of classical many-body systems from experimental data.
method Combines non-parametric Bayesian inference with physically-motivated constraints to automate the construction of approximate free energy functionals.
result Inference algorithms yield a probability distribution over free energy functionals, leading to highly accurate analytic expressions.
Bayesian analysis predicts properties of proton-emitting nuclei beyond the proton drip line.
problem Predicting properties of unstable nuclei in the proton-rich region.
method Bayesian Gaussian processes and mass models corrected with statistical emulators.
result Quantified predictions for separation energies and probabilities of proton emission.
Develops hyperparameter transfer methods for Dense Associative Memories.
problem Challenges in transferring hyperparameters for DenseAMs due to unique architecture and activation functions.
method Derives explicit prescriptions for hyperparameter transfer from small to large models.
result Excellent agreement between theoretical and empirical results.
New method simplifies optimization landscapes by transforming saddle points.
problem Saddle points hinder non-convex optimization in machine learning.
method Variable elimination algorithms, like VarPro, are compared to reveal geometric insights.
result Variable elimination reshapes critical point structure, creating local maxima from saddle points.