Proposes learning invariances in neural networks using a weight-space approach.
problem Learning invariances from data in neural networks remains an open problem.
method Minimizes a lower bound on the marginal likelihood in weight space.
result Results in higher performing models with naturally learned invariances.
New method approximates curvature from symmetries in deep networks.
problem Hard to approximate curvature in large deep networks.
method Analytically averaging over group actions that leave the loss invariant to construct structured Hessian approximations.
result Structured Hessian approximations from single gradients can be estimated, stored, and inverted.
Variational inference struggles with weight symmetries in neural networks, leading to biased posteriors.
problem Weight space symmetries in neural networks cause multimodal posteriors, challenging variational inference.
method Developed a symmetrization mechanism to create permutation invariant variational posteriors.
result Symmetrized variational posteriors have a better fit to the true posterior and improved predictive performance.
Empirical study shows removing neural parameter symmetries impacts model performance.
problem Understanding the impact of neural parameter symmetries on model performance.
method Developed two methods to reduce parameter space symmetries in neural networks.
result Removing parameter symmetries can lead to faster and more effective Bayesian neural network training.
DeepWeightFlow generates diverse neural network weights efficiently.
problem Generating complete neural network weights efficiently and accurately.
method Flow Matching in weight space with Git Re-Basin and TransFusion.
result DeepWeightFlow generates high-accuracy neural networks without fine-tuning.
This work explores how overparametrization and priors affect Bayesian neural network posteriors.
problem Symmetries, non-identifiabilities, and weight-space priors fragment and inflate BNN posteriors.
method We study the interplay between overparametrization and priors in BNN posteriors, deriving key phenomena and validating through experiments.
result Overparametrization induces structured, prior-aligned weight posterior distributions.
Functional input neural networks approximate continuous functions on weighted spaces.
problem Approximating continuous functions on infinite-dimensional weighted spaces.
method Additive family mapping, non-linear activation, linear readouts, Stone-Weierstrass theorem.
result Global universal approximation of continuous functions on weighted spaces.
Hypernetworks are neural networks that generate weights for another neural network. We formulate the hypernetwork training objective as a compromise between accuracy and diversity, where the diversity takes into account trivial symmetry transformations of the target network. We explain how this simple formulation gener…
This paper explores Bayesian Neural Network posteriors, uncovering symmetries and their impact.
problem Understanding the complex posterior distribution of deep Bayesian Neural Networks.
method Investigates optimal approaches for approximating posteriors, analyzes modes, and explores visualizations.
result Uncovered weight-space symmetries and their impact on the posterior, particularly scaling symmetries.
MLDS dataset reveals hidden model behavior via weight-space analysis.
problem Neural networks' opacity makes them hard to evaluate.
method Presented MLDS dataset of trained neural networks.
result Weight-space analysis reveals meaningful divergence with small changes in training data.
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima. In a network of d−1 hidden layers with nk neurons in layers $k = 1, \ldots, …
Improved loss functions adapt to weight-space anisotropy, outperforming isotropic counterparts.
problem Adapting to the anisotropic nature of deep weight spaces for better performance.
method Refined local entropic loss functions restricted to a subset of weights, exploiting anisotropy.
result Partial local entropies outperform isotropic counterparts on image classification tasks.
Data-driven model shows deep learning weights behave like a liquid.
problem Understanding the structure of deep neural network optimization landscapes.
method Statistical mechanics framework to model high-dimensional weight spaces.
result Deep networks' weight spaces are well-connected, not hierarchical, unlike shallow networks.
The paper investigates how neural network weights evolve to monitor training progress.
problem Monitoring the training progress of neural networks in a cost-effective manner.
method Investigates the evolution of neural network weights in weight space.
result DNN models evolve on unique, smooth trajectories in weight space that can be used to track training progress.
Paper presents a method to summarize HMC samples for neural networks, providing meaningful uncertainty estimates.
problem Lack of interpretable summary statistics for HMC samples in neural networks due to permutation symmetry.
method Introducing a transpositions metric to quantify permutations and using rebasin method to summarize HMC samples.
result Compact representation of HMC samples provides meaningful uncertainty estimates for each weight in a neural network.
Paper solves CR positive mass and Yamabe problems on weighted spaces.
problem CR positive mass and Yamabe problems on weighted spaces.
method Analyzes sub-Laplacian on Folland-Stein spaces.
result CR positive mass and Yamabe problems resolved.
Study of quantum spaces on Kähler manifolds with T-symmetry converging to a mixed polarization.
problem Quantum spaces on Kähler manifolds with T-symmetry and their convergence.
method Construction of a one-parameter family of Kähler structures and study of quantum spaces.
result Quantum spaces corresponding to different polarizations converge to a mixed polarization as the parameter goes to infinity.
The lottery ticket hypothesis finds multiple winning sub-networks in neural networks.
problem Finding a single winning sub-network in neural networks.
method Analyzing neural networks trained in isolation and on different tasks.
result Neural networks contain multiple sub-networks that match the accuracy of the original network, not just one.
Study on Bayesian transformers finds issues with weight-space inference and prior specification.
problem Challenges in obtaining meaningful uncertainty estimates for transformer models.
method Proposed a novel method based on implicit reparameterization of the Dirichlet distribution for variational inference on attention weights.
result Proposed method performs competitively with baselines in estimating predictive uncertainty.
LoRAs enable efficient adaptation of large models; this paper explores processing LoRA weights with machine learning.
problem Efficient processing of low-rank weight decompositions in large finetuned models.
method Developed symmetry-aware invariant and equivariant LoL models to process LoRA weights.
result LoL models can predict CLIP scores, finetuning data attributes, and accuracy on downstream tasks.
The paper studies gravitational instantons with flat limits and finds elliptic regularity estimates.
problem Analyzing gravitational instantons with flat limits using elliptic analysis.
method Establishing elliptic regularity estimates and showing uniform constants for a family of metrics.
result The Laplacian is Fredholm and an isomorphism between specific weighted spaces.
Extends Feller theory to non-locally compact spaces for stochastic equations.
problem Stochastic partial differential equations and fractional processes.
method Extended Feller processes and proofs of folklore results.
result No condition of generalized Feller semigroups can be dropped.
The study analyzes local minima in ReLU networks and finds low probability of bad local minima.
problem Understanding the existence and probability of local minima in ReLU networks.
method Theoretical analysis combined with linear programming and experiments on MNIST and CIFAR-10 datasets.
result No bad differentiable local minima found almost everywhere in weight space.
New method adapts neural networks without losing prior knowledge.
problem Understanding and enabling flexible adaptation of neural networks.
method Differential geometry framework, functionally invariant paths (FIP).
result Achieves comparable state-of-the-art performance on continual learning and sparsification tasks.
Study connects Gaussian processes and regularization for sequence-function mappings.
problem Understanding and interpreting sequence-function maps in biology.
method Relates Gaussian process priors, regularization, and gauge fixing in overparameterized weight space.
result Established the relationship between regularized regression and Gaussian processes in function space.
Study nonnegative solutions on Riemannian manifolds using fractional porous medium equation.
problem Analyzing solutions to fractional porous medium equation on noncompact Riemannian manifolds.
method Existence and smoothing estimates for weak solutions in L1 and weighted spaces. result Results hold for Euclidean and hyperbolic spaces, including larger data classes.
For a 2-periodic link L~ in the thickened annulus and its quotient link L, we exhibit a spectral sequence with E1≅AKh(L~)⊗F2F2[θ,θ−1]⇉E∞≅AKh(L)⊗F2F2[θ,θ−1]. This spectral sequence splits along qu…
Characterizes deep neural network weight space for adversarial attacks.
problem Poor performance of deep learning models in adversarial examples.
method Characterizes deep neural network solution space using two paradigms.
result Adversarial attacks are less successful against Associative Memory Models.
We propose regularizing the empirical loss for semi-supervised learning by acting on both the input (data) space, and the weight (parameter) space. We show that the two are not equivalent, and in fact are complementary, one affecting the minimality of the resulting representation, the other insensitivity to nuisance va…
Let g be a complex, simple Lie algebra with Cartan subalgebra h and Weyl group W. We construct a one-parameter family of flat connections D on h with values in any finite-dimensional h-module V and simple poles on the root hyperplanes. The corresponding monodromy representation of the braid group B of type g is a defor…
LGV boosts adversarial attacks by improving surrogate models.
problem Improving the transferability of black-box adversarial attacks.
method LGV uses a pretrained surrogate model and multiple weight sets from additional training epochs to generate an effective surrogate ensemble.
result LGV outperforms other test-time transformations by significant margins.
Paper studies solutions to a specific equation in conformal geometry with singular sets.
problem Singular solutions to a fully non-linear equation in conformal geometry.
method Uses a classical gluing method adapted to the fully non-linear setting.
result Shows the classical gluing method can be applied to the σ2--Yamabe equation. The paper studies the Heath-Jarrow-Morton-Musiela equation of the bond market. The equation is analyzed in weighted spaces of functions defined on [0,+∞). Sufficient conditions for local and global existence are obtained . For equation with the linear diffusion term the conditions for global existence are close …
LaLoRA prevents forgetting in LoRA fine-tuning.
problem Catastrophic forgetting in fine-tuned models.
method LaLoRA applies Laplace approximation to LoRA weights for regularization.
result Improved learning-forgetting trade-off with controllable regularization strength.
We propose Radial Bayesian Neural Networks (BNNs): a variational approximate posterior for BNNs which scales well to large models while maintaining a distribution over weight-space with full support. Other scalable Bayesian deep learning methods, like MC dropout or deep ensembles, have discrete support-they assign zero…
High order splitting schemes with complex timesteps are applied to Kolmogorov backward equations stemming from stochastic differential equations in Stratonovich form. In the setting of weighted spaces, the necessary analyticity of the split semigroups can be easily proved. A numerical example from interest rate theory,…
We give a topological interpretation of the space of L2-harmonic forms on finite-volume manifolds with sufficiently pinched negative curvature. We give examples showing that this interpretation fails if the curvature is not sufficiently pinched and that our result is sharp with respect to the pinching constants. The me…
Paper addresses variational inference issues in Bayesian neural networks.
problem Negative infinite ELBO for function-space priors in BNNs.
method Regularized KL divergence for well-defined function-space variational inference.
result Method provides competitive uncertainty estimates for BNNs.
Uniform elliptic theory for Dirac operators on orbifold resolutions.
problem Analyzing Dirac operators on orbifold resolutions.
method Viewing orbifolds as conically fibred singular spaces and resolving them by gluing asymptotically conical fibrations.
result Uniform index formula for Dirac operators on orbifold resolutions.
We introduce a quotient of the affine Temperley-Lieb category that encodes all weight-preserving linear maps between finite-dimensional sl(2)-representations. We study the diagrammatic idempotents that correspond to projections onto extremal weight spaces and find that they satisfy similar properties as Jones-Wenzl pro…
This paper is devoted to obtaining a wellposedness result for multidimensional BSDEs with possibly unbounded random time horizon and driven by a general martingale in a filtration only assumed to satisfy the usual hypotheses, i.e. the filtration may be stochastically discontinuous. We show that for stochastic Lipschitz…
The paper defines function spaces on manifolds with bounded or singular geometries.
problem Defining function spaces on manifolds with various geometries.
method Introduces and analyzes Sobolev, Besov, and Bessel potential spaces on uniformly regular and singular Riemannian manifolds.
result Demonstrates maximal regularity for a linear parabolic problem on singular manifolds.
In this paper, we investigate the behavior of ADM mass and Einstein-Hilbert functional under the Yamabe flow. Through studying the Yamabe flow by weighted spaces, we show that ADM mass and Einstein-Hilbert functional are well-defined and monotone non-increasing under the Yamabe flow on n-dimensional, n≥3, asymp…
Koopman mode analysis applied to neural networks for training optimization.
problem Optimizing neural network training, identifying issues, and speeding up learning.
method Koopman operator analysis of neural network dynamics.
result Spectral analysis of Koopman operator aids in determining network depth, initialization quality, and training termination.
The Dirac operator d+delta on the Hodge complex of a Riemannian manifold is regarded as an annihilation operator A. On a weighted space L_mu^2 Omega, [A,A*] acts as multiplication by a positive constant on excited states if and only if the logarithm of the measure density of mu satisfies a pair of equations. The equati…
CUTS removes corruption from models without clean data, improving utility and security.
problem Removing corruption from models without access to clean training data.
method CUTS uses a proxy set to amplify corruption and subtract it from model weights.
result CUTS recovers a large fraction of lost utility and nearly eliminates attacks with minimal damage.
The paper defines and analyzes abla-Sobolev spaces and operators on manifolds.
problem Defining and analyzing Sobolev spaces and differential operators on manifolds.
method Coordinate-free approach using connections, proving properties of abla-Sobolev spaces and operators. result Equivalent definitions of abla-Sobolev spaces and operators under certain conditions. Unified framework for Bayesian PDE-constrained inversion using physics-informed neural networks.
problem Incorporating prior distributions in function space into Bayesian PINN-based inversion.
method Functional-prior-based approaches (fpBPINN) to Bayesian PDE-constrained inversion using physics-informed neural networks (PINNs). Two complementary approaches: FPI-BPINN and fParVI-PINN.
result Accurate estimation of posterior distributions in seismic traveltime tomography and Darcy-flow permeability inversion.