Global implicit function theorem for Fréchet spaces, solving derivative loss problems.
problem Solving initial value problems with derivative loss in Fréchet spaces.
method Global implicit function theorems for Keller's Cc1-mappings in Fréchet spaces, applied through submersions and transversality. result Global existence and uniqueness of solutions to initial value problems with derivative loss.
We prove implicit function theorems for mappings on topological vector spaces over valued fields. In the real and complex cases, we obtain implicit function theorems for mappings from arbitrary (not necessarily locally convex) topological vector spaces to Banach spaces.
Deep equilibrium models converge globally without explicit computation.
problem Global convergence of deep learning models with implicit layers.
method Analysis of gradient dynamics and proof of convergence rate.
result Deep equilibrium models converge to global optimum at a linear rate.
This paper proves SGD converges to global minimum for over-parameterized ReLU networks.
problem Theoretical understanding of implicit neural networks is limited.
method Gradient flow analysis of ReLU activated implicit neural networks.
result Randomly initialized gradient descent converges to global minimum at a linear rate for square loss function in over-parameterized ReLU networks.
We prove an implicit function theorem for functions on infinite-dimensional Banach manifolds, invariant under the (local) action of a finite dimensional Lie group. Motivated by some geometric variational problems, we consider group actions that are not necessarily differentiable everywhere, but only on some dense subse…
Proposes a parametric modal regression method using the implicit function theorem.
problem Finding conditional modes for multi-modal conditional distributions.
method Uses the implicit function theorem to develop an objective function for learning a joint function over inputs and targets.
result Empirically demonstrates scalability and effectiveness in learning multi-valued functions and high-dimensional inputs.
Large learning rates lead to various implicit biases in nonconvex optimization.
problem Understanding the conditions under which large learning rates yield edge of stability, balancing, and catapult phenomena.
method Developed a global convergence theory for nonconvex functions without globally Lipschitz continuous gradient, focusing on functions with good regularity.
result These implicit biases are more likely to occur in functions with good regularity, and large learning rates favor flatter regions.
The paper details local forms of morphisms in colored supermanifolds.
problem Understanding local forms of morphisms in colored supermanifolds.
method Detailed account of Z2n-differential calculus and local theorems. result Detailed insights into local forms of morphisms in colored supermanifolds.
This paper shows how to train only the implicit layer of overparameterized implicit neural networks.
problem Understanding how the implicit layer contributes to the training of overparameterized implicit neural networks.
method Restricting training to only the implicit layer and analyzing the generalization error for ReLU-activated networks.
result Global convergence is guaranteed even if only the implicit layer is trained, and gradient flow with proper random initialization can achieve small generalization errors.
Proposes efficient, modular method for implicit differentiation.
problem Implicit differentiation of optimization problems.
method Automatic implicit differentiation using autodiff and implicit function theorem.
result Automatic differentiation of optimization problems is made easier and more modular.
Global inverse function theorem proved easily using Riemannian geometry.
problem Global inverse function theorem in Riemannian geometry.
method Hopf--Rinow theorem in Riemannian geometry.
result Hadamard's global inverse function theorem is proven easily.
Building upon ideas of Hironaka, Bierstone-Milman, Malgrange and others we generalize the inverse and implicit function theorem (in differential, analytic and algebraic setting) to sets of functions of larger multiplicities (or ideals). This allows one to describe singularities given by a finite set of generators or by…
A new method for optimizing non-decomposable metrics with constraints.
problem Optimizing complex machine learning objectives with thresholded constraints.
method Formulate rate-constrained optimization using the Implicit Function theorem and solve with gradient-based methods.
result Demonstrated effectiveness over existing methods on benchmark datasets.
The paper explores how the depth of neural networks affects their ability to represent data accurately.
problem Understanding the implicit bias and rank of neural networks with large depth.
method Analyzing the convergence of representation cost to a notion of rank as network depth increases, and investigating conditions for recovering the true rank of data.
result There is a range of network depths where the true rank of data is recovered, and this affects the topology of class boundaries.
We study the implicit bias of gradient descent methods in solving a binary classification problem over a linearly separable dataset. The classifier is described by a nonlinear ReLU model and the objective function adopts the exponential loss function. We first characterize the landscape of the loss function and show th…
The paper analyzes optimal implicit bias in linear regression for over-parameterized models.
problem Finding the best generalization performance in over-parameterized linear regression.
method Asymptotic analysis of generalization performance for convex functions/potentials.
result Optimal implicit bias that achieves the best generalization error under certain conditions.
New tensor formulation reveals gradient flow's bias in linear neural networks.
problem Understanding implicit bias in linear neural network training.
method Tensor formulation of neural networks, including fully-connected, diagonal, and convolutional networks.
result Gradient flow on linear tensor networks converges to solutions of specific optimization problems.
New dual formulation reduces generalization error for ERM-fDR.
problem Generalization error in constrained optimization problems.
method Introduces a dual formulation of ERM-fDR using Legendre-Fenchel transform and implicit function theorem.
result Explicit characterizations of generalization error for algorithms under mild conditions.
Bayesian optimization improved for large-scale problems.
problem Efficient optimization of expensive functions with many variables.
method Local Bayesian optimization using a collection of local models and a bandit approach for sample allocation.
result TuRBO algorithm outperforms state-of-the-art methods on various high-dimensional problems.
The paper explores moduli space of heterotic system using two deformation paths.
problem Exploring the moduli space of the heterotic system.
method Considering two dual deformation paths starting from a Kähler solution, one along Bott-Chern cohomology class and the other along Aeppli cohomology class. Using the implicit function theorem to prove local existence of heterotic solutions.
result Established an initial step to construct local moduli coordinates around a Kähler solution.
A method to visualize multidimensional local subspaces using implicit differentiation.
problem Understanding the effect of multidimensional projection on local subspaces.
method Implicit function differentiation to analyze local subspaces shaped by multidimensional ellipses.
result Visualization of local subspaces provides insights into the global structure of data.
Gradient descent converges to a global minimum in nonlinear ReLU implicit networks with linear width.
problem Understanding convergence of gradient methods in nonlinear, infinitely deep ReLU networks.
method Introduced a scaling constant to ensure well-posedness of the equilibrium equation, proving convergence to a global minimum for linear width networks.
result Gradient descent converges to a global minimum at a linear rate for nonlinear ReLU implicit networks with linear width.
Minimal surfaces with dihedral symmetry are studied as angles converge to zero.
problem Understanding minimal surfaces with dihedral symmetry as angles approach zero.
method Analyzing the limit of minimal surfaces in wedges with varying angles and using the implicit function theorem.
result New minimal surfaces are discovered and existence proofs are simplified.
Dual optimization connects ERM-fDR to normalization function.
problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.
New methods distill data for deep networks efficiently.
problem Reduction of training data cost and inconvenience.
method Generative teaching networks, gradient matching, Implicit Function Theorem.
result New methods are computationally more efficient and improve model performance.
Reduces function approximation dimensions from high to low with sparse data.
problem Function approximation from sparse data.
method Nonlinear Level Set Learning (NLL) with geometric information.
result Reduces input dimension to theoretical lower bound with minor accuracy loss.
Abstracts a theorem for non-smooth maps in infinite dimensions.
problem Generalizing inverse mapping theorem for non-smooth maps.
method Introduces property A and applies it to non-smooth maps.
result Generalized inverse mapping theorems for non-smooth maps.
We exhibit differential geometric structures that arise in numerical methods, based on the construction of Cauchy sequences, that are currently used to prove explicitly the existence of weak solutions to functional equations. We describe the geometric framework, highlight several examples and describe how two well-know…
Modern implicit generative models such as generative adversarial networks (GANs) are generally known to suffer from issues such as instability, uninterpretability, and difficulty in assessing their performance. If we see these implicit models as dynamical systems, some of these issues are caused by being unable to cont…
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems. We explore the question of whether the specifi…
Paper constructs solutions to a system using Aeppli class without auxiliary gauge connection.
problem Constructing solutions to the Hull-Strominger system without auxiliary gauge connection.
method Deforming conformally balanced metric and tuning by Aeppli class to satisfy anomaly cancellation condition.
result Existence of family of solutions obtained via implicit function theorem.
We introduce a proximal subdifferential and develop a calculus for nonsmooth functions defined on any Riemannian manifold M. We give several applications of this theory, concerning: 1) differentiability and geometrical properties of the distance function to a closed subset C of M; 2) solvability and implicit func…
Algorithm optimizes millions of hyperparameters efficiently.
problem Training modern network architectures with millions of hyperparameters.
method Combines implicit function theorem with efficient inverse Hessian approximations for gradient-based optimization.
result Jointly tuning weights and hyperparameters is only a few times more costly than standard training.
We prove a version of the implicit function theorem for Lipschitz mappings f:Rn+m⊃A→X into arbitrary metric spaces. As long as the pull-back of the Hausdorff content H∞n by f has positive upper n-density on a set of positive Lebesgue measure, then, there is a local diff…
Kuranishi's proof of complex deformation theory revisited
problem Existence of complex deformations on compact complex manifolds
method Hamilton-Nash-Moser implicit function theorem
result Revisits classical proof with modern tools
Our paper characterizes how ReLU affects GD's implicit bias in high-dimensional neural networks.
problem Understanding the implicit bias of gradient descent on neural networks.
method Novel primal-dual analysis tracking predictions and coefficients.
result The implicit bias approximates the minimum-ℓ2-norm solution with high probability. We establish a glueing theorem for the Ginzburg-Landau equations in dimension n>2. To this end, we consider a nondegenerate minimal submanifold of codimension 2, and construct a one-parameter family of solutions to the Ginzburg-Landau equations such that the energy density concentrates near this submanifold. The pr…
A geometric flow on (2,2)-forms is introduced which preserves the balanced condition of metrics, and whose stationary points satisfy the anomaly equation in Strominger systems. The existence of solutions for a short time is established, using Hamilton's version of the Nash-Moser implicit function theorem.
Learning algorithms for implicit generative models can optimize a variety of criteria that measure how the data distribution differs from the implicit model distribution, including the Wasserstein distance, the Energy distance, and the Maximum Mean Discrepancy criterion. A careful look at the geometries induced by thes…
Synthetic proof shows globally hyperbolic Lorentzian spaces with specific curvature are warped products.
problem Synthetic proof of rigidity for globally hyperbolic Lorentzian spaces.
method Synthetic geometry and warped product analysis.
result Spaces with specific curvature and distance realizer are warped products.
Harmonic functions stable under small changes.
problem Stability of multivalued harmonic functions under deformations.
method Application of Nash-Moser implicit function theorem.
result Stability of harmonic sections under small deformations.
Recently, there has been a growing interest in the problem of learning rich implicit models - those from which we can sample, but can not evaluate their density. These models apply some parametric function, such as a deep network, to a base measure, and are learned end-to-end using stochastic optimization. One strategy…
Kernel-guided training stabilizes GANs by controlling discrepancies.
problem Stability and interpretability issues in GANs.
method Kernel-based regularization to control discrepancies in GAN loss function.
result Theoretical guarantees on stability of the training dynamics.
Local gluing connects flow lines in finite time intervals.
problem Connecting flow lines in finite time intervals.
method Functional analytic approach to define local gluing map.
result Explicit construction of local gluing map in Euclidean case; intricate construction in non-Euclidean case.
Neural networks can approximate gradient of smooth functions, but with limitations.
problem Approximating gradient of smooth functions using neural networks.
method Proving limitations of neural networks with more than one hidden layer and introducing implicit parametrization.
result Neural networks with more than one hidden layer can only represent one feature in their first hidden layer.
We give sufficient conditions for a Cc1-local diffeomorphism between Fréchet spaces to be a global one. We extend the Clarke's theory of generalized gradients to the more general setting of Fréchet spaces. As a consequence, we define the Chang Palais-Smale condition for Lipschitz functions and show that a functio…
Local gap theorem for Ricci shrinkers ensures flatness if certain functionals are close to zero.
problem Understanding the global geometry of Ricci shrinkers from local information.
method Proving a local gap theorem using the local μ-functional. result Ricci shrinkers are flat if the local μ-functional is close to zero. We interpret the variational inference of the Stochastic Gradient Descent (SGD) as minimizing a new potential function named the \textit{quasi-potential}. We analytically construct the quasi-potential function in the case when the loss function is convex and admits only one global minimum point. We show in this case th…