Gradient estimates for solutions to a p-Laplacian equation on Riemannian manifolds.
problem Gradient estimates for positive weak solutions to a p-Laplacian equation on Riemannian manifolds.
method Morser iteration technique
result Gradient estimates show that positive weak solutions do not exist under certain conditions on manifolds with nonnegative Ricci curvature.
Paper presents a new policy gradient theorem using weak derivatives for reinforcement learning.
problem Continuous state-action reinforcement learning problems.
method Introduced an alternative policy gradient theorem using weak derivatives.
result The new approach yields algorithms that converge almost surely to stationary points of the value function.
Paper investigates conditions for independence of weak gradients on metric spaces.
problem Dependence of weak gradients on p in arbitrary metric measure spaces. method Investigates the Bounded Interpolation Property to ensure independence of weak gradients.
result Bounded Interpolation Property guarantees independence of weak gradients.
Weak correlations explain linear dynamics in deep learning models.
problem Understanding the linear structure in gradient-based learning algorithms.
method Characterization of weak correlations between derivatives and parameters.
result Weak correlations are the underlying principle for linearization in deep learning models.
Studied SGD convergence under weak conditions.
problem Convergence of SGD in nonconvex optimization.
method Analyzed biased nonconvex SGD under mild conditions.
result Provided convergence rates and complexities.
Stein variational gradient descent (SVGD) is a deterministic sampling algorithm that iteratively transports a set of particles to approximate given distributions, based on an efficient gradient-based update that guarantees to optimally decrease the KL divergence within a function space. This paper develops the first th…
Study shows how weak inverse anisotropic mean curvature flow behaves at infinity.
problem Understanding the asymptotic behavior of anisotropic mean curvature flow.
method Established local gradient estimates for anisotropic p-harmonic functions and weak solutions of IAMCF. result Weak IAMCF is asymptotic to the expanding Wulff shape solution at infinity.
Study on test risk dynamics in learning theory with stochastic gradient flow.
problem Understanding test risk in stochastic gradient flow dynamics.
method Path integral formulation for small learning rates, explicit computation for weak features.
result Explicit corrections due to stochastic term in dynamics, good agreement with simulations.
Optimizes shapes on non-standard manifolds.
problem Optimization on non-standard infinite-dimensional manifolds.
method Develops gradient descent on weak Riemannian manifolds.
result Establishes foundational properties for optimization on various weak Riemannian manifolds.
Self-test loss functions improve data-driven modeling of weak-form operators and gradient flows.
problem Challenges in selecting test functions for data-driven modeling involving weak-form operators and gradient flows.
method Introducing self-test loss functions that depend on unknown parameters and are quadratic.
result Self-test loss functions conserve energy for gradient flows and coincide with log-likelihood ratios for stochastic differential equations.
We prove the weak stability of expanding gradient Ricci solitons with positive curvature operator and quadratic curvature decay at infinity.
Proposes a constrained labeling method for weakly supervised learning.
problem Combining weak supervision signals while navigating misleading correlations.
method Randomized constrained labeling within a defined space.
result Randomized constrained labeling converges after few iterations and outperforms other methods.
GrowNet uses shallow neural networks for gradient boosting, outperforming existing methods.
problem Improving gradient boosting performance through shallow neural networks.
method Unified gradient boosting framework with shallow neural networks as weak learners, incorporating corrective steps.
result GrowNet outperformed state-of-the-art boosting methods in classification, regression, and learning to rank tasks.
Paper analyzes weak-to-strong generalization in CNNs, identifying data-scarce and data-abundant regimes.
problem Weak-to-strong generalization in CNNs trained on weak models.
method Formal analysis of gradient descent dynamics in data-scarce and data-abundant regimes.
result Identifies two regimes and distinct mechanisms of generalization in each.
Gradient descent benefits from tangent kernel advantages under specific conditions.
problem Comparing gradient descent with tangent kernel methods in learning.
method Analysis of gradient descent and tangent kernel methods under different conditions.
result Gradient descent can achieve small error only if tangent kernel methods have a non-trivial advantage, but this advantage can be very small.
Paper proves uniqueness of weak solutions for Plateau flow.
problem Proving uniqueness of weak solutions for Plateau flow.
method Used natural energy condition and alternative methods from Struwe.
result Proves uniqueness of weak solutions under natural condition.
New methods validate a hypothesis explaining how neural nets generalize well.
problem Why over-parameterized nets generalize well despite memorizing training data.
method Developed new algorithms to suppress weak gradient directions without per-example gradients.
result Validated a hypothesis about gradient directions and their role in generalization.
We propose a projected semi-stochastic gradient descent method with mini-batch for improving both the theoretical complexity and practical performance of the general stochastic gradient descent method (SGD). We are able to prove linear convergence under weak strong convexity assumption. This requires no strong convexit…
The paper classifies special hypersurfaces in space forms.
problem Investigating gradient Yamabe solitons in space forms.
method Using the weak Omori-Yau principle for the drifted Laplacian.
result Gradient Yamabe solitons are fully classified under certain conditions.
A new parametric method studies Willmore flows and energy quantization.
problem Understanding Willmore flows and their singularities.
method Parametric approach to Willmore gradient flows.
result For small-energy weak immersions, a unique solution exists.
Study of 4D Ricci solitons with symmetry, finding precise geometric asymptotics.
problem Classifying 4D gradient steady Ricci solitons and understanding their geometric properties.
method Analysis of 4D gradient steady Ricci solitons with O(3)-symmetry under a weak curvature decay condition.
result Find precise geometric asymptotics similar to 3D compact κ-solutions.
Gradient descent converges to minimum Bayes risk for two-layer ReLU networks in mean field regime.
problem Training two-layer ReLU networks using gradient descent in the mean field regime.
method Describes a condition for convergence to minimum Bayes risk, extending previous results to ReLU-activated networks.
result The condition for convergence does not depend on initialization and concerns weak convergence of network realization.
We prove that the evolution of weight vectors in online gradient descent can encode arbitrary polynomial-space computations, even in very simple learning settings. Our results imply that, under weak complexity-theoretic assumptions, it is impossible to reason efficiently about the fine-grained behavior of online gradie…
Diffusion approximation provides weak approximation for stochastic gradient descent algorithms in a finite time horizon. In this paper, we introduce new tools motivated by the backward error analysis of numerical stochastic differential equations into the theoretical framework of diffusion approximation, extending the …
Boosting is a popular way to derive powerful learners from simpler hypothesis classes. Following previous work (Mason et al., 1999; Friedman, 2000) on general boosting frameworks, we analyze gradient-based descent algorithms for boosting with respect to any convex objective and introduce a new measure of weak learner p…
The paper examines partial regularity of Lipschitz solutions to minimal surface system.
problem Understanding the regularity of solutions to the minimal surface system.
method Investigation of stationary, integral weak, and viscosity solutions; interior gradient estimate using maximum principle.
result Partial regularity results for Lipschitz solutions, including interior gradient estimate.
Gradient boosting is a prediction method that iteratively combines weak learners to produce a complex and accurate model. From an optimization point of view, the learning procedure of gradient boosting mimics a gradient descent on a functional variable. This paper proposes to build upon the proximal point algorithm, wh…
In this paper, we will establish an elliptic local Li-Yau gradient estimate for weak solutions of the heat equation on metric measure spaces with generalized Ricci curvature bounded from below. One of its main applications is a sharp gradient estimate for the logarithm of heat kernels. These results seem new even for s…
In dimension n=3, there is a complete theory of weak solutions of Ricci flow - the singular Ricci flows introduced by Kleiner and Lott - which are unique across singularities, as was proved by Bamler and Kleiner. We show that uniqueness should not be expected to hold for Ricci flow weak solutions in dimensions $n\geq…
Study weak f-K-contact manifolds, finding Einstein-type metrics and solitons.
problem Characterize and study geometric properties of weak f-K-contact manifolds. method Analyzing weak metric f-structures, using Killing vector fields, and Jacobi operators. result Einstein weak f-K-contact manifolds are Ricci flat. This article analyzes the weak error of SGD optimization schemes.
problem Analyzing the error in SGD optimization schemes with respect to a test function.
method Weak error analysis for SGD type optimization schemes.
result The weak error decays at the same speed as in the strong sense.
Large learning rates cause oscillations in NN weights that improve generalization.
problem Improving generalization of neural networks trained with large learning rates.
method Theoretical analysis and feature-noise data generation model.
result Oscillating SGD with large learning rates benefits NN generalization by effectively learning weak features.
Paper proposes a weak approximation of reflection coupling for non-convex optimization.
problem Non-convex optimization problems with different drift terms.
method Proposes an approximate reflection coupling (ARC) for stochastic differential equations (SDEs).
result ARC converges weakly to the reflection coupling and can be applied to non-convex optimization.
The problem of prescribing conformally the scalar curvature of a closed Riemannian manifold as a given Morse function reduces to solving an elliptic partial differential equation with critical Sobolev exponent. Two ways of attacking this problem consist in subcritical approximations or negative pseudo gradient flows. W…
Lipschitz regularity proved for harmonic map heat flows into CAT(0) spaces.
problem Proving Lipschitz regularity for harmonic map heat flows into CAT(0) spaces.
method Elliptic approximation method
result Every weak solution of the harmonic map heat flow into CAT(0) spaces is Lipschitz continuous in both space and time.
Faster weak supervision framework using triplet methods.
problem Computational inefficiency in weak supervision models.
method Closed-form solution for latent variable models, avoiding iterative methods.
result Orders of magnitude faster than previous approaches.
In this paper, we will study the (linear) geometric analysis on metric measure spaces. We will establish a local Li-Yau's estimate for weak solutions of the heat equation and prove a sharp Yau's gradient gradient for harmonic functions on metric measure spaces, under the Riemannian curvature-dimension condition $RCD^*(…
We develop the method of stochastic modified equations (SME), in which stochastic gradient algorithms are approximated in the weak sense by continuous-time stochastic differential equations. We exploit the continuous formulation together with optimal control theory to derive novel adaptive hyper-parameter adjustment po…
Gradient Boosting Machine (GBM) is an extremely powerful supervised learning algorithm that is widely used in practice. GBM routinely features as a leading algorithm in machine learning competitions such as Kaggle and the KDDCup. In this work, we propose Accelerated Gradient Boosting Machine (AGBM) by incorporating Nes…
In this paper we show how techniques coming from stochastic analysis, such as stochastic completeness (in the form of the weak maximum principle at infinity), parabolicity and Lp-Liouville type results for the weighted Laplacian associated to the potential may be used to obtain triviality, rigidity results, and scal…
Paper develops methods for statistical inference in SGD with infinite variance.
problem Challenges in statistical inference for SGD with infinite variance.
method Model-agnostic methodology based on weak convergence and subsampling calibration.
result Asymptotically valid confidence regions for SGD in both finite and infinite variance regimes.
Improved sampling method using regularized Stein Variational Gradient Flow.
problem Improving the accuracy of sampling methods in machine learning.
method Proposed Regularized Stein Variational Gradient Flow to interpolate between SVGD and Wasserstein Gradient Flow.
result Established theoretical properties and provided preliminary numerical evidence of improved performance.
New algorithm improves convergence of gradient boosting trees.
problem Global convergence of Newton boosting in tabular machine learning.
method Introduces Gradient Regularized Newton Descent for GBDTs, proving linear convergence for smooth, strongly convex losses and O(k21) rate for general convex losses. result Achieves globally convergent second-order GBDT algorithm with rate matching first-order boosting.
The goal of policy gradient approaches is to find a policy in a given class of policies which maximizes the expected return. Given a differentiable model of the policy, we want to apply a gradient-ascent technique to reach a local optimum. We mainly use gradient ascent, because it is theoretically well researched. The …
Improves data labeling efficiency in machine learning.
problem Efficiency in data labeling for machine learning models.
method Weakly supervised learning, active labeling, stochastic gradient descent.
result Derives a new algorithm for active labeling that scales better with input dimension.
Natural images are virtually surrounded by low-density misclassified regions that can be efficiently discovered by gradient-guided search --- enabling the generation of adversarial images. While many techniques for detecting these attacks have been proposed, they are easily bypassed when the adversary has full knowledg…
Bayesian priors improve neural network performance on weak signals.
problem Challenges in encoding domain knowledge for weak signals in neural networks.
method Proposed a new joint prior over local scale parameters for feature sparsity and signal-to-noise ratio, optimized with Stein gradient.
result Improved prediction accuracy on various datasets, including genetics applications with weak and sparse signals.
No expanding breathers found in noncompact Ricci flows with certain curvature conditions.
problem Finding expanding breathers in noncompact Ricci flows.
method Curvature positivity conditions (weak PIC-2 or nonnegative bisectional curvature).
result Complete noncompact expanding breathers are gradient solitons.