A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Derives gradient bounds for f-heat equations on manifolds with Bakry-Emery Ricci curvature.
problem Gradient estimates for positive solutions of f-heat equations on manifolds with specific curvature conditions.
method Applies Li-Yau gradient estimates to positive solutions of the f-heat equation on closed manifolds with Bakry-Emery Ricci curvature bounded below.
result Derives Li-Yau gradient bounds for positive solutions of the f-heat equation.
In this paper, we prove the compactness theorem for gradient Ricci solitons. Let (Mα,gα) be a sequence of compact gradient Ricci solitons of dimension n≥4, whose curvatures have uniformly bounded L2n norms, whose Ricci curvatures are uniformly bounded from below with uniformly lower bounded vol…
In this work we revisit gradient regularization for adversarial robustness with some new ingredients. First, we derive new per-image theoretical robustness bounds based on local gradient information. These bounds strongly motivate input gradient regularization. Second, we implement a scaleable version of input gradient…
We lower bound the complexity of finding ε-stationary points (with gradient norm at most ε) using stochastic first-order methods. In a well-studied model where algorithms access smooth, potentially non-convex functions through queries to an unbiased stochastic gradient oracle with bounded variance, we prove that (i…
New research shows existing information-theoretic methods can't establish minimax rates for gradient descent in stochastic convex optimization.
problem Establishing minimax rates for gradient descent in stochastic convex optimization using information-theoretic methods.
method Examined several information-theoretic frameworks including input-output mutual information bounds, conditional mutual information bounds, PAC-Bayes bounds, and their variants.
result Proved that none of the examined information-theoretic frameworks can establish minimax rates for gradient descent in stochastic convex optimization.
Many continuous control tasks have bounded action spaces. When policy gradient methods are applied to such tasks, out-of-bound actions need to be clipped before execution, while policies are usually optimized as if the actions are not clipped. We propose a policy gradient estimator that exploits the knowledge of action…
New bounds show BBVI's gradient variance matches SGD conditions, improving parameterization efficiency.
problem Understanding and improving the convergence of black-box variational inference (BBVI).
method Showed BBVI satisfies matching gradient variance bounds corresponding to the ABC condition for smooth and quadratically-growing log-likelihoods.
result Proven BBVI's gradient variance matches SGD conditions, with superior dimensional dependence for mean-field parameterization.
Researchers estimate gradients of solutions to a Finslerian Allen-Cahn equation.
problem Estimating gradients of solutions to a specific type of partial differential equation.
method Using the Finslerian Allen-Cahn equation as an Euler-Lagrange equation to a Liapunov entropy functional, proving gradient estimates on compact and noncompact Finsler metric measure spaces.
result Global and local gradient estimates of positive solutions to the Finslerian Allen-Cahn equation.
This work bounds the run-time of nonconvex optimization with early stopping.
problem Bounding the expected run-time of nonconvex optimization with early stopping.
method Derives conditions for well-defined early stopping based on validation function norms and bounds the expected number of iterations and gradient evaluations.
result Guarantees the validity of early stopping and provides bounds on the expected run-time for various optimization algorithms.
DIFF2 improves differential privacy in nonconvex optimization with better utility bounds.
problem Improving differential privacy in nonconvex optimization with better utility bounds.
method DIFF2 constructs a differential private global gradient estimator using gradient differences.
result DIFF2 achieves a utility of \(\widetilde O(d^{2/3}/(n\varepsilon_{\mathrm{DP}})^{4/3})\), significantly better than \(\widetilde O(\sqrt{d}/(n\varepsilon_{\mathrm{DP}}))\).