New approach connects stochastic gradient descent to ODE splitting schemes.
problem Improving convergence in stochastic optimization.
method Connection between stochastic gradient descent and ODE splitting schemes.
result Derive a new upper bound on global splitting error.
The paper proves conditions for Einstein solitons to split into line and manifold.
problem Conditions for Einstein solitons to split into line and manifold.
method Weighted Laplacian comparison of distance function and bounded integral condition on Ricci curvature.
result Gradient ρ-Einstein solitons split off a line isometrically under certain conditions.
Study shows splitting schemes can approximate WFR flows faster than the exact flow.
problem Improving sampling efficiency in Wasserstein-Fisher-Rao gradient flows.
method Investigates operator splitting techniques to numerically approximate WFR flows.
result A judicious choice of step size and operator ordering can lead to faster convergence of split schemes to the target distribution.
We discuss some geometric conditions under which a complete noncompact shrinking gradient Ricci soliton will split at infinity.
TSSM splits neural networks for parallel training with minimal accuracy loss.
problem Accuracy degradation in parallel training of deep neural networks.
method TSSM reformulates alternating minimization to achieve parallelism with minimal accuracy loss.
result TSSM achieves significant speedup without accuracy loss on multiple datasets.
A new method for growing neural networks by splitting neurons, improving efficiency.
problem Optimizing neural network structures, especially for lightweight architectures.
method A progressive training approach using steepest descent to adaptively grow and split neurons.
result The method provides a computationally efficient way to optimize neural network structures.
New minibatching strategy reduces stochastic gradient bias in optimisation.
problem Reducing bias in stochastic gradient descent.
method Symmetric Minibatching Strategy combined with momentum.
result Reduced stochastic gradient bias from O(h2) to O(h4). Backpropagation-free trunk training improves model performance on various benchmarks.
problem Memory inefficiency and noisy gradient estimates in deep network training.
method Split Forward Gradient (Split-FG) method that splits network into trunk and head, estimating only trunk gradient.
result Split-FG achieves better performance than pure forward-gradient training and backpropagation on various benchmarks.
LoBoost improves local conformal prediction for gradient-boosted trees without extra data splits.
problem Quantifying uncertainty in gradient-boosted tree predictions.
method Model-native local conformal prediction using leaf structure.
result Competitive interval quality and improved test MSE with large calibration speedups.
Gradient bounds and Liouville theorems for quasi-linear equations on manifolds with nonnegative Ricci curvature.
problem Establishing bounds and theorems for solutions to quasi-linear elliptic equations on compact manifolds with nonnegative Ricci curvature.
method Gradient bounds, Liouville-type theorems, local splitting theorem, Harnack-type inequality, ABP estimate.
result Gradient bounds and Liouville-type theorems for solutions to quasi-linear equations on compact manifolds with nonnegative Ricci curvature.
The works of Donaldson and Mark make the structure of the Seiberg-Witten invariant of 3-manifolds clear. It corresponds to certain torsion type invariants counting flow lines and closed orbits of a gradient flow of a circle-valued Morse map on a 3-manifold. We study these invariants using the Morse-Novikov theory and H…
Temporal difference learning explained through gradient splitting, improving convergence times.
problem Learning value functions in Markov Decision Processes with linear approximations.
method Interpreting TD learning as gradient splitting and applying convergence proofs from gradient descent.
result Improved convergence times for TD learning, especially with a minor variation.
The paper derives gradient estimates for porous medium and fast diffusion equations on metric measure spaces.
problem Gradient estimates for porous medium and fast diffusion equations on metric measure spaces.
method Derives Li-Yau and Souplet-Zhang type gradient estimates for the given equations.
result Gradient estimates for the equations on complete noncompact metric measure spaces with compact boundary.
We investigate the structure of a Finsler manifold of nonnegative weighted Ricci curvature including a straight line, and extend the classical Cheeger-Gromoll-Lichnerowicz splitting theorem. Such a space admits a diffeomorphic, measure-preserving splitting in general. As for a special class of Berwald spaces, we can pe…
We show that if a closed hyperbolic 3-manifold has infinitely many finite covers of bounded Heegaard genus, then it is virtually fibered. This generalizes a theorem of Lackenby, removing restrictions needed about the regularity of the covers. Furthermore, we can replace the assumption that the covers have bounded Heega…
Improved neural architecture optimization for energy efficiency.
problem Designing energy-efficient deep learning networks for mobile and edge devices.
method Incorporates energy cost in splitting process and uses a scalable stochastic gradient algorithm to speed up the splitting.
result Trains highly accurate and energy-efficient networks on challenging datasets like ImageNet.
We prove a splitting theorem for complete gradient Ricci soliton with nonnegative curvature and establish a rigidity theorem for codimension one complete shrinking gradient Ricci soliton in Rn+1 with nonnegative Ricci curvature.
New theorem splits weighted Lorentz-Finsler manifolds into simpler parts.
problem Understanding the geometry of weighted Lorentz-Finsler manifolds.
method Developed a splitting theorem using weighted Berwald spacetimes and Busemann functions.
result Weighted Lorentz-Finsler manifolds with certain properties split into simpler isometric translations.
Boosting as gradient descent algorithms is one popular method in machine learning. In this paper a novel Boosting-type algorithm is proposed based on restricted gradient descent with structural sparsity control whose underlying dynamics are governed by differential inclusions. In particular, we present an iterative reg…
The study examines gradient Ricci solitons with nonnegative curvature, proving properties of their blow-downs.
problem Characterizing gradient Ricci solitons with nonnegative curvature operator away from a compact set.
method Analyzing blow-downs and limits of Ricci flows to prove properties of solitons.
result No (n−1)-dimensional compact split limit Ricci flow can arise from the blow-down of (M,g). We prove a sharp integral gradient estimate for harmonic functions on noncompact Kähler manifolds. As application, we obtain a sharp estimate for the bottom of spectrum of the p-Laplacian and prove a splitting theorem for manifolds achieving this estimate.
agtboost speeds up gradient tree boosting with automatic complexity adjustment.
problem Speeding up and simplifying gradient tree boosting computations.
method Adaptive gradient tree boosting with automatic complexity adjustment and feature importance.
result Significant decrease in computation time and simplification of model complexity.
New method uses symmetric splitting for efficient HMC inference in large neural networks.
problem Efficient inference for Bayesian neural networks with large datasets.
method Introduces a symmetric integration scheme for Hamiltonian Monte Carlo (HMC) that does not rely on stochastic gradients.
result Symmetric splitting leads to more efficient HMC inference over large data sets.
New splitting theorems in a semi-Riemannian manifold which admits an irrotational vector field (not necessarily a gradient) with some suitable properties are obtained. According to the extras hypothesis assumed on the vector field, we can get twisted, warped or direct decompositions. Some applications to Lorentzian man…
We study a stochastic and distributed algorithm for nonconvex problems whose objective consists of a sum of N nonconvex Li/N-smooth functions, plus a nonsmooth regularizer. The proposed NonconvEx primal-dual SpliTTing (NESTT) algorithm splits the problem into N subproblems, and utilizes an augmented Lagrangian b…
A new transformer model accelerates training with optimization techniques.
problem Training deep neural networks efficiently and effectively.
method Interprets transformer layers as optimization steps, applying Nesterov acceleration.
result The new model outperforms existing models on benchmark datasets.
New algorithms solve monotone inclusions and convex-concave minimax problems.
problem Solving maximally monotone equations and inclusions.
method Developed new accelerated algorithms based on Halpern-type fixed-point iteration and Popov's past extra-gradient method.
result Achieved O(1/k) convergence rates for various problems. The paper investigates conditions for compactness of submanifolds in Kahler manifolds.
problem Conditions for compactness of submanifolds in Kahler manifolds.
method Analyzes the action of Lie groups on Kahler manifolds and investigates conditions for compactness.
result Proves a splitting result for real connected submanifolds of Kahler manifolds.
In this paper we obtain a splitting theorem for the symmetric diffusion operator Δφ=Δ−⟨∇φ,∇⟩ and a non-constant C3 function f in a complete Riemannian manifold M, under the assumptions that the Ricci curvature associated with Δφ satisfies Ricφ(∇f,∇f)≥0, that $|…
We characterize complete nonnegatively curved steady gradient soliton with curvature in L^1. We show that there are isometric to a product (R^2,g_{cigar}) times(R^{n-2}, eucl))/Gamma where Gamma is a Bieberbach group of rank n-2. We prove also a similar local splitting result under weaker curvature assumptions.
Optimization is at the heart of machine learning, statistics and many applied scientific disciplines. It also has a long history in physics, ranging from the minimal action principle to finding ground states of disordered systems such as spin glasses. Proximal algorithms form a class of methods that are broadly applica…
New algorithms split deep learning tasks into representation and uncertainty estimation.
problem Challenges in uncertainty quantification for deep learning models.
method Proposes a two-stage approach: representation learning and uncertainty estimation.
result Simple methods outperform complex uncertainty layers in selective classification and out-of-distribution detection.
CSE-FSL reduces communication and storage costs in federated learning.
problem High communication and storage costs in federated learning.
method CSE-FSL uses an auxiliary network to locally update client models and sends only selected epochs' smashed data.
result Significant communication reduction with state-of-the-art convergence and model accuracy.
Combining Bayesian deep learning and split conformal prediction affects out-of-distribution coverage.
problem Improving out-of-distribution coverage in multiclass image classification.
method Combining Bayesian deep learning with split conformal prediction methods.
result Combining methods can reduce out-of-distribution coverage in some cases.
Innovative method solves nonconvex optimization on manifolds.
problem Nonconvex optimization problems on Riemannian manifolds.
method Intrinsic Riemannian proximal gradient method.
result Converges for nonconvex or nonembedded problems.
Paper proves properties of minimal hypersurfaces in specific solitons.
problem Characterizing minimal hypersurfaces in shrinking gradient Ricci solitons.
method Analyzes stable minimal hypersurfaces with specific curvature conditions.
result Minimal hypersurfaces in these solitons have zero second fundamental form and normal Ricci curvature.
A novel gradient-based method optimizes decision trees for complex tasks.
problem Training decision trees with arbitrary differentiable loss functions.
method Gradient-based optimization using first and second derivatives of loss functions.
result Improves accuracy and flexibility in decision tree optimization.
Study minimal graphs on non-negative Ricci curvature manifolds.
problem Minimal graphs with linear growth on manifolds with non-negative Ricci curvature.
method New gradient estimate for minimal graphs and heat equation techniques.
result Non-constant minimal graphs force tangent cones to split off a line.
We generalize the splitting theorem of Cai-Galloway for complete Riemannian manifolds with $\Ric\geq-(n-1)$ admitting a family of compact hypersurfaces tending to infinity with mean curvatures tending to n−1 sufficiently fast to the setting of smooth metric measure spaces. This result complements and provides a new p…
SketchBoost accelerates GBDT for multioutput problems up to 40x.
problem Efficiently training GBDT for multioutput problems with high-dimensional outputs.
method Approximate computation of scoring function for faster decision tree splitting.
result SketchBoost speeds up GBDT training by up to 40 times.
Gradient Boosting Decision Tree (GBDT) are popular machine learning algorithms with implementations such as LightGBM and in popular machine learning toolkits like Scikit-Learn. Many implementations can only produce trees in an offline manner and in a greedy manner. We explore ways to convert existing GBDT implementatio…
New test improves tree ensemble pruning for better model performance.
problem Lack of robust theoretical justification for penalty terms in tree ensembles.
method Developed a novel hypothesis test for tree ensemble split quality.
result Significant reduction in out-of-sample loss using the new test.
New sampling technique improves boosting model accuracy.
problem Improving generalization performance and learning time in stochastic gradient boosting.
method Formulated optimization problem to maximize estimation accuracy, leading to Minimal Variance Sampling (MVS).
result MVS significantly increases model quality and reduces the number of examples needed.
The asymptotic behavior of the stochastic gradient algorithm with a biased gradient estimator is analyzed. Relying on arguments based on the dynamic system theory (chain-recurrence) and the differential geometry (Yomdin theorem and Lojasiewicz inequality), tight bounds on the asymptotic bias of the iterates generated b…
The study generalizes curvature bounds for manifolds with boundary.
problem Proving curvature bounds for manifolds with boundary.
method Bakry-Émery curvature bounds and splitting theorems.
result Proves curvature bounds for manifolds with boundary.
Method solves optimisation problems on non-Riemannian surfaces with bilateral curvature bounds.
problem Optimisation problems on non-Riemannian surfaces with sharp edges.
method Forward-backward splitting in Alexandrov spaces with bilateral curvature bounds.
result Convergence of the forward-backward method in Alexandrov spaces with bilateral curvature bounds.
LEARN-SAM improves RL from sub-optimal demonstrations by localizing expert policies and selectively using demonstrations.
problem Improving RL from sub-optimal or sparse demonstrations.
method Local Ensemble and Reparameterization with Split and Merge of expert policies (LEARN-SAM).
result LEARN-SAM boosts learning speed and accuracy by selectively using demonstrations.
This paper explores how train-validation splits help in NAS to prevent overfitting.
problem NAS overfits with train-validation splits and needs better generalization guarantees.
method Established refined properties of validation loss and risk for NAS.
result NAS with train-validation splits can select the most generalizable model.