Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

98196294392 · Jun 202019922001200920172026
48 results for gradient refinement

Refines neural network predictions using background knowledge for improved accuracy.

problem Compensate for lack of labeled data in neural networks.
method Introduces differentiable refinement functions and Iterative Local Refinement (ILR) algorithm to refine predictions efficiently and accurately.
result ILR finds competitive results in MNIST addition task and refines predictions on complex SAT formulas.

Paper proposes SDRL to improve continual learning with less computational cost.

problem Catastrophic forgetting in continual learning.
method SDRL method that refines gradients from memorized samples to reduce gradient diversity.
result SDRL shows better performance than state-of-the-art methods on multiple benchmark tasks.

Paper introduces a new GG^\star regret measure for online convex optimization with smooth losses.

problem Online convex optimization with smooth losses.
method Introduces a new GG^\star regret measure that depends on the cumulative squared gradient norm.
result The GG^\star regret can be arbitrarily sharper than existing measures when losses have vanishing curvature.

Classifies gradient Ricci solitons with harmonic Weyl curvature in dimensions 5 and above.

problem Classifying gradient Ricci solitons with harmonic Weyl curvature in higher dimensions.
method Developed a novel method of refined adapted frame fields and used geometric arguments.
result Local and complete classifications of gradient Ricci solitons with harmonic Weyl curvature.

Improved analysis for clipped gradient methods in nonsmooth convex optimization under heavy-tailed noise.

problem Optimization under heavy-tailed noise in nonsmooth convex problems.
method Refined analysis of Clipped Stochastic Gradient Descent (Clipped SGD) with new rates and improved utilization of Freedman's inequality.
result New rates O(σldmeff1/2pln11/p(1/δ)T1/p1){\cal O}(σ_{\frak l}d_{ m eff}^{-1/2{\frak p}}\ln^{1-1/{\frak p}}(1/δ)T^{1/{\frak p}-1}) and O(σl2dmeff1/pln22/p(1/δ)T2/p2){\cal O}(σ_{\frak l}^2d_{ m eff}^{-1/{\frak p}}\ln^{2-2/{\frak p}}(1/δ)T^{2/{\frak p}-2}) for nonsmooth convex and strongly convex problems, respectively.

Proves convergence of gradient Ricci shrinkers with uniform bounds.

problem Compactness and energy concentration in gradient Ricci shrinkers.
method Bubble-tree convergence and local energy analysis.
result No energy concentrates in neck regions, leading to a local diffeomorphism finiteness theorem.

Over the past decade there has been considerable interest in spectral algorithms for learning Predictive State Representations (PSRs). Spectral algorithms have appealing theoretical guarantees; however, the resulting models do not always perform well on inference tasks in practice. One reason for this behavior is the m…

2017-02-14abs ↗pdf ↗

Maxout networks study gradients and propose initialization strategies.

problem Complexity in input-output Jacobian distribution complicates stable parameter initialization.
method Obtained bounds on moments of gradients and formulated initialization strategies.
result Parameter initialization strategies improve training of deep maxout networks.

Study on 3-manifolds with nonnegative scalar curvature and positive harmonic functions.

problem Characterizing 3-manifolds with nonnegative scalar curvature.
method Exhaustions by level sets of harmonic functions and refined average gradient estimates.
result Contractible 3-manifolds are diffeomorphic to R^3, and handlebodies have genus at most 1.

Improved lower bound for first Dirichlet eigenvalue using variance refinement.

problem Finding a more precise lower bound for the first Dirichlet eigenvalue.
method Refined Jensen-Hölder averaging using variance term.
result Explicit closed-form in-diameter bound strictly stronger than previous estimates.

Study on gradient h-almost Yamabe solitons with scalar curvature estimation.

problem Exploring triviality and scalar curvature estimation of gradient h-almost Yamabe solitons.
method Established sufficient conditions for triviality and scalar curvature estimation under integral inequalities involving the scalar curvature and soliton function.
result Extended and refined former works on almost and h-almost Yamabe solitons, characterizing their geometric structures.

New algorithm improves online learning with reduced discretization.

problem Improving adaptive online learning with refined discretization.
method Continuous time approach to online learning, followed by a new discretization argument.
result Optimal regret bound with O(VT)O(\sqrt{V_T}) dependence on gradient variance.

New method improves online covariance estimation for SGD.

problem Improving online covariance estimation for SGD.
method Proposes a de-biased covariance estimator that eliminates second-order derivatives.
result Achieves a convergence rate of n(α1)/2lognn^{(α-1)/2} \sqrt{\log n}, outperforming existing methods.

Gradient descent with small random init mimics spectral methods for low-rank matrix recovery.

problem Reconstructing a low-rank matrix from few measurements.
method Gradient descent with small random initialization followed by a few iterations.
result Gradient descent from small random init converges to a well-generalizing solution.

The paper analyzes stability and generalization of shallow neural networks using gradient methods.

problem Understanding the generalization of overparameterized shallow neural networks.
method The paper uses gradient descent and stochastic gradient descent to study shallow neural networks, developing consistent excess risk bounds.
result The analysis improves on existing methods by providing a refined estimation of iterates and Hessian eigenvalues, leading to better excess risk bounds.

The challenge of assigning importance to individual neurons in a network is of interest when interpreting deep learning models. In recent work, Dhamdhere et al. proposed Total Conductance, a "natural refinement of Integrated Gradients" for attributing importance to internal neurons. Unfortunately, the authors found tha…

2018-07-26abs ↗pdf ↗

Unified bounds for random subset generalization error and improved SGD Langevin dynamics.

problem Generalization error bounds for random subsets and stochastic gradient Langevin dynamics.
method Unified framework based on Hellström and Durisi's work, extending bounds for Langevin dynamics.
result Unified and refined bounds for generalization error in stochastic gradient Langevin dynamics.

The standard practice in Generative Adversarial Networks (GANs) discards the discriminator during sampling. However, this sampling method loses valuable information learned by the discriminator regarding the data distribution. In this work, we propose a collaborative sampling scheme between the generator and the discri…

2019-02-02abs ↗pdf ↗

Classically, the time complexity of a first-order method is estimated by its number of gradient computations. In this paper, we study a more refined complexity by taking into account the `lingering' of gradients: once a gradient is computed at xkx_k, the additional time to compute gradients at xk+1,xk+2,x_{k+1},x_{k+2},\dots m…

2019-01-09abs ↗pdf ↗

SVGD algorithm converges at rate 1/sqrt(log log n) for sub-Gaussian distributions.

problem Approximating a probability distribution with particles.
method Stein variational gradient descent (SVGD) with finite particles and sub-Gaussian target distribution.
result SVGD achieves a convergence rate of 1/sqrt(log log n) for sub-Gaussian distributions.

Improved Liouville theorems for ancient solutions to V-harmonic map heat flows.

problem Establishing Liouville theorems for ancient solutions to V-harmonic map heat flows.
method Refined gradient estimates and exponential growth conditions.
result Better Liouville theorems for ancient solutions to V-harmonic map heat flows.

The study analyzes how many neurons are needed for two-layer neural networks trained with gradient descent.

problem Determining the minimum number of neurons required for effective training of shallow neural networks.
method Analyzes two-layer neural networks in the NTK regime, trained with gradient descent. Derives fast rates of convergence and tracks the number of hidden neurons required for generalization.
result Derives fast rates of convergence and improves on existing results for the number of hidden neurons needed for generalization.

This paper analyzes convergence of large-scale Transformers with weight decay.

problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.

New findings show margins are not sufficient for explaining gradient boosting performance.

problem The inadequacy of margin explanations in explaining the performance of gradient boosting.
method Demonstrated and proved a stronger margin-based generalization bound for boosted classifiers.
result Proved a stronger margin-based generalization bound that explains the performance of modern gradient boosters.

New inequality for refined knot invariants in a specific space.

problem General adjunction inequality for refined ss-invariants does not hold.
method Introduced an adjunction inequality for a specific spatial refinement in kCP2k\overline{\mathbb{CP}^2}.
result An adjunction inequality holds for the ss-version of the Sq1Sq^1-refinement in kCP2k\overline{\mathbb{CP}^2}.

We study refined topological string theory in the presence of orientifolds by counting second-quantized BPS states in M-theory. This leads us to propose a new integrality condition for both refined and unrefined topological strings when orientifolds are present. We define the SO(2N) refined Chern-Simons theory which co…

2012-02-20abs ↗pdf ↗

Gradient flow of ReLU networks converges in low-correlation high-dimensional data.

problem Convergence of shallow ReLU networks trained on weakly interacting data.
method Gradient flow analysis with Polyak-Łojasiewicz viewpoint.
result Network width of order log(n) neurons suffices for global convergence with high probability.

In this paper, we consider unregularized online learning algorithms in a Reproducing Kernel Hilbert Spaces (RKHS). Firstly, we derive explicit convergence rates of the unregularized online learning algorithms for classification associated with a general gamma-activating loss (see Definition 1 in the paper). Our results…

2015-03-02abs ↗pdf ↗

New research shows label refinement and weak training have limitations for aligning LLMs.

problem Limitations of refinement methods for aligning large language models.
method Analyzed probabilistic assumptions and alternative approaches to label refinement and weak training.
result Label refinement and weak training suffer from irreducible error, leaving a performance gap.

Generative model initializes 2-layer network weights for small datasets.

problem Approximating functions with 2-layer networks using small datasets and gradient-based training.
method Initialize hidden weights with a learned proposal distribution parameterized as a deep generative model. Refine with gradient-based post-processing and regularization.
result Demonstrates effectiveness of the approach with numerical examples.