Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3977951,1921,589 · Jun 202019922001200920182026
48 results for non-gradient descent learning

PILAE learns DNNs without gradient descent, achieving better performance.

problem Training deep feedforward neural networks efficiently and accurately.
method PILAE uses a pseudoinverse learning algorithm for autoencoder building blocks of MLP DNNs.
result PILAE achieves better performance on tradeoff between training efficiency and accuracy.

In this paper, we briefly review the basic scheme of the pseudoinverse learning (PIL) algorithm and present some discussions on the PIL, as well as its variants. The PIL algorithm, first presented in 1995, is a non-gradient descent and non-iterative learning algorithm for multi-layer neural networks and has several adv…

2018-05-20abs ↗pdf ↗

Study defends shallow neural networks from data-poisoning attacks.

problem Protecting shallow neural networks from adversarial attacks during training.
method Developed a non-gradient stochastic algorithm for depth-2 neural networks, proving near-optimal trade-offs.
result Demonstrated improved performance over stochastic gradient descent under various data distributions.

Study of special Lorentzian Lie groups with 4D isometry group, finding all are non-gradient expanding Ricci solitons.

problem Characterizing homogeneous Lorentzian three-manifolds with a 4D isometry group.
method Explicit global coordinate description and proof of Ricci soliton properties.
result All special examples are non-gradient expanding Ricci solitons.

Gradient Ricci solitons can be extended to non-gradient Ricci solitons using energy function.

problem Extending the geometry of gradient Ricci solitons to non-gradient Ricci solitons.
method Using energy function EE to study the geometry.
result A non-steady Ricci soliton with symmetric covariant derivative is gradient.

This research analyzes how input and output layers affect deep neural networks' resistance to adversarial attacks.

problem The vulnerability of deep neural networks to adversarial inputs, especially non-gradient based attacks.
method Analysis of three different fully connected dense network classes with manipulated input and output layers.
result Manipulating input and output layers can significantly enhance a deep neural network's robustness against adversarial attacks.

We study 33-dimensional Ricci solitons which project via a semi-conformal mapping to a surface. We reformulate the equations in terms of parameters of the map; this enables us to give an ansatz for constructing solitons in terms of data on the surface. A complete description of the soliton structures on all the 33-di…

2005-10-14abs ↗pdf ↗

The three-dimensional Heisenberg group H3H_3 has three left-invariant Lorentz metrics g1g_1, g2g_2 and g3g_3. They are not isometric each other. In this paper, we characterize the left-invariant Lorentzian metric g1g_1 as a Lorentz Ricci soliton. This Ricci soliton g1g_1 is a shrinking non-gradient Ricci soliton. Likew…

2009-06-01abs ↗pdf ↗

The paper classifies Ricci solitons and studies harmonic vector fields on a specific Thurston geometry.

problem Classifying Ricci solitons and studying harmonic vector fields in a specific Thurston geometry.
method Left-invariant Riemannian metric classification and analysis of harmonic maps and vector fields.
result All Ricci solitons on (F4,g)(F^4,g) are expanding and non-gradient.

A general Boltzmann machine with continuous visible and discrete integer valued hidden states is introduced. Under mild assumptions about the connection matrices, the probability density function of the visible units can be solved for analytically, yielding a novel parametric density function involving a ratio of Riema…

2017-12-20abs ↗pdf ↗

SGLD proves geometric ergodicity via reflection coupling for nonconvex log-concave distributions.

problem Proving geometric ergodicity of SGLD in nonconvex, log-concave settings.
method Reflection coupling technique to handle SGLD's time discretization and minibatch issues.
result SGLD has an invariant distribution and geometric ergodicity in W1W_1 distance.

With a f-left-invariant Riemannian metric on a Lie group GG, we mean a Riemannian metric which is conformally equivalent to a left-invariant Riemannian metric, with the conformal factor ff. In this article, we study the geometry of such metrics and give a necessary and sufficient condition for an f-left-invariant Rie…

2014-01-03abs ↗pdf ↗

New algorithm improves online learning with reduced discretization.

problem Improving adaptive online learning with refined discretization.
method Continuous time approach to online learning, followed by a new discretization argument.
result Optimal regret bound with O(VT)O(\sqrt{V_T}) dependence on gradient variance.

SOLO uses DNN to optimize complex topology problems with reduced FEM calculations.

problem Optimizing materials distribution in complex domains with high computational cost.
method Integrates DNN with FEM calculations to learn and substitute objective functions dynamically.
result Optimum predicted by DNN converges to true global optimum through iterations.

Reparameterizes mirror descent as gradient descent for efficient sparse learning.

problem Efficiently training small sparse networks with mirror descent.
method Develops a framework to convert mirror descent updates into gradient descent updates on different parameters.
result Mirror descent can be reparameterized as gradient descent on modified parameters, facilitating standard backpropagation.

Killing fields on compact m-quasi-Einstein manifolds are shown under specific curvature conditions.

problem Characterizing Killing fields on compact m-quasi-Einstein manifolds.
method Extending a result by Bahuaud-Gunasekaran-Kunduri-Woolgar, the approach involves proving the existence of Killing fields under certain curvature conditions.
result A sufficient condition for a compact, non-gradient m-quasi-Einstein metric to admit a Killing field is provided, extending the original result to the m = -2 case.

The paper introduces various gradient descent algorithms for training deep learning models.

problem Training deep neural networks is challenging due to their complexity.
method Gradient descent and its variants are discussed for optimizing deep learning models.
result Gradient descent and its variants improve the training performance of deep learning models.

Any gradient descent optimization requires to choose a learning rate. With deeper and deeper models, tuning that learning rate can easily become tedious and does not necessarily lead to an ideal convergence. We propose a variation of the gradient descent algorithm in the which the learning rate is not fixed. Instead, w…

2018-01-27abs ↗pdf ↗

Proposes Hebbian-descent for neural network learning, addressing Hebbian and gradient descent issues.

problem Learning issues with correlated data and vanishing error term in gradient descent.
method Integrates Hebbian and gradient descent principles without activation function derivatives, centering neural activities.
result Biologically plausible, convergent, and effective in online learning with correlated data.

Foundation for robust finance using rough path theory.

problem Mathematical models of financial markets under Knightian uncertainty.
method Introducing Property (RIE) for càdlàg paths, proving existence of rough integrals, verifying admissibility of trading strategies.
result Existence and stability of rough path integrals for non-gradient integrands.

Double descent phenomenon explained in simple terms.

problem Understanding the surprising drop in test error in overparameterized models.
method Informal explanation using linear algebra and probability, visual intuition with polynomial regression, mathematical analysis with ordinary linear regression.
result Three factors create double descent: data undersampling, model size, and parameter count. Ablating any one of these factors prevents double descent.

Gradient descent benefits from tangent kernel advantages under specific conditions.

problem Comparing gradient descent with tangent kernel methods in learning.
method Analysis of gradient descent and tangent kernel methods under different conditions.
result Gradient descent can achieve small error only if tangent kernel methods have a non-trivial advantage, but this advantage can be very small.

The study explores (m,ρ)(m,ρ)-quasi-Einstein structures on contact metric manifolds.

problem Exploring (m,ρ)(m,ρ)-quasi-Einstein structures in contact geometry.
method Proving properties of (m,ρ)(m,ρ)-quasi-Einstein structures on contact metric manifolds.
result Compact contact or HH-contact metric manifolds with (m,ρ)(m,ρ)-quasi-Einstein structures have specific properties.

Proposes a continuous flow model to understand and control instability in gradient descent for deep learning.

problem Understanding and controlling the instability of gradient descent in deep learning.
method Introduces the Principal Flow (PF), a continuous time flow that approximates gradient descent dynamics.
result The PF captures divergent and oscillatory behaviors of gradient descent, including escaping local minima and saddle points.

New method reconstructs non-equilibrium stochastic systems from data.

problem Reconstructing non-equilibrium stochastic systems from ensemble measurements.
method Schrödinger bridge problem with multivariate Ornstein-Uhlenbeck process.
result Simulation-free algorithm achieves higher accuracy than competing methods.

Discovering quasipotential equations from data using machine learning.

problem Understanding escape mechanisms from metastable states in nonlinear systems.
method Combining neural networks and sparse regression to symbolically reconstruct quasipotential equations.
result Model-unbiased analytical forms of quasipotential discovered directly from data.

Gradient descent finds halfspaces with low error for agnostic learning.

problem Agnostic learning of linear halfspaces with convex surrogates.
method Gradient descent on convex surrogates for zero-one loss.
result Gradient descent finds halfspaces with error O(OPT1/2+ε)O(\mathsf{OPT}^{1/2} + \varepsilon) in poly time and sample complexity.

Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.

problem Achieving zero loss minimizers in deep learning networks.
method Analysis of gradient descent algorithm in deep learning, focusing on underparametrized networks.
result Zero loss minimization cannot be achieved generically in deep learning networks.

Study shows double and triple descent in unsupervised autoencoders, improving performance in various tasks.

problem Exploring the phenomenon of double descent in unsupervised learning.
method Analytical demonstration and extensive experiments on synthetic and real datasets.
result Over-parameterized unsupervised autoencoders exhibit double and triple descent, enhancing performance in downstream tasks.

This paper explores a new framework for reinforcement learning based on online convex optimization, in particular mirror descent and related algorithms. Mirror descent can be viewed as an enhanced gradient method, particularly suited to minimization of convex functions in highdimensional spaces. Unlike traditional grad…

2012-10-16abs ↗pdf ↗

Gradient descent optimizes deep ReLU networks with proper initialization.

problem Training deep neural networks with ReLU activation.
method Gradient descent and stochastic gradient descent with proper random weight initialization.
result Gradient descent finds global minima for over-parameterized deep ReLU networks.

Kalman Gradient Descent optimizes machine learning models by reducing variance in stochastic optimization.

problem Reducing variance in stochastic gradient descent to improve optimization performance.
method Uses Kalman filtering to adaptively reduce gradient variance in stochastic gradient descent.
result Improved performance on various machine learning tasks including neural networks and black box variational inference.