PILAE learns DNNs without gradient descent, achieving better performance.
problem Training deep feedforward neural networks efficiently and accurately.
method PILAE uses a pseudoinverse learning algorithm for autoencoder building blocks of MLP DNNs.
result PILAE achieves better performance on tradeoff between training efficiency and accuracy.
PairNets optimize AI models for fast IoT applications.
problem Slow training and high memory usage of deep neural networks.
method Developed Pairwise Neural Networks (PairNets) with low memory and fast training.
result PairNets achieve faster training (one epoch) and lower prediction errors.
In this paper, we briefly review the basic scheme of the pseudoinverse learning (PIL) algorithm and present some discussions on the PIL, as well as its variants. The PIL algorithm, first presented in 1995, is a non-gradient descent and non-iterative learning algorithm for multi-layer neural networks and has several adv…
Study defends shallow neural networks from data-poisoning attacks.
problem Protecting shallow neural networks from adversarial attacks during training.
method Developed a non-gradient stochastic algorithm for depth-2 neural networks, proving near-optimal trade-offs.
result Demonstrated improved performance over stochastic gradient descent under various data distributions.
Study of special Lorentzian Lie groups with 4D isometry group, finding all are non-gradient expanding Ricci solitons.
problem Characterizing homogeneous Lorentzian three-manifolds with a 4D isometry group.
method Explicit global coordinate description and proof of Ricci soliton properties.
result All special examples are non-gradient expanding Ricci solitons.
Study on non-gradient Ricci almost solitons in warped products.
problem Understanding non-gradient Ricci almost solitons.
method Construction method and explicit example in warped products.
result Rigidity result for Gaussian soliton.
Gradient Ricci solitons can be extended to non-gradient Ricci solitons using energy function.
problem Extending the geometry of gradient Ricci solitons to non-gradient Ricci solitons.
method Using energy function E E E to study the geometry. result A non-steady Ricci soliton with symmetric covariant derivative is gradient.
The paper classifies quasi-Einstein 3-manifolds and their properties.
problem Classifying compact locally homogeneous non-gradient quasi-Einstein 3-manifolds.
method Analyzing quotient spaces of Lie groups and using properties of quasi-Einstein metrics.
result Identifies conditions for the existence of nontrivial quasi-Einstein metrics.
This research analyzes how input and output layers affect deep neural networks' resistance to adversarial attacks.
problem The vulnerability of deep neural networks to adversarial inputs, especially non-gradient based attacks.
method Analysis of three different fully connected dense network classes with manipulated input and output layers.
result Manipulating input and output layers can significantly enhance a deep neural network's robustness against adversarial attacks.
A1GM method improves efficiency in reconstructing missing data using KL divergence.
problem Efficiently reconstructing missing data in matrices.
method Fast non-gradient-based rank-1 NMF using KL divergence.
result A1GM outperforms gradient methods in efficiency with competitive reconstruction errors.
We study 3 3 3 -dimensional Ricci solitons which project via a semi-conformal mapping to a surface. We reformulate the equations in terms of parameters of the map; this enables us to give an ansatz for constructing solitons in terms of data on the surface. A complete description of the soliton structures on all the 3 3 3 -di…
The three-dimensional Heisenberg group H 3 H_3 H 3 has three left-invariant Lorentz metrics g 1 g_1 g 1 , g 2 g_2 g 2 and g 3 g_3 g 3 . They are not isometric each other. In this paper, we characterize the left-invariant Lorentzian metric g 1 g_1 g 1 as a Lorentz Ricci soliton. This Ricci soliton g 1 g_1 g 1 is a shrinking non-gradient Ricci soliton. Likew…
Introduces new info-geometric structure for dynamics on graphs and hypergraphs.
problem Modeling dynamics on discrete structures like graphs and hypergraphs.
method Introduces two dually flat structures: one on vertex space and another on edge space.
result Extends gradient flows to include nonequilibrium dynamics.
The paper classifies Ricci solitons and studies harmonic vector fields on a specific Thurston geometry.
problem Classifying Ricci solitons and studying harmonic vector fields in a specific Thurston geometry.
method Left-invariant Riemannian metric classification and analysis of harmonic maps and vector fields.
result All Ricci solitons on ( F 4 , g ) (F^4,g) ( F 4 , g ) are expanding and non-gradient. Paper studies non-gradient almost Yamabe solitons and their properties.
problem Characterizing structures of non-gradient almost Yamabe solitons.
method Investigates conditions for trivial solitons and local warped product structures.
result Almost Yamabe solitons with closed vector fields admit local warped product structures.
A general Boltzmann machine with continuous visible and discrete integer valued hidden states is introduced. Under mild assumptions about the connection matrices, the probability density function of the visible units can be solved for analytically, yielding a novel parametric density function involving a ratio of Riema…
New quasi-Einstein metrics found on a sphere.
problem Finding quasi-Einstein metrics on a sphere.
method Constructing axi-symmetric non-gradient m m m -quasi-Einstein structures using hypergeometric functions. result Found new regular metrics on a two-sphere, including the extreme Kerr black hole horizon.
Study of Bach flow on specific nilmanifolds, converging to a soliton.
problem Analyzing the Bach flow on specific nilmanifolds.
method Fourth order geometric flow on four-dimensional simply connected nilmanifolds.
result The Bach flow converges to an expanding Bach soliton on these manifolds.
SGLD proves geometric ergodicity via reflection coupling for nonconvex log-concave distributions.
problem Proving geometric ergodicity of SGLD in nonconvex, log-concave settings.
method Reflection coupling technique to handle SGLD's time discretization and minibatch issues.
result SGLD has an invariant distribution and geometric ergodicity in W 1 W_1 W 1 distance. Proposes KDA to protect deep nets from adversarial attacks.
problem Machine learning system vulnerability to adversarial attacks.
method Key based diversified aggregation with pre-filtering.
result Demonstrates high robustness and universality against various attacks.
With a f-left-invariant Riemannian metric on a Lie group G G G , we mean a Riemannian metric which is conformally equivalent to a left-invariant Riemannian metric, with the conformal factor f f f . In this article, we study the geometry of such metrics and give a necessary and sufficient condition for an f-left-invariant Rie…
New algorithm improves online learning with reduced discretization.
problem Improving adaptive online learning with refined discretization.
method Continuous time approach to online learning, followed by a new discretization argument.
result Optimal regret bound with O ( V T ) O(\sqrt{V_T}) O ( V T ) dependence on gradient variance. SOLO uses DNN to optimize complex topology problems with reduced FEM calculations.
problem Optimizing materials distribution in complex domains with high computational cost.
method Integrates DNN with FEM calculations to learn and substitute objective functions dynamically.
result Optimum predicted by DNN converges to true global optimum through iterations.
The purpose of this article is to study the existence and uniqueness of quasi-Einstein structures on 3 3 3 -dimensional homogeneous Riemannian manifolds. To this end, we use the eight model geometries for 3-dimensional manifolds identified by Thurston. First, we present here a complete description of quasi-Einstein metric…
Blind Descent avoids gradient issues, using a different learning approach.
problem Gradient issues like exploding and vanishing gradients.
method Does not use gradients to guide learning; instead, it is a more fundamental learning process.
result Gradient descent is a specific case of Blind Descent.
E-LDA offers faster, interpretable LDA topic models.
problem Inferring topics in LDA topic models with strong guarantees.
method Non-gradient combinatorial approach for faster convergence.
result Logarithmic parallel computation time and interpretability.
Reparameterizes mirror descent as gradient descent for efficient sparse learning.
problem Efficiently training small sparse networks with mirror descent.
method Develops a framework to convert mirror descent updates into gradient descent updates on different parameters.
result Mirror descent can be reparameterized as gradient descent on modified parameters, facilitating standard backpropagation.
Killing fields on compact m-quasi-Einstein manifolds are shown under specific curvature conditions.
problem Characterizing Killing fields on compact m-quasi-Einstein manifolds.
method Extending a result by Bahuaud-Gunasekaran-Kunduri-Woolgar, the approach involves proving the existence of Killing fields under certain curvature conditions.
result A sufficient condition for a compact, non-gradient m-quasi-Einstein metric to admit a Killing field is provided, extending the original result to the m = -2 case.
The paper introduces various gradient descent algorithms for training deep learning models.
problem Training deep neural networks is challenging due to their complexity.
method Gradient descent and its variants are discussed for optimizing deep learning models.
result Gradient descent and its variants improve the training performance of deep learning models.
Any gradient descent optimization requires to choose a learning rate. With deeper and deeper models, tuning that learning rate can easily become tedious and does not necessarily lead to an ideal convergence. We propose a variation of the gradient descent algorithm in the which the learning rate is not fixed. Instead, w…
Proposes Hebbian-descent for neural network learning, addressing Hebbian and gradient descent issues.
problem Learning issues with correlated data and vanishing error term in gradient descent.
method Integrates Hebbian and gradient descent principles without activation function derivatives, centering neural activities.
result Biologically plausible, convergent, and effective in online learning with correlated data.
Foundation for robust finance using rough path theory.
problem Mathematical models of financial markets under Knightian uncertainty.
method Introducing Property (RIE) for càdlàg paths, proving existence of rough integrals, verifying admissibility of trading strategies.
result Existence and stability of rough path integrals for non-gradient integrands.
Accelerates coordinate descent methods for machine learning problems.
problem Slowness of coordinate descent methods in machine learning.
method Extrapolation-based accelerated coordinate descent.
result Significant speed-up in practice compared to existing methods.
New insights into double descent phenomenon in neural networks.
problem Understanding the double descent behavior in deep learning models.
method Linear teacher-student setup and tools from statistical physics.
result Distinct features are learned at different scales, leading to epoch-wise double descent.
Double descent phenomenon explained in simple terms.
problem Understanding the surprising drop in test error in overparameterized models.
method Informal explanation using linear algebra and probability, visual intuition with polynomial regression, mathematical analysis with ordinary linear regression.
result Three factors create double descent: data undersampling, model size, and parameter count. Ablating any one of these factors prevents double descent.
Gradient descent benefits from tangent kernel advantages under specific conditions.
problem Comparing gradient descent with tangent kernel methods in learning.
method Analysis of gradient descent and tangent kernel methods under different conditions.
result Gradient descent can achieve small error only if tangent kernel methods have a non-trivial advantage, but this advantage can be very small.
The study explores ( m , ρ ) (m,ρ) ( m , ρ ) -quasi-Einstein structures on contact metric manifolds.
problem Exploring ( m , ρ ) (m,ρ) ( m , ρ ) -quasi-Einstein structures in contact geometry. method Proving properties of ( m , ρ ) (m,ρ) ( m , ρ ) -quasi-Einstein structures on contact metric manifolds. result Compact contact or H H H -contact metric manifolds with ( m , ρ ) (m,ρ) ( m , ρ ) -quasi-Einstein structures have specific properties. Proposes a continuous flow model to understand and control instability in gradient descent for deep learning.
problem Understanding and controlling the instability of gradient descent in deep learning.
method Introduces the Principal Flow (PF), a continuous time flow that approximates gradient descent dynamics.
result The PF captures divergent and oscillatory behaviors of gradient descent, including escaping local minima and saddle points.
New method reconstructs non-equilibrium stochastic systems from data.
problem Reconstructing non-equilibrium stochastic systems from ensemble measurements.
method Schrödinger bridge problem with multivariate Ornstein-Uhlenbeck process.
result Simulation-free algorithm achieves higher accuracy than competing methods.
Discovering quasipotential equations from data using machine learning.
problem Understanding escape mechanisms from metastable states in nonlinear systems.
method Combining neural networks and sparse regression to symbolically reconstruct quasipotential equations.
result Model-unbiased analytical forms of quasipotential discovered directly from data.
Unified framework for gradient descent variants in machine learning.
problem Understanding and comparing various gradient descent methods.
method A unified framework interpreting 6 gradient descent variants.
result Some variants coincide under specific conditions.
Gradient descent finds halfspaces with low error for agnostic learning.
problem Agnostic learning of linear halfspaces with convex surrogates.
method Gradient descent on convex surrogates for zero-one loss.
result Gradient descent finds halfspaces with error O ( O P T 1 / 2 + ε ) O(\mathsf{OPT}^{1/2} + \varepsilon) O ( OPT 1/2 + ε ) in poly time and sample complexity. Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.
problem Achieving zero loss minimizers in deep learning networks.
method Analysis of gradient descent algorithm in deep learning, focusing on underparametrized networks.
result Zero loss minimization cannot be achieved generically in deep learning networks.
We propose accelerated randomized coordinate descent algorithms for stochastic optimization and online learning. Our algorithms have significantly less per-iteration complexity than the known accelerated gradient algorithms. The proposed algorithms for online learning have better regret performance than the known rando…
Study shows double and triple descent in unsupervised autoencoders, improving performance in various tasks.
problem Exploring the phenomenon of double descent in unsupervised learning.
method Analytical demonstration and extensive experiments on synthetic and real datasets.
result Over-parameterized unsupervised autoencoders exhibit double and triple descent, enhancing performance in downstream tasks.
This paper explores a new framework for reinforcement learning based on online convex optimization, in particular mirror descent and related algorithms. Mirror descent can be viewed as an enhanced gradient method, particularly suited to minimization of convex functions in highdimensional spaces. Unlike traditional grad…
Gradient descent optimizes deep ReLU networks with proper initialization.
problem Training deep neural networks with ReLU activation.
method Gradient descent and stochastic gradient descent with proper random weight initialization.
result Gradient descent finds global minima for over-parameterized deep ReLU networks.
Kalman Gradient Descent optimizes machine learning models by reducing variance in stochastic optimization.
problem Reducing variance in stochastic gradient descent to improve optimization performance.
method Uses Kalman filtering to adaptively reduce gradient variance in stochastic gradient descent.
result Improved performance on various machine learning tasks including neural networks and black box variational inference.