Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified analysis of conjugate gradients and accelerated methods using duality gap.
NCG methods improve shape optimization efficiency.
The natural gradient method has been used effectively in conjugate Gaussian process models, but the non-conjugate case has been largely unexplored. We examine how natural gradients can be used in non-conjugate stochastic settings, together with hyperparameter learning. We conclude that the natural gradient can signific…
We address the challenge of effective exploration while maintaining good performance in policy gradient methods. As a solution, we propose diverse exploration (DE) via conjugate policies. DE learns and deploys a set of conjugate policies which can be conveniently generated as a byproduct of conjugate gradient descent. …
A new method speeds up deep neural network training.
Proposes an extension of BayesCG for solving multiple linear systems.
Improved kernel ridge regression using conjugate gradients.
Improved Gaussian process regression with tighter log marginal likelihood bounds.
This manuscript proposes a probabilistic framework for algorithms that iteratively solve unconstrained linear problems with positive definite for . The goal is to replace the point estimates returned by existing methods with a Gaussian posterior belief over the elements of the inverse of , which can …
Conjugate gradient methods improve efficiency for high-dimensional GLMMs.
A method uses CG to create efficient channels for ideal observers.
The study proves triviality and rigidity results for Ricci solitons and estimates their conjugate radius.
Harmonic functions of two variables are exactly those that admit a conjugate, namely a function whose gradient has the same length and is everywhere orthogonal to the gradient of the original function. We show that there are also partial differential equations controlling the functions of three variables that admit a c…
A new method reduces the cost of solving large-scale linear models.
NeuralIF uses neural networks to improve preconditioning for faster CG convergence.
We propose a conjugate gradient type optimization technique for the computation of the Karcher mean on the set of complex linear subspaces of fixed dimension, modeled by the so-called Grassmannian. The identification of the Grassmannian with Hermitian projection matrices allows an accessible introduction of the geometr…
New method for identifying autoregressive systems on manifolds.
The computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computa…
Deep learning designs effective preconditioners for water engineering problems.
New recommendations improve Gaussian process accuracy and stability.
We present a general method for deriving collapsed variational inference algo- rithms for probabilistic models in the conjugate exponential family. Our method unifies many existing approaches to collapsed variational inference. Our collapsed variational inference leads to a new lower bound on the marginal likelihood. W…
Nowadays stochastic approximation methods are one of the major research direction to deal with the large-scale machine learning problems. From stochastic first order methods, now the focus is shifting to stochastic second order methods due to their faster convergence and availability of computing resources. In this pap…
Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functi…
Nonlinear conjugate gradient (NLCG) based optimizers have shown superior loss convergence properties compared to gradient descent based optimizers for traditional optimization problems. However, in Deep Neural Network (DNN) training, the dominant optimization algorithm of choice is still Stochastic Gradient Descent (SG…
We propose and study kernel conjugate gradient methods (KCGM) with random projections for least-squares regression over a separable Hilbert space. Considering two types of random projections generated by randomized sketches and Nyström subsampling, we prove optimal statistical results with respect to variants of norms …
In this paper we introduce a parameter dependent class of Krylov-based methods, namely CD, for the solution of symmetric linear systems. We give evidence that in our proposal we generate sequences of conjugate directions, extending some properties of the standard Conjugate Gradient (CG) method, in order to preserve the…
Improves understanding of stochastic NGVI convergence rates.
New method improves calibration of BayesCG for better uncertainty quantification.
BONG optimizes Bayesian inference online with natural gradient descent.
The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradi…
Two methods solve kernel ridge regression problems efficiently.
A new iterative K-FAC algorithm reduces training time and memory usage.
This work proves convergence of adaptive resampling for random Fourier features.
Recent advances in policy gradient methods and deep learning have demonstrated their applicability for complex reinforcement learning problems. However, the variance of the performance gradient estimates obtained from the simulation is often excessive, leading to poor sample efficiency. In this paper, we apply the stoc…
New method uses robust estimators for Newton's method in empirical risk minimization.
New method tackles high-dimensional SBL without covariance matrices.
The extragradient method accelerates convergence in complex game dynamics.
A new network reduces MIMO detection complexity.
Proves FR-NGD optimally approximates evolutionary dynamics and continuous Bayesian inference.
The paper analyzes contraction rates for GP regression approximations.
The paper improves sparse Gaussian processes by optimizing predictive loss.
Let , , , be a compact -dimensional manifold, , with metric evolving by the Ricci flow such that the second fundamental form of with respect to the unit outward normal of is uniformly bounded below on . We will pr…
Novel probabilistic solver speeds up solving related linear systems.
Several recent works have explored stochastic gradient methods for variational inference that exploit the geometry of the variational-parameter space. However, the theoretical properties of these methods are not well-understood and these methods typically only apply to conditionally-conjugate models. We present a new s…
We establish a point-wise gradient estimate for positive solutions of the conjugate heat equation. This contrasts to Perelman's point-wise gradient estimate which works mainly for the fundamental solution rather than all solutions. Like Perelman's estimate, the most general form of our gradient estimate does not …
New algorithm selects best preconditioner for iterative methods.
We prove statistical rates of convergence for kernel-based least squares regression from i.i.d. data using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is related to Kernel Partial Least Squares, a regression method that combines supervised dimensio…