NCG methods improve shape optimization efficiency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Nonlinear conjugate gradient (NLCG) based optimizers have shown superior loss convergence properties compared to gradient descent based optimizers for traditional optimization problems. However, in Deep Neural Network (DNN) training, the dominant optimization algorithm of choice is still Stochastic Gradient Descent (SG…
This manuscript proposes a probabilistic framework for algorithms that iteratively solve unconstrained linear problems with positive definite for . The goal is to replace the point estimates returned by existing methods with a Gaussian posterior belief over the elements of the inverse of , which can …
Generalizes Fenchel conjugation to nonlinear functions on arbitrary sets.
Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functi…
Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.
The natural gradient method has been used effectively in conjugate Gaussian process models, but the non-conjugate case has been largely unexplored. We examine how natural gradients can be used in non-conjugate stochastic settings, together with hyperparameter learning. We conclude that the natural gradient can signific…
We address the challenge of effective exploration while maintaining good performance in policy gradient methods. As a solution, we propose diverse exploration (DE) via conjugate policies. DE learns and deploys a set of conjugate policies which can be conveniently generated as a byproduct of conjugate gradient descent. …
The study proves triviality and rigidity results for Ricci solitons and estimates their conjugate radius.
Harmonic functions of two variables are exactly those that admit a conjugate, namely a function whose gradient has the same length and is everywhere orthogonal to the gradient of the original function. We show that there are also partial differential equations controlling the functions of three variables that admit a c…
We propose a novel Riemannian manifold preconditioning approach for the tensor completion problem with rank constraint. A novel Riemannian metric or inner product is proposed that exploits the least-squares structure of the cost function and takes into account the structured symmetry that exists in Tucker decomposition…
A new method speeds up deep neural network training.
Bayesian model uses simple functions to forecast macroeconomic data.
The computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computa…
Improved Gaussian process regression with tighter log marginal likelihood bounds.
This note provides a novel, simple analysis of the method of conjugate gradients for the minimization of convex quadratic functions. In contrast with standard arguments, our proof is entirely self-contained and does not rely on the existence of Chebyshev polynomials. Another advantage of our development is that it clar…
We would like to congratulate the authors of "A Bayesian Conjugate Gradient Method" on their insightful paper, and welcome this publication which we firmly believe will become a fundamental contribution to the growing field of probabilistic numerical methods and in particular the sub-field of Bayesian numerical methods…
In this work we systematically analyze general properties of differential equations used as machine learning models. We demonstrate that the gradient of the loss function with respect to to the hidden state can be considered as a generalized momentum conjugate to the hidden state, allowing application of the tools of c…
Study of eigenvalues in nonlinear kernels for classification of separable data.
Existence of a conjugate point in the incompressible Euler flow on a sphere and an ellipsoid is considered. Misiolek (1996) formulated a differential-geometric criterion (we call M-criterion) for the existence of a conjugate point in a fluid flow. In this paper, it is shown that no zonal flow (stationary Euler flow) sa…
A method uses CG to create efficient channels for ideal observers.
Conjugate gradient methods improve efficiency for high-dimensional GLMMs.
The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradi…
New method for identifying autoregressive systems on manifolds.
By restricting the iterate on a nonlinear manifold, the recently proposed Riemannian optimization methods prove to be both efficient and effective in low rank tensor completion problems. However, existing methods fail to exploit the easily accessible side information, due to their format mismatch. Consequently, there i…
Regularized least-squares (kernel-ridge / Gaussian process) regression is a fundamental algorithm of statistics and machine learning. Because generic algorithms for the exact solution have cubic complexity in the number of datapoints, large datasets require to resort to approximations. In this work, the computation of …
We propose a conjugate gradient type optimization technique for the computation of the Karcher mean on the set of complex linear subspaces of fixed dimension, modeled by the so-called Grassmannian. The identification of the Grassmannian with Hermitian projection matrices allows an accessible introduction of the geometr…
NeuralIF uses neural networks to improve preconditioning for faster CG convergence.
Let , , , be a compact -dimensional manifold, , with metric evolving by the Ricci flow such that the second fundamental form of with respect to the unit outward normal of is uniformly bounded below on . We will pr…
We establish a point-wise gradient estimate for positive solutions of the conjugate heat equation. This contrasts to Perelman's point-wise gradient estimate which works mainly for the fundamental solution rather than all solutions. Like Perelman's estimate, the most general form of our gradient estimate does not …
Improves understanding of stochastic NGVI convergence rates.
New recommendations improve Gaussian process accuracy and stability.
We present a general method for deriving collapsed variational inference algo- rithms for probabilistic models in the conjugate exponential family. Our method unifies many existing approaches to collapsed variational inference. Our collapsed variational inference leads to a new lower bound on the marginal likelihood. W…
The paper introduces a new framework to understand and optimize deep neural networks.
Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations.
The paper proves gradient estimates for nonlinear parabolic equations on smooth metric measure spaces.
The paper provides gradient estimates for nonlinear heat-type equations on smooth metric measure spaces.
We prove Gaussian type bounds for the fundamental solution of the conjugate heat equation evolving under the Ricci flow. As a consequence, for dimension 4 and higher, we show that the backward limit of type I -solutions of the Ricci flow must be a non-flat gradient shrinking Ricci soliton. This extends Perelman's pr…
We propose and study kernel conjugate gradient methods (KCGM) with random projections for least-squares regression over a separable Hilbert space. Considering two types of random projections generated by randomized sketches and Nyström subsampling, we prove optimal statistical results with respect to variants of norms …
In this paper, we study elliptic gradient estimates for a nonlinear -heat equation, which is related to the gradient Ricci soliton and the weighted log-Sobolev constant of smooth metric measure spaces. Precisely, we obtain Hamilton's and Souplet-Zhang's gradient estimates for positive solutions to the nonlinear -…
We propose a novel Bayesian approach to solve stochastic optimization problems that involve finding extrema of noisy, nonlinear functions. Previous work has focused on representing possible functions explicitly, which leads to a two-step procedure of first, doing inference over the function space and second, finding th…
Gradient descent and SGD solve nonlinear inverse problems efficiently.
In this work, we consider the use of model-driven deep learning techniques for massive multiple-input multiple-output (MIMO) detection. Compared with conventional MIMO systems, massive MIMO promises improved spectral efficiency, coverage and range. Unfortunately, these benefits are coming at the cost of significantly i…
SING improves state inference in latent SDE models for better drift function estimation.
We prove certain localized and global differential Harnack inequality for all positive solutions to the geometric conjugate heat equation coupled to the forward in time Ricci flow. In this case, the diffusion operator is perturbed with the curvature operator, precisely, the Laplace-Beltrami operator is replaced with "$…
In this paper we introduce a parameter dependent class of Krylov-based methods, namely CD, for the solution of symmetric linear systems. We give evidence that in our proposal we generate sequences of conjugate directions, extending some properties of the standard Conjugate Gradient (CG) method, in order to preserve the…
Nowadays stochastic approximation methods are one of the major research direction to deal with the large-scale machine learning problems. From stochastic first order methods, now the focus is shifting to stochastic second order methods due to their faster convergence and availability of computing resources. In this pap…
Let be a solution to the Ricci flow on a closed Riemannian manifold. In this paper, we prove differential Harnack inequalities for positive solutions of nonlinear parabolic equations of the type $$\ppt f=Δf-f \ln f +Rf.$$ We also comment on an earlier result of the first author on positive solutions of the c…