Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

112224335447 · Jun 202019922001200920172026
48 results for nonlinear conjugate gradient

This manuscript proposes a probabilistic framework for algorithms that iteratively solve unconstrained linear problems Bx=bBx = b with positive definite BB for xx. The goal is to replace the point estimates returned by existing methods with a Gaussian posterior belief over the elements of the inverse of BB, which can …

2014-02-10abs ↗pdf ↗

Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functi…

2017-10-27abs ↗pdf ↗

Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.

problem Comparing statistical properties of different gradient methods in ridge regression.
method Explicit non-standard error decomposition to bound prediction error of conjugate gradient iterates.
result Conjugate gradient iterates share optimality properties with gradient flow and ridge regression up to a constant factor.

We address the challenge of effective exploration while maintaining good performance in policy gradient methods. As a solution, we propose diverse exploration (DE) via conjugate policies. DE learns and deploys a set of conjugate policies which can be conveniently generated as a byproduct of conjugate gradient descent. …

2019-02-10abs ↗pdf ↗

Harmonic functions of two variables are exactly those that admit a conjugate, namely a function whose gradient has the same length and is everywhere orthogonal to the gradient of the original function. We show that there are also partial differential equations controlling the functions of three variables that admit a c…

2012-05-30abs ↗pdf ↗

Bayesian model uses simple functions to forecast macroeconomic data.

problem Forecasting large datasets in macroeconomics with complex nonlinear relationships.
method Sum of simple two-component location mixtures, logistic function threshold, conjugate priors.
result Accurate point and density forecasts in US macroeconomic aggregates.

The computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computa…

2016-02-22abs ↗pdf ↗

Improved Gaussian process regression with tighter log marginal likelihood bounds.

problem Improving predictive performance in Gaussian process regression models.
method Lower bound on log marginal likelihood using conjugate gradients.
result Improved predictive performance compared to other conjugate gradient based approaches.

We would like to congratulate the authors of "A Bayesian Conjugate Gradient Method" on their insightful paper, and welcome this publication which we firmly believe will become a fundamental contribution to the growing field of probabilistic numerical methods and in particular the sub-field of Bayesian numerical methods…

2019-08-08abs ↗pdf ↗

In this work we systematically analyze general properties of differential equations used as machine learning models. We demonstrate that the gradient of the loss function with respect to to the hidden state can be considered as a generalized momentum conjugate to the hidden state, allowing application of the tools of c…

2019-09-09abs ↗pdf ↗

Study of eigenvalues in nonlinear kernels for classification of separable data.

problem Understanding the applicability of linear equivalents in nonlinearly separable data classification.
method Analysis of conjugate kernels and their quadratic equivalents for a canonical nonlinearly separable dataset (XOR problem).
result Identification of regimes where nonlinear kernels deviate from linear equivalents, leading to label-aligned eigenspaces.

The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradi…

2018-02-07abs ↗pdf ↗

New method for identifying autoregressive systems on manifolds.

problem Identifying autoregressive systems on Stiefel and Grassmann manifolds.
method Defining parameters as orthogonal group elements, averaging over observations, conjugate gradient descent on manifolds.
result System parameters can be estimated efficiently using the proposed algorithm.

By restricting the iterate on a nonlinear manifold, the recently proposed Riemannian optimization methods prove to be both efficient and effective in low rank tensor completion problems. However, existing methods fail to exploit the easily accessible side information, due to their format mismatch. Consequently, there i…

2016-11-12abs ↗pdf ↗

Regularized least-squares (kernel-ridge / Gaussian process) regression is a fundamental algorithm of statistics and machine learning. Because generic algorithms for the exact solution have cubic complexity in the number of datapoints, large datasets require to resort to approximations. In this work, the computation of …

2019-11-14abs ↗pdf ↗

We propose a conjugate gradient type optimization technique for the computation of the Karcher mean on the set of complex linear subspaces of fixed dimension, modeled by the so-called Grassmannian. The identification of the Grassmannian with Hermitian projection matrices allows an accessible introduction of the geometr…

2012-09-14abs ↗pdf ↗

NeuralIF uses neural networks to improve preconditioning for faster CG convergence.

problem Improving convergence of conjugate gradient method for large-scale sparse systems.
method Data-driven approach using graph neural networks to generate incomplete factorization.
result Data-driven preconditioners accelerate convergence of conjugate gradient method.

Let (M,g(t))(M,g(t)), 0tT0\le t\le T, Mφ\partial M\neφ, be a compact nn-dimensional manifold, n2n\ge 2, with metric g(t)g(t) evolving by the Ricci flow such that the second fundamental form of M\partial M with respect to the unit outward normal of M\partial M is uniformly bounded below on M×[0,T]\partial M\times [0,T]. We will pr…

2008-01-23abs ↗pdf ↗

Improves understanding of stochastic NGVI convergence rates.

problem Lack of knowledge about non-asymptotic convergence rates in stochastic NGVI.
method Proved non-asymptotic convergence rates for conjugate likelihoods and showed implicit optimization for non-conjugate likelihoods.
result First O(1T)\mathcal{O}(\frac{1}{T}) non-asymptotic convergence rate for stochastic NGVI in conjugate likelihoods.

New recommendations improve Gaussian process accuracy and stability.

problem Numerical instabilities and poor test likelihoods in iterative Gaussian process learning.
method Investigated CG tolerance, preconditioner rank, and Lanczos decomposition rank. Recommended small CG tolerance and large root decomposition size.
result L-BFGS-B optimizer achieves convergence with fewer gradient updates, improving Gaussian process accuracy.

We present a general method for deriving collapsed variational inference algo- rithms for probabilistic models in the conjugate exponential family. Our method unifies many existing approaches to collapsed variational inference. Our collapsed variational inference leads to a new lower bound on the marginal likelihood. W…

2012-06-22abs ↗pdf ↗

The paper introduces a new framework to understand and optimize deep neural networks.

problem Understanding and optimizing the trainability and generalization of deep neural networks.
method Developed a conjugate learning theoretical framework based on convex conjugate duality.
result Demonstrated that training deep neural networks with SGD achieves global optima of empirical risk.

Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations.

problem Understanding moduli of continuity for fully nonlinear parabolic equations.
method Proving moduli of continuity of viscosity solutions are subsolutions of one-dimensional parabolic equations.
result Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations with bounded initial data.

The paper proves gradient estimates for nonlinear parabolic equations on smooth metric measure spaces.

problem Proving gradient estimates for nonlinear parabolic equations on smooth metric measure spaces.
method Using Souplet-Zhang type estimates and properties of Bakry-Emery Ricci tensor and weighted mean curvature.
result Gradient estimates for nonlinear parabolic equations on smooth metric measure spaces with Dirichlet boundary condition.

The paper provides gradient estimates for nonlinear heat-type equations on smooth metric measure spaces.

problem Proving gradient estimates for nonlinear heat-type equations on smooth metric measure spaces.
method Using Hamilton type and Li-Yau type estimates, the paper proves gradient estimates on positive solutions to generalized nonlinear parabolic equations on smooth metric measure spaces with compact boundary.
result Gradient estimates for nonlinear heat-type equations on smooth metric measure spaces.

We prove Gaussian type bounds for the fundamental solution of the conjugate heat equation evolving under the Ricci flow. As a consequence, for dimension 4 and higher, we show that the backward limit of type I κκ-solutions of the Ricci flow must be a non-flat gradient shrinking Ricci soliton. This extends Perelman's pr…

2010-06-03abs ↗pdf ↗

We propose and study kernel conjugate gradient methods (KCGM) with random projections for least-squares regression over a separable Hilbert space. Considering two types of random projections generated by randomized sketches and Nyström subsampling, we prove optimal statistical results with respect to variants of norms …

2018-11-05abs ↗pdf ↗

In this work, we consider the use of model-driven deep learning techniques for massive multiple-input multiple-output (MIMO) detection. Compared with conventional MIMO systems, massive MIMO promises improved spectral efficiency, coverage and range. Unfortunately, these benefits are coming at the cost of significantly i…

2019-06-10abs ↗pdf ↗

Let (M,g(t))(M,g(t)) be a solution to the Ricci flow on a closed Riemannian manifold. In this paper, we prove differential Harnack inequalities for positive solutions of nonlinear parabolic equations of the type $$\ppt f=Δf-f \ln f +Rf.$$ We also comment on an earlier result of the first author on positive solutions of the c…

2010-01-28abs ↗pdf ↗