Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

6.2%12.4%18.7%24.9% · Jun 202019922001200920182026
48 results for matrix stochastic gradient

An ADRC-incorporated SGD algorithm improves latent factor analysis speed and accuracy.

problem Slow convergence in standard SGD for HDI matrix analysis.
method Incorporates ADRC principles to refine historical and future learning error states.
result Empirically outperforms state-of-the-art LFA models in HDI matrix prediction.

Efficient algorithm estimates low-rank matrices from noisy measurements.

problem Estimating low-rank matrices from noisy linear measurements.
method Stochastic variance-reduced gradient descent algorithm.
result Algorithm converges to the unknown low-rank matrix at a linear rate up to the minimax optimal statistical error.

New algorithms optimize matrix manifolds, converging faster than existing methods.

problem Optimizing on Riemannian matrix manifolds with constraints.
method Adaptive stochastic gradient algorithms for row and column subspaces.
result Converges faster with rate O(log(T)/T)\mathcal{O}(\log (T)/\sqrt{T}).

Paper shows robustness of gradient descent in matrix sensing despite perturbations.

problem Understanding robustness of gradient descent in matrix sensing.
method Developed perturbed gradient flow to capture noise and improve robustness.
result Gradient descent is robust to perturbations in matrix sensing.

We propose a fast algorithm for spectral embedding using stochastic gradient descent.

problem Scalability issue in spectral embedding due to eigendecomposition bottleneck.
method Reformulate spectral embedding as a stochastic optimization problem, replacing orthogonality constraint with an orthogonalization matrix.
result Efficient algorithm based on mini-batch gradient descent that outperforms existing techniques in execution speed.

Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2…

2014-12-03abs ↗pdf ↗

Paper proposes an online estimator for covariance matrix of SGD iterates.

problem Quantifying variability and randomness of SGD-based estimates in online learning.
method Proposes a fully online estimator for covariance matrix of ASGD using SGD iterates.
result Establishes consistency of the online estimator and shows comparable convergence rate to offline methods.

Paper develops a method to construct confidence regions for model parameters using batch means method.

problem Constructing confidence regions for model parameters in stochastic gradient descent.
method Batch means method to cancel out covariance matrix, using Polyak-Ruppert averaging.
result Established process-level functional central limit theorem for stochastic gradient descent estimators.

SGD with mini-batches can solve convex low-rank matrix problems efficiently.

problem Solving large-scale convex low-rank matrix problems efficiently.
method Stochastic Gradient Descent with mini-batches and low-rank projections.
result SGD with mini-batches produces low-rank iterates with high probability.

Unified derivation of high-dimensional linear models using stochastic gradient descent.

problem Performance analysis of high-dimensional linear models trained with stochastic gradient descent.
method Derivation of a deterministic equivalence for the two-point function of a random matrix resolvent.
result Unified understanding of model performance including previously known and novel results.

NSGLD improves SGLD for non-convex optimization problems.

problem Optimizing non-convex objectives efficiently.
method Introducing non-reversible SGLD by adding an anti-symmetric matrix to the drift term of the Langevin diffusion.
result NSGLD converges faster to the same stationary distribution with non-asymptotic guarantees.

We study PCA as a stochastic optimization problem and propose a novel stochastic approximation algorithm which we refer to as "Matrix Stochastic Gradient" (MSG), as well as a practical variant, Capped MSG. We study the method both theoretically and empirically.

2013-07-05abs ↗pdf ↗

The paper simplifies the Fisher information matrix for random deep networks, speeding up learning.

problem Learning deep neural networks efficiently with large parameter spaces.
method Statistical neurodynamical method to reveal Fisher information properties, proving unit-wise block diagonal structure and explicit inverse.
result Explicit natural gradient formula without matrix inversion, speeding up learning.

Stochastic gradient descent regularizes least squares problems by smoothing large singular values.

problem Regularization of least squares problems using stochastic gradient descent.
method Analysis of stochastic gradient descent applied to least squares problems, showing a regularization effect.
result Stochastic gradient descent leads to a quick regularization effect, smoothing large singular values.

SGD's training dynamics align with Hessian and gradient spectra in high-dimensional classification tasks.

problem Understanding the spectra of Hessian and gradient matrices in high-dimensional classification tasks.
method Rigorous analysis of SGD dynamics and spectra of Hessian and gradient matrices.
result SGD trajectory and emergent outlier eigenspaces align with a common low-dimensional subspace in multi-class high-dimensional mixtures and neural networks.

New framework shows all local minima are globally optimal in non-convex low-rank problems.

problem Non-convex low-rank problems, including matrix sensing, completion, and robust PCA.
method Developed a new framework to analyze the optimization landscapes of these problems.
result All local minima are also globally optimal and no high-order saddle points exist.

A new method for machine learning updates reduces complexity and improves robustness.

problem Stochastic gradient updates are inefficient and sensitive to feature scaling.
method Incremental Gauss-Newton Descent (IGND) reduces the need for matrix operations and improves robustness.
result IGND improves robustness to sensitivity scaling and can be competitive with common stochastic optimizers.

Vanilla SGD learns SIM from anisotropic data without explicit covariance estimation.

problem Learning SIM from anisotropic Gaussian inputs.
method Vanilla Stochastic Gradient Descent (SGD) trained on SIM with anisotropic input.
result Vanilla SGD adapts to anisotropic data's covariance structure.

Study shows how SGD's implicit regularization relates to ridge regression.

problem Least squares regression optimization with mini-batch SGD.
method Analyzes stochastic gradient flow as a continuous-time model of SGD.
result Bound on excess risk of SGD flow over ridge regression, revealing how parameters drive risk.

New insights into learning rates and batch sizes for neural networks using random matrix theory.

problem Understanding how batch size affects learning rates in neural networks.
method Random matrix theory applied to spiked, field-dependent random matrices.
result Analytical expressions for maximal learning rates as a function of batch size.

New method uses Hessian to track gradients, improving variance reducing stochastic methods.

problem Improving variance reducing stochastic methods for faster convergence.
method Proposes a modified SVRG method using the Hessian for better control variates and accurate approximations.
result Demonstrates faster theoretical convergence and effectiveness on various problems.

We consider the problem of learning a Gaussian variational approximation to the posterior distribution for a high-dimensional parameter, where we impose sparsity in the precision matrix to reflect appropriate conditional independence structure in the model. Incorporating sparsity in the precision matrix allows the Gaus…

2016-05-18abs ↗pdf ↗

A new algorithm reduces variance in Riemannian stochastic quasi-Newton methods.

problem Minimizing the average of many loss functions on Riemannian manifolds.
method R-SQN-VR algorithm with variance reduction for non-convex and retraction-convex functions.
result The algorithm outperforms existing methods on manifold computations.

The study improves generalization in large-batch training by adding structured covariance noise to gradients.

problem Improving generalization in large-batch training while maintaining optimal convergence.
method Adding covariance noise to the gradients to improve generalization performance.
result The method improves generalization performance without degrading optimization performance and training duration.

Exploring how noise and curvature affect optimization and generalization.

problem The interaction between noise and curvature in optimization and generalization.
method Analyzing the speed of minimizing expected loss with stochastic methods, distinguishing between Fisher, Hessian, and gradient covariance matrices.
result Clarifying the role of curvature and noise in estimating the generalization gap.

RES, a regularized stochastic version of the Broyden-Fletcher-Goldfarb-Shanno (BFGS) quasi-Newton method is proposed to solve convex optimization problems with stochastic objectives. The use of stochastic gradient descent algorithms is widespread, but the number of iterations required to approximate optimal arguments c…

2014-01-29abs ↗pdf ↗

Researchers interpret SGD using diffusion metrics for clearer geometric understanding.

problem Elusiveness of geometrical significance in stochastic gradient descent.
method Study a deterministic model with geodesics of diffusion metrics.
result Establishes parallel with General Relativity models.

Stochastic gradient descent optimizes Nyström samples for kernel matrix approximation.

problem Optimizing Nyström samples for kernel matrix approximation.
method Stochastic gradient descent applied to multisets of landmark points (Nyström samples) using a surrogate criterion (radial SKD).
result Local minimization of the radial SKD yields improved Nyström approximation accuracy.

NeuralIF uses neural networks to improve preconditioning for faster CG convergence.

problem Improving convergence of conjugate gradient method for large-scale sparse systems.
method Data-driven approach using graph neural networks to generate incomplete factorization.
result Data-driven preconditioners accelerate convergence of conjugate gradient method.

New method for zeroth-order stochastic gradient algorithms provides confidence intervals.

problem Lack of inferential capabilities for zeroth-order stochastic gradient algorithms.
method Established central limit theorem and provided online estimators for asymptotic covariance matrix.
result Asymptotically valid confidence sets for parameter estimation and prediction.

New stochastic gradient descent with random search directions improves efficiency and convergence.

problem Efficiency and convergence of stochastic gradient descent methods.
method Developed a new class of stochastic gradient descent algorithms with random search directions.
result Established almost sure convergence and provided Lp\mathbb{L}^p rates of convergence.

We use matricial free energy to regularize autoencoders, producing Gaussian-like codes.

problem Generating Gaussian-like codes for autoencoders.
method Define a differentiable loss function based on singular values of the code matrix, minimizing matricial free energy.
result Minimizing matricial free energy results in Gaussian-like codes that generalize.