Three methods for tuning HMC diagonal scale matrices compared.
problem Improving Hamiltonian Monte Carlo efficiency with diagonal scale matrices.
method Three approaches: ISG, median crossing frequency, and estimated marginal standard deviations.
result ISG method leads to more efficient sampling in many cases.
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
problem Scaling Gaussian processes for large datasets.
method Structured diagonal scaling matrix and Power-EP framework.
result Structured approximations improve performance without increasing computational cost.
This paper improves linear system solving by optimizing matrix diagonal scaling.
problem Improving the condition number of a matrix for faster iterative methods.
method Left or right diagonal rescaling of the matrix A, with new bounds and algorithms.
result Jacobi preconditioning reduces A's condition number to within a quadratic factor of the best possible scaling.
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
problem Optimizing diagonal preconditioning to reduce matrix condition numbers.
method Reformulated as a quasi-convex problem, solved with bisection and Newton updates.
result Optimal diagonal preconditioners can significantly improve iterative methods.
A new method solves diagonally constrained SDPs quickly and accurately.
problem Solving large-scale diagonally constrained SDPs efficiently.
method Combines momentum from convex optimization with coordinate descent and matrix factorization.
result Local linear convergence and first-order critical point convergence proved.
New adaptive and accelerated SGD methods achieve optimal convergence rates.
problem Optimizing convergence rates of stochastic gradient descent methods.
method Integrates diagonal scaling and momentum into accelerated SGD.
result Achieves optimal sampling and iteration complexity for smooth stochastic optimization.
Gradient methods work well on overparameterized diagonal linear networks.
problem Understanding why gradient-based methods work well in overparameterized models.
method Study of Deep Diagonal Linear Networks with gradient flow analysis.
result Gradient flow on layer parameters induces a mirror-flow dynamic in the effective parameter space, leading to explicit convergence guarantees.
Unified framework for scale-invariant representation learning using MAPCA.
problem Learning invariant representations in data.
method Metric-Aware Principal Component Analysis (MAPCA) based on generalized eigenproblem.
result MAPCA provides a unified geometric language for various self-supervised learning objectives.
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank represent…
New unitary RNN architecture using complex Cayley transform outperforms existing methods.
problem Vanishing or exploding gradient problem in RNNs.
method Developed a unitary RNN architecture based on a complex scaled Cayley transform.
result scuRNN achieves comparable or better results than existing unitary RNNs.
This article concerns new off-diagonal estimates on the remainder and its derivatives in the pointwise Weyl law on a compact n-dimensional Riemannian manifold. As an application, we prove that near any non self-focal point, the scaling limit of the spectral projector of the Laplacian onto frequency windows of constant …
Recurrent Neural Networks (RNNs) are powerful models that achieve exceptional performance on several pattern recognition problems. However, the training of RNNs is a computationally difficult task owing to the well-known "vanishing/exploding" gradient problem. Algorithms proposed for training RNNs either exploit no (or…
Study on special symmetries in biwarped product 3-manifolds.
problem Characterizing Killing vector fields on biwarped product-type 3-manifolds.
method Derived system of equations for Killing fields and described their structure.
result Families of solutions found, including explicit examples.
Sharp pseudospectral bounds prevent transient amplification in coupled gradient descent.
problem Transient amplification in coupled gradient descent systems.
method Developed a sharp pseudospectral theory for block-triangular Jacobians, proving Kreiss constant bounds and matching minimax lower bounds.
result Obtained a finite-horizon iteration-complexity bound of O(K(J)2log(1/δ)) for stochastic coupled descent. We attempt to unveil the fine structure of volatility feedback effects in the context of general quadratic autoregressive (QARCH) models, which assume that today's volatility can be expressed as a general quadratic form of the past daily returns. The standard ARCH or GARCH framework is recovered when the quadratic kern…
Study reveals how initialization scale affects training accuracy in linear networks.
problem Understanding implicit bias in linear classification models.
method Asymptotic analysis of gradient flow trajectories and training loss minimization.
result Implicit bias is more complex at reasonable initialization scales and training accuracies.
This paper solves matrix blind joint block diagonalization with noise.
problem Identifying the diagonalizer and block diagonal structure of matrices under noise.
method Bi-block diagonalization method.
result The method can identify the exact solution under certain conditions.
Diagonalizes metrics of 3D Lorentzian manifolds.
problem Diagonalizing metrics of 3D Lorentzian manifolds.
method Applying the technique of moving frames.
result Every smooth Lorentzian 3-manifold admits an atlas with a diagonal metric.
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence. But they are rarely applied to deep learning in practice because of high computational cost and th…
Smooth manifolds have been always understood intuitively as spaces with an affine geometry on the infinitesimal scale. In Synthetic Differential Geometry this can be made precise by showing that a smooth manifold carries a natural structure of an infinitesimally affine space. This structure is comprised of two pieces o…
Study Ricci vector fields on 2D space with diagonal metrics.
problem Understanding Ricci vector fields on 2D space with specific metrics.
method Examined Ricci vector fields on R2 with a diagonal metric. result Characterized Ricci vector fields on R2 with a diagonal metric. Develops a novel stochastic algorithm for diagonal estimation of large matrices.
problem Efficient diagonal estimation for large or implicit matrices.
method Adaptive parameter selection in a stochastic algorithm.
result Lower bound on random query vectors needed for estimation.
Octagon map accelerates diagonal changes algorithm.
problem Improving the efficiency of diagonal changes algorithm.
method Octagon Farey map as an acceleration.
result Octagon map accelerates diagonal changes algorithm.
Unified analysis of parameter norms in overparameterized linear models, revealing scaling laws and thresholds.
problem Understanding the scaling of parameter norms in overparameterized linear models.
method Simple dual-ray analysis revealing competition between signal spike and bulk of null coordinates.
result Unified closed-form predictions for parameter norm scaling, including elbow and threshold laws.
Study finds symmetries in a special 3D space with a diagonal metric.
problem Identifying symmetries in a specific 3D space.
method Determining Killing vector fields on a diagonal metric in R3. result Killing vector fields on the space R3 with a diagonal metric have been identified. New diagonal knots found with non-torus structure.
problem Identifying knots with diagonal grid diagrams.
method Analysis of knots represented by diagonal grid diagrams.
result All diagonal knots are positive, and a new non-torus example is found.
We use mathematical induction to prove that the horizontal composition in the class of coherently diagonal complexes is indeed a binary operation. That is to say, the embedding of two coherently diagonal complexes in an alternating planar diagram produces a coherently diagonal complex.
Most known four-dimensional cohomogeneity-one Einstein metrics are diagonal in the basis defined by the left-invariant one-forms, though some essentially non-diagonal ones are known. We consider the problem of explicitly seeking non-diagonal Einstein metrics, and we find solutions which in some cases exhaust the possib…
New method improves Pham's algorithm for joint diagonalization.
problem Optimizing joint diagonalization of matrices for statistical learning.
method Quasi-Newton method for Pham's diagonalization criterion.
result Proposed method outperforms Pham's algorithm in experiments.
Equal diagonal energies proven on Liouville surfaces.
problem Diagonal energies on Liouville surfaces.
method Analyzing parameter curves and rectangles on Liouville surfaces.
result Diagonal energies are equal in n-dimensional Liouville manifolds.
Efficiently approximates Sparse PCA with significant speedups and minor error.
problem Sparse Principal Component Analysis (Sparse PCA) is NP-hard and computationally expensive.
method Approximates the covariance matrix with block-diagonal form, solves sub-problems in each block, and reconstructs the solution.
result Significant computational speedups with minor additive error.
Diagonal linear networks converge to lasso regularization path during training.
problem Understanding the regularization behavior of diagonal linear networks.
method Analyzing the training trajectory of diagonal linear networks and comparing it to the lasso regularization path.
result The training trajectory of diagonal linear networks is closely related to the lasso regularization path.
We show that a basis of a semisimple Lie algebra of compact type, for which any diagonal left-invariant metric has a diagonal Ricci tensor, is characterized by the Lie algebraic condition of being "nice". Namely, the bracket of any two basis elements is a multiple of another basis element. This extends the work of Laur…
New diagonal move simplifies knots and links efficiently.
problem Efficiently unknotting knots and links.
method Introduces diagonal move, proving its effectiveness for classical and welded knots.
result Diagonal move reduces any knot or link to the unknot or unlink with fewer operations.
Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
problem Modeling non-Gaussian time-series with stationary increments.
method Complex wavelet transform for scale variations, joint correlation matrix for scale dependencies, second wavelet transform for diagonalization, maximum entropy models conditioned by scattering spectra coefficients.
result Scattering spectra of self-similar processes are scale invariant, allowing statistical testing and generation of new time-series.
Study uncovers scaling laws and spectral properties of shallow neural networks.
problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.
We obtain the natural diagonal almost product and locally product structures on the total space of the cotangent bundle of a Riemannian manifold. We find the Riemannian almost product (locally product) and the (almost) para-Hermitian cotangent bundles of natural diagonal lift type. We prove the characterization theorem…
Improved neural network inference with eigenvalue correction.
problem Inference of flexible variational posteriors is computationally expensive.
method Eigenvalue correction to matrix-variate Gaussian posterior.
result Empirically, the method outperforms existing algorithms.
Study grid homology of diagonal knots, finding key terms related to prime factors and decompositions.
problem Determine grid homology of diagonal knots and compare them to other knot types.
method Use grid diagrams and combinatorial knot Floer homology to analyze diagonal knots.
result Grid homology detects the number of prime factors and decompositions of the knot into non-integer tangles.
Gradient descent optimally trains RNNs without overparameterization.
problem Training recurrent neural networks (RNNs) with gradient descent.
method Nonasymptotic analysis of gradient descent for RNNs with diagonal weight matrices.
result Gradient descent can achieve optimality in RNNs with a network size scaling logarithmically with the number of samples.
Study on stability of non-diagonal Einstein metrics on specific homogeneous spaces.
problem Stability analysis of non-diagonal Einstein metrics on HimesH/ΔK. method Formula for scalar curvature, study of stability with Hilbert action.
result Non-diagonal Einstein metrics on M are unstable with different coindexes. The author connects Poincaré embeddings to Reidemeister traces and diagonal maps.
problem Existence of Poincaré embeddings for specific spaces.
method Relates total obstruction to Reidemeister trace and uses Poincaré duality.
result Diagonal maps admit Poincaré embeddings under certain conditions.
We consider a general Hermitian holomorphic line bundle L on a compact complex manifold M and let □pq be the Kodaira Laplacian on (0,q) forms with values in Lp. The main result is a complete asymptotic expansion for the semi-classically scaled heat kernel exp(−u□pq/p)(x,x) along the diagonal…
Conditions for flat 3-manifolds with diagonal metrics are identified.
problem Characterizing flat 3-manifolds with diagonal metrics.
method Provided necessary and sufficient conditions for flatness.
result Characterized flat manifolds of warped product-type.
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
problem Diagonalizing the Toda flow on matrices with simple spectrum.
method Lie theoretic methods applied to complex semisimple Lie algebras and their real forms.
result Decouples the Toda vector field into simpler components.
New phase harmonic covariance models capture non-Gaussian properties of stationary processes.
problem Capturing non-Gaussian properties of stationary processes using Fourier phase.
method Introduce phase harmonic covariance moments and maximum entropy models conditioned by these moments.
result Maximum entropy models from phase harmonic covariances improve image synthesis of turbulent flows.
Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
problem Subspace clustering with block diagonal structure for noisy data.
method ABDR explicitly pursues block diagonality without sacrificing convexity, using a specially designed convex regularizer.
result Experimental results show ABDR outperforms state-of-the-arts.