OPT framework improves neural network generalization by learning an orthogonal transformation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We develop the idea of using an algebraic-geometry approach to classical differential geometry problems. Consider an orthogonal net constructed according to algebraic-geometric data we obtain a set of smooth orthogonal nets that are Ribaucour transformations of the initial orthogonal net.
LOFT separates subspace rotation and transformation for orthogonal fine-tuning.
New method enforces orthogonality in convolutional layers for improved robustness.
Transformers learn to recall with non-orthogonal embeddings in realistic settings.
Pion optimizes LLMs by preserving weight matrix singular values.
This paper explores how Transformers predict next tokens in autoregressive tasks.
Constructs orthogonal coordinates in curved spaces.
Nonnegative matrix factorization (NMF) is a popular method for audio spectral unmixing. While NMF is traditionally applied to off-the-shelf time-frequency representations based on the short-time Fourier or Cosine transforms, the ability to learn transforms from raw data attracts increasing attention. However, this adds…
Study isotropy groups for complex orthogonal and skew-symmetric matrices.
New algorithms learn sparse set functions in non-orthogonal Fourier bases.
OSA overcomes instability in skipless Transformers.
Discrete conjugate systems are quadrilateral nets with all planar faces. Discrete orthogonal systems are defined by the additional property of all faces being concircular. Their geometric properties allow one to consider them as proper discretization of conjugate, resp. orthogonal coordinate systems of classical differ…
Batch normalization makes deep neural networks' representations increasingly orthogonal.
OLinear forecasts time series more efficiently by transforming data orthogonally.
New method for curve comparison using iterated integrals and moving frames.
A Clifford algebra model for M"obius geometry is presented. The notion of Ribaucour pairs of orthogonal systems in arbitrary dimensions is introduced, and the structure equations for adapted frames are derived. These equations are discretized and the geometry of the occuring discrete nets and sphere congruences is disc…
Proposes -PCA to learn identifiable linear transformations without whitening.
A new hashing method improves accuracy by learning an orthogonal transform.
The theory of slice regular functions of a quaternion variable is applied to the study of orthogonal complex structures on domains Ω of R^4. When Ω is a symmetric slice domain, the twistor transform of such a function is a holomorphic curve in the Klein quadric. The case in which Ω is the complement of a parabola is st…
We show how Ramond free neutral Fermi fields lead to a -function theory of BKP type which describes iso-orthogonal deformations of systems of ortogonal curvilinear coordinates. We also provide a vertex operator representation for the classical Ribaucour transformation.
Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients. Short of designing new RNN architectures, previous methods for dealing with this problem usually boil down to orthogonalization of the recurrent dynamics,…
Recurrent neural networks (RNNs) have been successfully used on a wide range of sequential data problems. A well known difficulty in using RNNs is the \textit{vanishing or exploding gradient} problem. Recently, there have been several different RNN architectures that try to mitigate this issue by maintaining an orthogo…
The vectorial fundamental transformation for the Darboux equations is reduced to the symmetric case. This is combined with the orthogonal reduction of Lame type to obtain those vectorial Ribaucour transformations which preserve the Egoroff reduction. We also show that a permutability property holds for all these transf…
OPAA estimates probability densities using functional analysis.
Overparameterized models improve performance in sequential learning tasks.
Study geodesics on flat tori, focusing on orthogonal lengths and their distribution.
We study the problem of approximating orthogonal matrices so that their application is numerically fast and yet accurate. We find an approximation by solving an optimization problem over a set of structured matrices, that we call extended orthogonal Givens transformations, including Givens rotations as a special case. …
The paper tackles learning symmetries in data without expert knowledge.
A diagonal metric sum_{i=1}^n g_{ii} dx_i^2 is termed Guichard_k if sum_{i=1}^{n-k}g_{ii}-sum_{i=n-k+1}^n g_{ii}=0. A hypersurface in R^{n+1} is isothermic_k if it admits line of curvature co-ordinates such that its induced metric is Guichard_k. Isothermic_1 surfaces in R^3 are the classical isothermic surfaces in R^3.…
Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simp…
We survey the existing parts of a classification of finite groups generated by orthogonal transformations in a finite-dimensional Euclidean space whose fixed point subspace has codimension one or two and extend it to a complete classification. These groups naturally arise in the study of the quotient of a Euclidean spa…
Finsler space is differentiable manifold for which Minkowski space is the fiber of the tangent bundle. To understand structure of the reference frame in Finsler space, we need to understand the structure of orthonormal basis in Minkowski space. In this paper, I considered the definition of orthonormal basis in Minkowsk…
Different neural networks trained on the same dataset often learn similar input-output mappings with very different weights. Is there some correspondence between these neural network solutions? For linear networks, it has been shown that different instances of the same network architecture encode the same representatio…
We introduce a novel approach to perform first-order optimization with orthogonal and unitary constraints. This approach is based on a parametrization stemming from Lie group theory through the exponential map. The parametrization transforms the constrained optimization problem into an unconstrained one over a Euclidea…
A machine learning method selects optimal orthonormal bases for functional data analysis.
Initialization of parameters in deep neural networks has been shown to have a big impact on the performance of the networks (Mishkin & Matas, 2015). The initialization scheme devised by He et al, allowed convolution activations to carry a constrained mean which allowed deep networks to be trained effectively (He et al.…
Muon optimizer simplifies matrix optimization with spectral orthogonalization.
A recent strategy to circumvent the exploding and vanishing gradient problem in RNNs, and to allow the stable propagation of signals over long time scales, is to constrain recurrent connectivity matrices to be orthogonal or unitary. This ensures eigenvalues with unit norm and thus stable dynamics and training. However …
DFRot improves LLMs by reducing outlier and massive activation effects.
Transitive consistency is an intrinsic property for collections of linear invertible transformations between Euclidean coordinate frames. In practice, when the transformations are estimated from data, this property is lacking. This work addresses the problem of synchronizing transformations that are not transitively co…
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
Geometrically transforms word embeddings into a common space for better comparison.
Efficiently optimizes orthogonal and Stiefel matrices on parallel units.
New methods create full discretized isothermic tori in Euclidean spaces.
Transformation models are a very important tool for applied statisticians and econometricians. In many applications, the dependent variable is transformed so that homogeneity or normal distribution of the error holds. In this paper, we analyze transformation models in a high-dimensional setting, where the set of potent…
A new optimizer preserves orthogonality constraints on matrices efficiently.
Spectral method for joint community detection and group synchronization.