Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

265278104 · Jun 202019922001200920172026
48 results for diagonal patterns

The paper explores how sinks and diagonal patterns prevent attention oversmoothing.

problem Preventing attention oversmoothing in neural networks.
method Analyzing geometric conditions and conditions for dense vs. sparse attention, proving equivalence between sinks and hard attention switch, and comparing the costs of sinks vs. diagonal patterns.
result Sinks and diagonal patterns effectively prevent attention oversmoothing, and diagonal patterns provide a more flexible approach.

Khovanov homology of a link and chromatic graph homology are known to be isomorphic in a range of homological gradings that depend on the girth of a graph. We discuss patterns shared by these two homology theories. In particular, we improve the bounds for the homological span of chromatic homology by Helme-Guizon, Przy…

2018-01-04abs ↗pdf ↗

New framework extends ICA for non-independent variables, identifying pairwise mean independence.

problem Non-independent variables complicating ICA recovery.
method Algebraic recovery algorithm based on least-squares optimization over the orthogonal group.
result Pairwise mean independence is identifiable, robust to independence constraints.

Recurrent Neural Networks (RNNs) are powerful models that achieve exceptional performance on several pattern recognition problems. However, the training of RNNs is a computationally difficult task owing to the well-known "vanishing/exploding" gradient problem. Algorithms proposed for training RNNs either exploit no (or…

2015-11-04abs ↗pdf ↗

We demystify attention patterns in multi-head softmax models for linear data.

problem Understanding the training dynamics and emergent patterns in multi-head softmax attention models.
method Extensive empirical experiments and rigorous theoretical analysis.
result Multi-head softmax attention models approximate a debiased gradient descent predictor, outperforming single-head attention and achieving near-Bayesian optimality.

Diagonal linear networks converge to lasso regularization path during training.

problem Understanding the regularization behavior of diagonal linear networks.
method Analyzing the training trajectory of diagonal linear networks and comparing it to the lasso regularization path.
result The training trajectory of diagonal linear networks is closely related to the lasso regularization path.

We show that a basis of a semisimple Lie algebra of compact type, for which any diagonal left-invariant metric has a diagonal Ricci tensor, is characterized by the Lie algebraic condition of being "nice". Namely, the bracket of any two basis elements is a multiple of another basis element. This extends the work of Laur…

2019-12-29abs ↗pdf ↗

Study grid homology of diagonal knots, finding key terms related to prime factors and decompositions.

problem Determine grid homology of diagonal knots and compare them to other knot types.
method Use grid diagrams and combinatorial knot Floer homology to analyze diagonal knots.
result Grid homology detects the number of prime factors and decompositions of the knot into non-integer tangles.

Study on stability of non-diagonal Einstein metrics on specific homogeneous spaces.

problem Stability analysis of non-diagonal Einstein metrics on HimesH/ΔKH imes H/ΔK.
method Formula for scalar curvature, study of stability with Hilbert action.
result Non-diagonal Einstein metrics on MM are unstable with different coindexes.

Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…

2019-05-26abs ↗pdf ↗

Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.

problem Subspace clustering with block diagonal structure for noisy data.
method ABDR explicitly pursues block diagonality without sacrificing convexity, using a specially designed convex regularizer.
result Experimental results show ABDR outperforms state-of-the-arts.

The approximate joint diagonalization of a set of matrices consists in finding a basis in which these matrices are as diagonal as possible. This problem naturally appears in several statistical learning tasks such as blind signal separation. We consider the diagonalization criterion studied in a seminal paper by Pham (…

2018-11-28abs ↗pdf ↗

We prove a number of convexity results for strata of the diagonal pants graph of a surface, in analogy with the extrinsic geometric properties of strata in the Weil-Petersson completion. As a consequence, we exhibit convex flat subgraphs of every possible rank inside the diagonal pants graph.

2011-11-04abs ↗pdf ↗

Classify projective subvarieties in Bogomolov-Guan manifolds using quasi-diagonals.

problem Classify projective subvarieties in non-Kahler holomorphically symplectic manifolds.
method Use quasi-diagonals to classify projective subvarieties.
result Prove that any projective subvariety belongs to a fiber of the Lagrangian fibration.

In this paper, we study deep diagonal circulant neural networks, that is deep neural networks in which weight matrices are the product of diagonal and circulant ones. Besides making a theoretical analysis of their expressivity, we introduced principled techniques for training these models: we devise an initialization s…

2019-01-29abs ↗pdf ↗

The main purpose of this note is to prove that any basis of a nilpotent Lie algebra for which all diagonal left-invariant metrics have diagonal Ricci tensor necessarily produce quite a simple set of structural constants; namely, the bracket of any pair of elements of the basis must be a multiple of some of them and onl…

2011-10-18abs ↗pdf ↗

New method improves deep learning model robustness and accuracy for long sequences.

problem Challenges in learning long-range sequence tasks using state-space models.
method Proposes a perturb-then-diagonalize (PTD) methodology to address ill-posed diagonalization problems in SSMs.
result Demonstrates improved robustness and accuracy of S5-PTD model on Long-Range Arena benchmark.

We give a quadratic lower bound on the dimension of the space of conjugacy classes of subgroups of SL(n,R) that are limits under conjugacy of the diagonal subgroup. We give the first explicit examples of abelian n-1 dimensional subgroups of SL(n,R) which are not such a limit, however all such abelian groups are limits …

2014-12-17abs ↗pdf ↗

Diagonal metrics solve Hermitian-Einstein equations for decomposed Higgs bundles.

problem Existence of diagonal pluriharmonic metrics in GG-Higgs bundles.
method Analyzes Higgs bundles over compact Kähler manifolds, decomposes vector bundles, and uses torus action to relate stability and conditions.
result Necessary and sufficient conditions for the existence of diagonal metrics solving Hermitian-Einstein equations.

In this paper, we propose a new Recurrent Neural Network (RNN) architecture. The novelty is simple: We use diagonal recurrent matrices instead of full. This results in better test likelihood and faster convergence compared to regular full RNNs in most of our experiments. We show the benefits of using diagonal recurrent…

2017-04-18abs ↗pdf ↗

We introduce the notion of Haantjes algebra: It consists of an assignment of a family of operator fields on a differentiable manifold, each of them with vanishing Haantjes torsion. They are also required to satisfy suitable compatibility conditions. Haantjes algebras naturally generalize several known interesting geome…

2017-10-12abs ↗pdf ↗

Diagonal transformations preserve independence structures in non-Gaussian distributions.

problem Preserving independence structures in non-Gaussian distributions.
method Diagonal nonlinear transformations of multivariate normal variables.
result Independence structures are preserved in non-Gaussian distributions under diagonal transformations.

New insights into Hessian structure of neural networks reveal two forces.

problem Understanding the Hessian structure of neural networks.
method Analyzing the static and dynamic forces, comparing limit distributions using random matrix theory.
result The Hessian structure arises from a combination of static and dynamic forces, with CC being a primary driver.