Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

70139209278 · Jun 202019922001200920172026
48 results for specialized attention

SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.

problem Overreacting to peak events in demand forecasting leads to biased forecasts.
method SPADE splits forecasting into two tasks: one for peak events and another for post-peak events, using masked convolution filters and a specialized Peak Attention module.
result Overall PPE improvement of 4.5%, 30% improvement for most affected forecasts after promotions and holidays, and 3.9% improvement in PE accuracy.

Although group convolutional networks are able to learn powerful representations based on symmetry patterns, they lack explicit means to learn meaningful relationships among them (e.g., relative positions and poses). In this paper, we present attentive group equivariant convolutions, a generalization of the group convo…

2020-02-07abs ↗pdf ↗

Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…

2017-05-22abs ↗pdf ↗

We study families of submanifolds in symmetric spaces of compact type arising as exponential images of s-orbits of variable radii. Special attention is given to the cases where the s-orbits are symmetric.

2005-05-25abs ↗pdf ↗

New analysis shows how cross-entropy training shapes attention in transformers.

problem Understanding how gradient-based learning creates the required internal geometry in transformers.
method Developed a first-order analysis of cross-entropy training effects on attention scores and values in a transformer attention head.
result Introduced an advantage-based routing law and responsibility-weighted update for attention scores and values, respectively.

We review geometrical properties of a static spacetime (M,g)(M,g), including geodesic completeness, causality, standard splittings, compact MM, closed geodesics and geodesic connectedness. We pay special attention to the critical quadratic behavior at infinity of the coefficients ββ, β1β^{-1} (β=g(K,K)β= -g(K,K), being KK a …

2004-06-16abs ↗pdf ↗

Generalized are the investigated in other works of the author transports along paths in fibre bundles to transports along arbitrary maps in them. Their structure and some properties are studied. Special attention is paid to the linear case and the case when the map's domain is a Cartesian product of two sets. Also cons…

1997-09-20abs ↗pdf ↗

Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed attention. The multiple heads learn diverse types of word relationships. However, with standard softmax attention, all attention heads are de…

2019-08-30abs ↗pdf ↗

Recent links between Finsler Geometry and the geometry of spacetimes are briefly revisited, and prospective ideas and results are explained. Special attention is paid to geometric problems with a direct motivation in Relativity and other parts of Physics.

2013-11-19abs ↗pdf ↗

Most previous studies on multi-agent reinforcement learning focus on deriving decentralized and cooperative policies to maximize a common reward and rarely consider the transferability of trained policies to new tasks. This prevents such policies from being applied to more complex multi-agent tasks. To resolve these li…

2019-09-27abs ↗pdf ↗

Transformers learn to integrate information from past positions incrementally, specializing heads in distinct patterns.

problem How transformers learn to integrate information from multiple past positions with varying statistical significance.
method High-order Markov chain task, incremental learning, sparse attention patterns, simplified differential equations, stage-wise convergence, early stopping as regularizer.
result Transformers learn to specialize heads in distinct patterns, shifting from competitive to cooperative learning dynamics.

SurvBESA predicts survival times using ensemble methods with self-attention.

problem Challenges in survival analysis due to censored data and unstable predictions.
method SurvBESA combines Beran estimators with a self-attention mechanism to predict survival times.
result SurvBESA outperforms state-of-the-art models in predicting survival times.

Study shows how specialized attention circuits emerge during transformer training.

problem Understanding the mechanisms of transformer training dynamics at large scales.
method Controlled sparse modular addition task; monitoring token evolution via visual sandbox.
result Specialized attention circuits (clustering heads) naturally emerge during training.

Paper analyzes infinite-width attention layers using Tensor Programs.

problem Capturing the infinite-width limit of attention layers.
method Tensor Programs framework to rigorously identify the limit distribution.
result Derives exact form of infinite-width limit distribution without Gaussian approximations.

Recent work Bobienski-Nurowski on 5-dimensional Riemannian manifolds with an SO(3) structure prompts us to investigate which Lie groups admit such a geometry. The case in which the SO(3) structure admits a compatible connection with torsion is considered. This leads to a classification under special behaviour of the on…

2006-07-17abs ↗pdf ↗

We focus our attention on the notion of intrinsic Lipschitz graphs, inside a special class of metric spaces i.e. the Carnot groups. More precisely, we provide a characterization of locally intrinsic Lipschitz functions in Carnot groups of step 2 in terms of their intrinsic distributional gradients.

2019-03-06abs ↗pdf ↗

Transformers cluster meaningless words around leaders for sentiment analysis.

problem Capturing context in sentiment analysis using transformers.
method Characterized transformers with hardmax self-attention and normalization, showing asymptotic convergence to clustered equilibrium.
result Transformers can effectively capture context by clustering meaningless words around leader words.

Unified framework for sequence models using test-time regression.

problem Designing efficient sequence models with associative memory.
method Formalizing associative recall as regression over input tokens, deriving various sequence models.
result Clarifies the effectiveness of query-key normalization in softmax attention and offers new generalizations.

The paper explores invariant vs non-invariant complex structures on Lie groups.

problem Understanding complex structures on Lie groups and their properties.
method Analysis of invariant and non-invariant almost complex structures on compact quotients of Lie groups.
result New computations of Kodaira dimension for invariant and non-invariant structures.

In this article we study the role of the Green function for the Laplacian in a compact Riemannian manifold as a tool for obtaining well-distributed points. In particular, we prove that a sequence of minimizers for the Green energy is asymptotically uniformly distributed. We pay special attention to the case of locally …

2017-02-02abs ↗pdf ↗

The paper glosses different forms of an introducing of higher order tangent-like functors, especially functors derived from higher order nonholonomic tangent functors. A special attention is devoted to higher order osculating bundles: their identification with higher order tangent bundles is demonstrated as the main re…

2012-02-13abs ↗pdf ↗

This paper is devoted to the study of the global properties of harmonically immersed Riemann surfaces in R3.\mathbb{R}^3. We focus on the geometry of complete harmonic immersions with quasiconformal Gauss map, and in particular, of those with finite total curvature. We pay special attention to the construction of new ex…

2011-02-21abs ↗pdf ↗

AGNN improves network localization accuracy by 37-53% in NLOS conditions.

problem Massive network localization under Non-Line-of-Sight conditions.
method Attentional Graph Neural Network (AGNN) with Adjacency Learning Module (ALM) and Multiple Graph Attention Layers (MGAL).
result Significant improvement in localization accuracy, approaching fundamental lower bounds.

This is a survey of old and new results on the problem when a compatible almost complex structure on a Riemannian manifold is a harmonic section or a harmonic map from the manifold into its twistor space. In this context, a special attention is paid to the Atiyah-Hitchin-Singer and Eells-Salamon almost complex structur…

2016-05-22abs ↗pdf ↗

New calculus framework for vector bundles with metrics.

problem Developing calculus for vector bundles with fiber metrics.
method Adapting differential calculus to graded commutative algebras and focusing on diole and triole algebras.
result Triole algebra provides a suitable environment for vector bundle calculus with fiber metrics.

Solves relative isoperimetric problem on polygonal domains, focusing on corners.

problem Relative isoperimetric problem on polygonal domains in R2\mathbb{R}^2.
method Developed techniques for polygonal domains, with special attention to corners.
result Solved the relative isoperimetric problem for a square with a square corner removed.

A generalized notion of a Lie algebroid is presented. Using this, the Lie algebroid generalized tangent bundle is obtained. A new point of view over (linear) connections theory on a fiber bundle is presented. These connections are characterized by o horizontal distribution of the Lie algebroid generalized tangent bundl…

2011-01-05abs ↗pdf ↗

New theory shows how multi-head attention reduces variance and decorrelates outputs.

problem Understanding and optimizing multi-head attention in neural networks.
method Developed a statistical theory linking multi-head attention to ensemble Nadaraya-Watson estimators.
result MHA variance reduction depends on head decorrelation, not just head count.

Hydra boosts efficiency for long-context reasoning in resource-constrained settings.

problem Quadratic complexity of transformers limits long-context reasoning in resource-constrained systems.
method Hydra uses a modular architecture with adaptive routing between sparse global attention, mixture-of-experts, and dual memories.
result Hydra achieves significant throughput and accuracy improvements for long-context reasoning.