New model mimics neural next item recommendation using Hankel matrices.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.
New algorithms improve multi-task learning across different environments.
Although group convolutional networks are able to learn powerful representations based on symmetry patterns, they lack explicit means to learn meaningful relationships among them (e.g., relative positions and poses). In this paper, we present attentive group equivariant convolutions, a generalization of the group convo…
OLS is a special case of Transformer, revealing its linear nature.
Slot Attention extracts object-centric representations from images.
Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…
We derive the stress-energy tensor for polyharmonic maps between Riemannian manifolds. Moreover, we employ the stress-energy tensor to characterize polyharmonic maps where we pay special attention to triharmonic maps.
We study families of submanifolds in symmetric spaces of compact type arising as exponential images of s-orbits of variable radii. Special attention is given to the cases where the s-orbits are symmetric.
New analysis shows how cross-entropy training shapes attention in transformers.
We review geometrical properties of a static spacetime , including geodesic completeness, causality, standard splittings, compact , closed geodesics and geodesic connectedness. We pay special attention to the critical quadratic behavior at infinity of the coefficients , (, being a …
Generalized are the investigated in other works of the author transports along paths in fibre bundles to transports along arbitrary maps in them. Their structure and some properties are studied. Special attention is paid to the linear case and the case when the map's domain is a Cartesian product of two sets. Also cons…
Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed attention. The multiple heads learn diverse types of word relationships. However, with standard softmax attention, all attention heads are de…
Recent links between Finsler Geometry and the geometry of spacetimes are briefly revisited, and prospective ideas and results are explained. Special attention is paid to geometric problems with a direct motivation in Relativity and other parts of Physics.
Most previous studies on multi-agent reinforcement learning focus on deriving decentralized and cooperative policies to maximize a common reward and rarely consider the transferability of trained policies to new tasks. This prevents such policies from being applied to more complex multi-agent tasks. To resolve these li…
Transformers learn to integrate information from past positions incrementally, specializing heads in distinct patterns.
In this survey article we gather classical as well as recent results on minimal geodesics of Riemannian or Finsler metrics, giving special attention to the two-dimensional case. Moreover, we present open problems together with some first ideas as to the solutions.
SurvBESA predicts survival times using ensemble methods with self-attention.
Study smoothings of singular intersections of ellipsoids.
Unified perspective on Hopfield networks with attention module.
Study shows how specialized attention circuits emerge during transformer training.
Paper analyzes infinite-width attention layers using Tensor Programs.
Modelling and exploiting teammates' policies in cooperative multi-agent systems have long been an interest and also a big challenge for the reinforcement learning (RL) community. The interest lies in the fact that if the agent knows the teammates' policies, it can adjust its own policy accordingly to arrive at proper c…
Recent work Bobienski-Nurowski on 5-dimensional Riemannian manifolds with an SO(3) structure prompts us to investigate which Lie groups admit such a geometry. The case in which the SO(3) structure admits a compatible connection with torsion is considered. This leads to a classification under special behaviour of the on…
We focus our attention on the notion of intrinsic Lipschitz graphs, inside a special class of metric spaces i.e. the Carnot groups. More precisely, we provide a characterization of locally intrinsic Lipschitz functions in Carnot groups of step 2 in terms of their intrinsic distributional gradients.
Transformers cluster meaningless words around leaders for sentiment analysis.
Unified framework for sequence models using test-time regression.
The paper explores invariant vs non-invariant complex structures on Lie groups.
Researchers found differential invariants for Kundt spacetimes.
In this article we study the role of the Green function for the Laplacian in a compact Riemannian manifold as a tool for obtaining well-distributed points. In particular, we prove that a sequence of minimizers for the Green energy is asymptotically uniformly distributed. We pay special attention to the case of locally …
New Ising models improve consensus clustering on specialized hardware.
Transformers learn causal structure through gradient descent on self-attention mechanisms.
In this paper, we generalize the parametric delta-VaR method from portfolios with normally distributed risk factors to portfolios with elliptically distributed ones. We treat both the expected shortfall and the Value-at-Risk of such portfolios. Special attention is given to the particular case of a multivariate t-distr…
The paper glosses different forms of an introducing of higher order tangent-like functors, especially functors derived from higher order nonholonomic tangent functors. A special attention is devoted to higher order osculating bundles: their identification with higher order tangent bundles is demonstrated as the main re…
This paper is devoted to the study of the global properties of harmonically immersed Riemann surfaces in We focus on the geometry of complete harmonic immersions with quasiconformal Gauss map, and in particular, of those with finite total curvature. We pay special attention to the construction of new ex…
In this paper, we generalize the parametric Delta-VaR methods from portfolios with elliptic distributed risk factors to portfolios with mixture of elliptically distributed ones. We treat both the Expected Shortfall and the Value-at-Risk of such portfolios. Special attention is given to the particular case of the mixtur…
AGNN improves network localization accuracy by 37-53% in NLOS conditions.
While deep learning has achieved great success in computer vision and many other fields, currently it does not work very well on patient genomic data with the "big p, small N" problem (i.e., a relatively small number of samples with high-dimensional features). In order to make deep learning work with a small amount of …
Study of symplectic manifolds degenerating into singular spaces.
This is a survey of old and new results on the problem when a compatible almost complex structure on a Riemannian manifold is a harmonic section or a harmonic map from the manifold into its twistor space. In this context, a special attention is paid to the Atiyah-Hitchin-Singer and Eells-Salamon almost complex structur…
New calculus framework for vector bundles with metrics.
Develops models for temporally abstract reasoning and attention.
Solves relative isoperimetric problem on polygonal domains, focusing on corners.
A generalized notion of a Lie algebroid is presented. Using this, the Lie algebroid generalized tangent bundle is obtained. A new point of view over (linear) connections theory on a fiber bundle is presented. These connections are characterized by o horizontal distribution of the Lie algebroid generalized tangent bundl…
Study conjugate locus in convex 3-manifolds using Jacobi fields.
Streamable model improves speech recognition performance.
New theory shows how multi-head attention reduces variance and decorrelates outputs.
Hydra boosts efficiency for long-context reasoning in resource-constrained settings.