The study reveals how attention paths in Transformers influence learning outcomes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SigMA uses signatures and attention to estimate parameters in fBm-driven SDEs.
A new network learns to prioritize messages for efficient multi-robot path planning.
New method for learning on heterogeneous graphs without meta-paths.
Transformers can solve complex filtering problems for non-Gaussian signals.
In this paper, we focus on graph representation learning of heterogeneous information network (HIN), in which various types of vertices are connected by various types of relations. Most of the existing methods conducted on HIN revise homogeneous graph embedding models via meta-paths to learn low-dimensional vector spac…
Transformer model improves asset allocation by unifying forecasting and optimization.
A concise discussion of the axiomatic approach to the concept of parallel transport is presented. Attention is drawn to a bijective map between the sets of connections and (axiomatically defined) parallel transports. The transports along paths are pointed as a generalization of the (axiomatically defined) parallel tran…
Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attent…
This article reports on an investigation of the use of convolutional neural networks to predict the visual attention of chess players. The visual attention model described in this article has been created to generate saliency maps that capture hierarchical and spatial features of chessboard, in order to predict the pro…
In this paper we use a time-evolving graph which consists of a sequence of graph snapshots over time to model many real-world networks. We study the path classification problem in a time-evolving graph, which has many applications in real-world scenarios, for example, predicting path failure in a telecommunication netw…
Dupire's functional Itô calculus provides an alternative approach to the classical Malliavin calculus for the computation of sensitivities, also called Greeks, of path-dependent derivatives prices. In this paper, we introduce a measure of path-dependence of functionals within the functional Itô calculus framework. Name…
Paper proposes HIDAM model to improve MSE default risk assessment using heterogeneous information networks.
Hierarchical pretraining with slow-fast ODEs
Much of the recent work on learning molecular representations has been based on Graph Convolution Networks (GCN). These models rely on local aggregation operations and can therefore miss higher-order graph properties. To remedy this, we propose Path-Augmented Graph Transformer Networks (PAGTN) that are explicitly built…
We propose the Gaussian attention model for content-based neural memory access. With the proposed attention model, a neural network has the additional degree of freedom to control the focus of its attention from a laser sharp attention to a broad attention. It is applicable whenever we can assume that the distance in t…
Transformer adapts to graphs with adaptive attention and auto-regressive decoding.
ABIForest improves anomaly detection using attention weights.
Neural A* uses machine learning to improve path planning efficiency.
Improved speech enhancement with MNTFA using time-frequency attention.
Generative models use kernel smoothing for conditioning on small example sets.
GAT-RWOS uses graph attention to improve imbalanced data classification.
A major drawback of backpropagation through time (BPTT) is the difficulty of learning long-term dependencies, coming from having to propagate credit information backwards through every single step of the forward computation. This makes BPTT both computationally impractical and biologically implausible. For this reason,…
As adversarial attacks pose a serious threat to the security of AI system in practice, such attacks have been extensively studied in the context of computer vision applications. However, few attentions have been paid to the adversarial research on automatic path finding. In this paper, we show dominant adversarial exam…
Proposes PGPS for efficient Bayesian inference.
Paper tackles offline SSP with value iteration for policy evaluation and learning.
Heterogeneous information network (HIN) embedding has gained increasing interests recently. However, the current way of random-walk based HIN embedding methods have paid few attention to the higher-order Markov chain nature of meta-path guided random walks, especially to the stationarity issue. In this paper, we system…
A new system recommends specific knowledge concepts in MOOCs based on student interests.
Without relevant human priors, neural networks may learn uninterpretable features. We propose Dynamics of Attention for Focus Transition (DAFT) as a human prior for machine reasoning. DAFT is a novel method that regularizes attention-based reasoning by modelling it as a continuous dynamical system using neural ordinary…
MAGNA improves graph neural networks by incorporating multi-hop context information.
Generalized are the investigated in other works of the author transports along paths in fibre bundles to transports along arbitrary maps in them. Their structure and some properties are studied. Special attention is paid to the linear case and the case when the map's domain is a Cartesian product of two sets. Also cons…
ie-HGCN addresses HIN challenges by efficiently learning node representations.
Paper proposes a method to evaluate SME credit risk using meta paths.
A new router uses attention-based reinforcement learning to solve detailed routing problems efficiently.
A brain-inspired spiking Transformer reduces energy consumption and enhances interpretability.
The existing literature provides evidence that limit order book data can be used to predict short-term price movements in stock markets. This paper proposes a new neural network architecture for predicting return jump arrivals in equity markets with high-frequency limit order book data. This new architecture, based on …
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
New model forecasts stock market volatility better than existing methods.
Unified theory explains two failure modes of deep transformers and provides initialisation guidelines.
Transformers learn multi-step reasoning through gradient descent.
New imputation strategies improve signature models for irregular time series.
Flow prediction (e.g., crowd flow, traffic flow) with features of spatial-temporal is increasingly investigated in AI research field. It is very challenging due to the complicated spatial dependencies between different locations and dynamic temporal dependencies among different time intervals. Although measurements of …
Rough Transformers improve time series modeling with lower costs and better performance.
A new method extracts events and their arguments efficiently from text.
This research develops a new model for cyber risk and insurance pricing.
The celebrated Sequence to Sequence learning (Seq2Seq) technique and its numerous variants achieve excellent performance on many tasks. However, many machine learning tasks have inputs naturally represented as graphs; existing Seq2Seq models face a significant challenge in achieving accurate conversion from graph form …
Unified theory of optimal transport for random measures.
Machine learning models have become more and more complex in order to better approximate complex functions. Although fruitful in many domains, the added complexity has come at the cost of model interpretability. The once popular k-nearest neighbors (kNN) approach, which finds and uses the most similar data for reasonin…