Quaternion self-attention reduces computational cost and improves performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method uses shared attention for multi-task time series forecasting.
In recent years, dock-less shared bikes have been widely spread across many cities in China and facilitate people's lives. However, at the same time, it also raises many problems about dock-less shared bike management due to the mismatching between demands and real distribution of bikes. Before deploying dock-less shar…
This work proposes a collaborative multi-head attention layer to reduce model size without sacrificing accuracy.
Bayesian neural networks improve knowledge sharing across networks.
DAF uses attention sharing to adapt forecasts from abundant to scarce data.
Reinforcement learning in multi-agent scenarios is important for real-world applications but presents challenges beyond those seen in single-agent settings. We present an actor-critic algorithm that trains decentralized policies in multi-agent settings, using centrally computed critics that share an attention mechanism…
We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. This is made possible by having a single attention mechanism that is …
The prevalence of social media has made information sharing possible across the globe. The downside, unfortunately, is the wide spread of misinformation. Methods applied in most previous rumor classifiers give an equal weight, or attention, to words in the microblog, and do not take the context beyond microblog content…
Transformers learn to generalize unseen tasks by composing self-attention layers.
Mathematical framework for understanding attention in neural networks.
Introduces Causal Energy Minimization to understand Transformer layers.
SCHA-VAE generates novel data from limited examples using hierarchical context aggregation.
Multi-task learning shares information between related tasks, sometimes reducing the number of parameters required. State-of-the-art results across multiple natural language understanding tasks in the GLUE benchmark have previously used transfer from a single large task: unsupervised pre-training with BERT, where a sep…
Proposes vMF distribution for skewed elliptical distributions.
Most companies utilize demographic information to develop their strategy in a market. However, such information is not available to most retail companies. Several studies have been conducted to predict the demographic attributes of users from their transaction histories, but they have some limitations. First, they focu…
Survey of multimodal deep generative models for diverse data types.
New method estimates treatment effects across different populations.
We statistically investigate the distribution of share price and the distributions of three common financial indicators using data from approximately 8,000 companies publicly listed worldwide for the period 2004-2013. We find that the distribution of share price follows Zipf's law; that is, it can be approximated by a …
This paper investigates weight-sharing in NAS, revealing its impact and providing solutions.
Paper proposes deep learning model for dynamic stock repurchase forecasting.
Federated learning interprets temporal dynamics across clients with graph attention.
GSAN learns adaptive node representations using geometric scattering and attention.
Extends inf-convolution to countable risk measures for risk sharing.
GAttNHP predicts future events in temporal knowledge graphs by encoding long-range dependencies and handling mutual excitation.
In a Markovian stochastic volatility model, we consider financial agents whose investment criteria are modelled by forward exponential performance processes. The problem of contingent claim indifference valuation is first addressed and a number of properties are proved and discussed. Special attention is given to the c…
Recommender systems (RS), which have been an essential part in a wide range of applications, can be formulated as a matrix completion (MC) problem. To boost the performance of MC, matrix completion with side information, called inductive matrix completion (IMC), was further proposed. In real applications, the factorize…
Time series prediction with deep learning methods, especially long short-term memory neural networks (LSTMs), have scored significant achievements in recent years. Despite the fact that the LSTMs can help to capture long-term dependencies, its ability to pay different degree of attention on sub-window feature within mu…
This paper proposes a new D2D data sharing approach to improve distributed machine learning training speed.
We propose a GAN design which models multiple distributions effectively and discovers their commonalities and particularities. Each data distribution is modeled with a mixture of generator distributions. As the generators are partially shared between the modeling of different true data distributions, shared ones ca…
Paper studies federated nonparametric testing with privacy constraints, achieving optimal rates and adaptive testing.
Paper introduces a novel framework for set input tasks in meta-learning.
3D Axial-Attention improves lung nodule classification accuracy.
Aligns attention distributions for improved accuracy and robustness.
InGRA models for efficient Granger causality learning in multivariate time series.
This paper improves transportation efficiency by teaching automated vehicles to cooperate.
We study the problem of distributed multi-task learning with shared representation, where each machine aims to learn a separate, but related, task in an unknown shared low-dimensional subspaces, i.e. when the predictor matrix has low rank. We consider a setting where each task is handled by a different machine, with sa…
Online learning with streaming data in a distributed and collaborative manner can be useful in a wide range of applications. This topic has been receiving considerable attention in recent years with emphasis on both single-task and multitask scenarios. In single-task adaptation, agents cooperate to track an objective o…
This paper describes the participation of Amobee in the shared sentiment analysis task at SemEval 2018. We participated in all the English sub-tasks and the Spanish valence tasks. Our system consists of three parts: training task-specific word embeddings, training a model consisting of gated-recurrent-units (GRU) with …
Meta learns low-rank covariance factors for better uncertainty estimation.
Sharing knowledge between tasks is vital for efficient learning in a multi-task setting. However, most research so far has focused on the easier case where knowledge transfer is not harmful, i.e., where knowledge from one task cannot negatively impact the performance on another task. In contrast, we present an approach…
Looped Transformers improve robustness and expressivity in in-context learning for diverse tasks.
Distributed machine learning has been widely studied in order to handle exploding amount of data. In this paper, we study an important yet less visited distributed learning problem where features are inherently distributed or vertically partitioned among multiple parties, and sharing of raw data or model parameters amo…
LieTransformer extends self-attention to Lie groups for improved deep learning tasks.
Variational Autoencoders (VAEs) are powerful in data representation inference, but it cannot learn relations between features with its vanilla form and common variations. The ability to capture relations within data can provide the much needed inductive bias necessary for building more robust Machine Learning algorithm…
NeoMLP improves neural fields by adding self-attention for better downstream tasks.
Study compares deep learning models for traffic forecasting, highlighting graph elements' impact.
Deep generative models are attracting great attention as a new promising approach for molecular design. All models reported so far are based on either variational autoencoder (VAE) or generative adversarial network (GAN). Here we propose a new type model based on an adversarially regularized autoencoder (ARAE). It basi…