Enhances sequence memory capacity in neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This research proves guarantees on sequence models' generalization to longer and novel sequences.
BestChanID identifies the channel with maximal capacity using training sequences.
Learning to remember long sequences remains a challenging task for recurrent neural networks. Register memory and attention mechanisms were both proposed to resolve the issue with either high computational cost to retain memory differentiability, or by discounting the RNN representation learning towards encoding shorte…
We study the computational capacity of a model neuron, the Tempotron, which classifies sequences of spikes by linear-threshold operations. We use statistical mechanics and extreme value theory to derive the capacity of the system in random classification tasks. In contrast to its static analog, the Perceptron, the Temp…
In this paper, the echo state network (ESN) memory capacity, which represents the amount of input data an ESN can store, is analyzed for a new type of deep ESNs. In particular, two deep ESN architectures are studied. First, a parallel deep ESN is proposed in which multiple reservoirs are connected in parallel allowing …
Study semicontinuity of capacity in non-smooth spaces using intrinsic flat convergence.
gLSTM improves graph neural networks by increasing storage capacity to prevent over-squashing.
Constrained sequence codes have been widely used in modern communication and data storage systems. Sequences encoded with constrained sequence codes satisfy constraints imposed by the physical channel, hence enabling efficient and reliable transmission of coded symbols. Traditional encoding and decoding of constrained …
Deep conditional generative models are developed to simultaneously learn the temporal dependencies of multiple sequences. The model is designed by introducing a three-way weight tensor to capture the multiplicative interactions between side information and sequences. The proposed model builds on the Temporal Sigmoid Be…
The study compares different scRNA sequencing methods using a high-dimensional dataset.
Globally normalized neural sequence models are considered superior to their locally normalized equivalents because they may ameliorate the effects of label bias. However, when considering high-capacity neural parametrizations that condition on the whole input sequence, both model classes are theoretically equivalent in…
Neural attention (NA) has become a key component of sequence-to-sequence models that yield state-of-the-art performance in as hard tasks as abstractive document summarization (ADS) and video captioning (VC). NA mechanisms perform inference of context vectors; these constitute weighted sums of deterministic input sequen…
Although Recurrent Neural Network (RNN) has been a powerful tool for modeling sequential data, its performance is inadequate when processing sequences with multiple patterns. In this paper, we address this challenge by introducing a novel mixture layer and constructing an adaptive RNN. The mixture layer augmented RNN (…
Deep dynamic generative models are developed to learn sequential dependencies in time-series data. The multi-layered model is designed by constructing a hierarchy of temporal sigmoid belief networks (TSBNs), defined as a sequential stack of sigmoid belief networks (SBNs). Each SBN has a contextual hidden state, inherit…
Transformers learn to recall with non-orthogonal embeddings in realistic settings.
The main theme of this paper is a relative version of the almost existence theorem for periodic orbits of autonomous Hamiltonian systems. We show that almost all low levels of a function on a geometrically bounded symplectically aspherical manifold carry contractible periodic orbits of the Hamiltonian flow, provided th…
Deep imagination optimizes decision-making in large trees with limited resources.
Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network can be increased by widening and adding layers. However, usually the former introduces additional parameters, while the latter increases the…
Funnel-Transformer reduces computation by compressing sequence data.
Accurately predicting the future health of batteries is necessary to ensure reliable operation, minimise maintenance costs, and calculate the value of energy storage investments. The complex nature of degradation renders data-driven approaches a promising alternative to mechanistic modelling. This study predicts the ch…
Can machines trace human knowledge like humans? Knowledge tracing (KT) is a fundamental task in a wide range of applications in education, such as massive open online courses (MOOCs), intelligent tutoring systems, educational games, and learning management systems. It models dynamics in a student's knowledge states in …
Global Memory Augmentation (GMAT) improves Transformer performance on long documents.
Many state-of-the-art results obtained with deep networks are achieved with the largest models that could be trained, and if more computation power was available, we might be able to exploit much larger datasets in order to improve generalization ability. Whereas in learning algorithms such as decision trees the ratio …
Reservoir Computing (RC) refers to a Recurrent Neural Networks (RNNs) framework, frequently used for sequence learning and time series prediction. The RC system consists of a random fixed-weight RNN (the input-hidden reservoir layer) and a classifier (the hidden-output readout layer). Here we focus on the sequence lear…
TVS-FNNs can approximate any continuous function on expanded input spaces.
Generative autoencoders offer a promising approach for controllable text generation by leveraging their latent sentence representations. However, current models struggle to maintain coherent latent spaces required to perform meaningful text manipulations via latent vector operations. Specifically, we demonstrate by exa…
Recent work has shown that topological enhancements to recurrent neural networks (RNNs) can increase their expressiveness and representational capacity. Two popular enhancements are stacked RNNs, which increases the capacity for learning non-linear functions, and bidirectional processing, which exploits acausal informa…
Variable order sequence modeling is an important problem in artificial and natural intelligence. While overcomplete Hidden Markov Models (HMMs), in theory, have the capacity to represent long-term temporal structure, they often fail to learn and converge to local minima. We show that by constraining HMMs with a simple …
Study on RNNs' ability to approximate past-dependent Hölder functions and their application to regression.
In this paper, we define a new capacity which allows us to control the behaviour of the Dirichlet spectrum of a compact Riemannian manifold with boundary, with "small" subsets (which may intersect the boundary) removed. This result generalizes a classical result of Rauch and Taylor ("the crushed ice theorem"). In the s…
Transformers can predict pseudo-random sequences from LCGs with unseen parameters and moduli.
This study examines how sequential correlations affect in-context learning in sequence models.
Paper extends learning theory to dependent data with uniform risk bounds.
We study various capacities on compact Kähler manifolds which generalize the Bedford-Taylor Monge-Ampère capacity. We then use these capacities to study the existence and the regularity of solutions of complex Monge-Ampère equations.
Solves a discrete logarithmic Minkowski problem for electrostatic p-capacity.
Neural networks struggle with extrapolation, but a new framework allows them to learn counterfactual invariances.
CapOptix uses options theory to price capacity in electricity markets.
Single-head transformers with a single self-attention layer can approximate any sequence-to-sequence function and are efficient under certain conditions.
In this article, we propose the notion of the general -affine capacity and prove some basic properties for the general -affine capacity, such as affine invariance and monotonicity. The newly proposed general -affine capacity is compared with several classical geometric quantities, e.g., the volume, the -var…
While symplectic manifolds have no local invariants, they do admit many global numerical invariants. Prominent among them are the so-called symplectic capacities. Different capacities are defined in different ways, and so relations between capacities often lead to surprising relations between different aspects of sympl…
Study excess capacity in neural networks using Rademacher complexity.
Study rigidity by logarithmic capacity and related functions.
Study binary perceptrons' capacity using random duality theory.
Study capacity constraints in continual learning with a simple model.
New analysis shows RPE-based Transformers can't approximate all functions.
New complete panel dataset for LMICs helps analyze innovation and development.
Upper bounds for Lagrangian capacities of Liouville domains