Minimum attention improves reinforcement learning performance in high-dimensional dynamics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Sequence-to-sequence models, such as attention-based models in automatic speech recognition (ASR), are typically trained to optimize the cross-entropy criterion which corresponds to improving the log-likelihood of the data. However, system performance is usually measured in terms of word error rate (WER), not log-likel…
Attention models can overfit without harming test performance.
In general, a self-attention mechanism has been applied for speaker embedding encoding. Previous studies focused on training the self-attention in a high-level layer, such as the last pooling layer. However, the effect of low-level features was reduced in the speaker embedding encoding. Therefore, we propose masked cro…
Inspired by the adaptation phenomenon of neuronal firing, we propose the regularity normalization (RN) as an unsupervised attention mechanism (UAM) which computes the statistical regularity in the implicit space of neural networks under the Minimum Description Length (MDL) principle. Treating the neural network optimiz…
Over the past decade, multivariate time series classification has received great attention. We propose transforming the existing univariate time series classification models, the Long Short Term Memory Fully Convolutional Network (LSTM-FCN) and Attention LSTM-FCN (ALSTM-FCN), into a multivariate time series classificat…
A-MMSE uses attention to learn efficient OFDM channel estimation.
Approximations of loopy belief propagation, including expectation propagation and approximate message passing, have attracted considerable attention for probabilistic inference problems. This paper proposes and analyzes a generalization of Opper and Winther's expectation consistent (EC) approximate inference method. Th…
In links with two components there are three different types of crossings: self-crossings in the first component, self crossings in the second component, and crossings between components. In this paper we examine the minimum number of crossing changes needed to unlink without changing the crossings between components. …
Transformers learn linear models in-context without updates.
CutMix training technique improves spatial locality in Vision Transformers.
Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch methods tend to converge to sharp minimizers has received increasing attention. We theoretically justify this hypothesis by providing new pro…
A new optimizer DDC improves deep learning models by respecting symmetries.
This article is written for the Proceedings of the Conference on Current Developments in Mathematics in Harvard University, November 16-17, 2007. It is an exposition of the analytic proof of the finite generation of the canonical ring for a compact complex algebraic manifold of general type. It lists and discusses the …
ESE-FN improves elderly activity recognition accuracy.
Paper proposes MWDE for estimating finite location-scale mixtures.
This paper optimizes portfolios using HRP and CLA algorithms on NIFTY 50 stocks.
Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures are comparable to stat…
Paper proposes a generalized precision matrix for t-Student distributions to improve portfolio optimization.
Reformulates binary classification on manifolds using Yang-Mills-Higgs theory.
Transfer learning improves MNI's performance in high-dimensional linear regression.
In this paper we continue to study (`strong') Nielsen coincidence numbers (which were introduced recently for pairs of maps between manifolds of arbitrary dimensions) and the corresponding minimum numbers of coincidence points and pathcomponents. We explore compatibilities with fibrations and, more specifically, with c…
Paper uses algebraic signatures to identify probabilistic structures in empirical data.
Deep convolutional semantic segmentation (DCSS) learning doesn't converge to an optimal local minimum with random parameters initializations; a pre-trained model on the same domain becomes necessary to achieve convergence.In this work, we propose a joint cooperative end-to-end learning method for DCSS. It addresses man…
A novel hypergraph partitioning method using tensor eigenvalue decomposition captures super-dyadic interactions.
Paper analyzes robustness of MDPDE under INH setups.
This paper proves SGD converges to global minimum for over-parameterized ReLU networks.
This paper studies long term investing by an investor that maximizes either expected utility from terminal wealth or from consumption. We introduce the concepts of a generalized stochastic discount factor (SDF) and of the minimum price to attain target payouts. The paper finds that the dynamics of the SDF needs to be c…
This paper proves IRM minimizes o.o.d. risk under certain conditions.
Minimum Description Length prevents overfitting in noisy data.
Verifying robustness of neural network classifiers has attracted great interests and attention due to the success of deep neural networks and their unexpected vulnerability to adversarial perturbations. Although finding minimum adversarial distortion of neural networks (with ReLU activations) has been shown to be an NP…
The minimum number of colors is a challenging knot invariant since, by definition, its calculation requires taking the minimum over infinitely many minima. In this article we estimate and in some cases calculate the minimum number of colors for the Turk's head knots on three strands.
Minimum braids are a complete invariant of knots and links. This paper defines minimum braids, describes how they can be generated, presents tables for knots up to ten crossings and oriented links up to nine crossings, and uses minimum braids to study graph trees, amphicheirality, unknotting numbers, and periodic table…
Study tightens bounds for interpolating noisy data using minimum l1-norm.
A new classification method based on Minimum Spanning Trees
The paper finds minimum Dehn colors for knots and defines useful graphs for coloring.
The paper calculates genus bounds for multibranched surfaces.
Knots are commonly found in molecular chains such as DNA and proteins, and they have been considered to be useful models for structural analysis of these molecules. One interested quantity is the minimum number of monomers necessary to realize a molecular knot. The minimum lattice length $\mbox{Len}(K)$ of a knot i…
Study shows how networks converge to minimum norm solutions with regularization.
We find the minimum dilatation of pseudo-Anosov braids with many strands.
When an AI system interacts with multiple users, it frequently needs to make allocation decisions. For instance, a virtual agent decides whom to pay attention to in a group setting, or a factory robot selects a worker to deliver a part. Demonstrating fairness in decision making is essential for such systems to be broad…
Paper proposes neural network for efficient MIMO channel estimation and pilot reduction.
Study introduces AMVP and AMRR for dynamic portfolio optimization in volatile markets.
Minimum algebraic intersection found in hyperbolic surfaces, growing with genus.
Computed minimum crossing numbers for Turaev genus 2 links.
We refine Expected Shortfall by controlling different tail portions, offering tailored risk assessments.
This paper studies the geometry of minimum-volume confidence sets for multinomial parameters.
Building on previous results on the quadratic helicity in magnetohydrodynamics (MHD) we investigate particular minimum helicity states. Those are eigenfunctions of the curl operator and are shown to constitute solutions of the quasi-stationary incompressible ideal MHD equations. We then show that these states have inde…