Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

54109163217 · Jun 202019922001200920172026
48 results for minimum attention

Minimum attention improves reinforcement learning performance in high-dimensional dynamics.

problem Improving reinforcement learning performance in high-dimensional nonlinear dynamics.
method Applying minimum attention as a regularization technique in reinforcement learning, including model-based and model-free approaches.
result Minimum attention outperforms state-of-the-art algorithms in few-shot adaptation and variance reduction.

In general, a self-attention mechanism has been applied for speaker embedding encoding. Previous studies focused on training the self-attention in a high-level layer, such as the last pooling layer. However, the effect of low-level features was reduced in the speaker embedding encoding. Therefore, we propose masked cro…

2020-01-28abs ↗pdf ↗

Over the past decade, multivariate time series classification has received great attention. We propose transforming the existing univariate time series classification models, the Long Short Term Memory Fully Convolutional Network (LSTM-FCN) and Attention LSTM-FCN (ALSTM-FCN), into a multivariate time series classificat…

2018-01-14abs ↗pdf ↗

A-MMSE uses attention to learn efficient OFDM channel estimation.

problem Accurate OFDM channel estimation requires second-order statistics, which are hard to obtain in practice.
method A-MMSE is a model-based DNN framework that learns linear MMSE filters via Attention Transformer, reducing inference complexity.
result A-MMSE outperforms other methods in normalized MSE across various SNR conditions.

In links with two components there are three different types of crossings: self-crossings in the first component, self crossings in the second component, and crossings between components. In this paper we examine the minimum number of crossing changes needed to unlink without changing the crossings between components. …

2019-06-29abs ↗pdf ↗

CutMix training technique improves spatial locality in Vision Transformers.

problem Improving spatial locality in Vision Transformers trained from scratch.
method Comparison of Baseline and Modern training protocols on CIFAR-10, CIFAR-100, and Tiny-ImageNet.
result CutMix training component significantly reduces Mean Attention Distance (MAD) in early layers of Vision Transformers.

A new optimizer DDC improves deep learning models by respecting symmetries.

problem Deep networks' loss is invariant to continuous symmetries, leading to optimization issues.
method DDC builds a Dead-Direction Conditioner that lifts a base optimizer into a G-equivariant one, preserving the quotient geometry.
result DDCAdam and DDCMuon outperform standard optimizers in various tasks, improving validation-train loss gaps and learning dynamics.

ESE-FN improves elderly activity recognition accuracy.

problem Recognizing individual actions and human-object interactions in elderly activities.
method Exploits multi-modal features from RGB videos and skeleton sequences using ESE attentions and a new Multi-modal Loss.
result ESE-FN achieves best accuracy on ETRI-Activity3D dataset.

This paper optimizes portfolios using HRP and CLA algorithms on NIFTY 50 stocks.

problem Designing an optimal stock portfolio with accurate forecasting of future returns and risks.
method Uses hierarchical risk parity and critical line algorithms on NIFTY 50 stocks.
result Hierarchical risk parity algorithm outperformed the critical line algorithm on test data.

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures are comparable to stat…

2017-12-05abs ↗pdf ↗

Paper proposes a generalized precision matrix for t-Student distributions to improve portfolio optimization.

problem Limitations of inverse covariance matrix in non-Gaussian settings.
method Exploits local dependence function to define generalized precision matrix (GPM) for multivariate t-Student distribution.
result GPM leads to statistically significant lower out-of-sample variances in minimum-variance portfolios.

Reformulates binary classification on manifolds using Yang-Mills-Higgs theory.

problem Binary classification on non-contractible spaces.
method Formulates binary classification as a Yang-Mills-Higgs variational problem, encoding data as a functor.
result Reveals a geometric interpretation of binary classification and solves XOR on the torus.

Transfer learning improves MNI's performance in high-dimensional linear regression.

problem Improving model performance in high-dimensional linear regression with diverse data.
method Proposes a Transfer MNI approach, analyzing its excess risk and conditions for outperformance.
result Identifies free-lunch covariate shift regimes where knowledge transfer benefits.

In this paper we continue to study (`strong') Nielsen coincidence numbers (which were introduced recently for pairs of maps between manifolds of arbitrary dimensions) and the corresponding minimum numbers of coincidence points and pathcomponents. We explore compatibilities with fibrations and, more specifically, with c…

2006-06-01abs ↗pdf ↗

Paper uses algebraic signatures to identify probabilistic structures in empirical data.

problem Identifying probabilistic structure from observed binomials in empirical probability tensors.
method Treating vanishing binomials as algebraic signatures, matching signatures to identify models without parameter estimation.
result The method successfully identified rank-one structures in real language data, revealing interpretable sets of words.

Deep convolutional semantic segmentation (DCSS) learning doesn't converge to an optimal local minimum with random parameters initializations; a pre-trained model on the same domain becomes necessary to achieve convergence.In this work, we propose a joint cooperative end-to-end learning method for DCSS. It addresses man…

2017-10-22abs ↗pdf ↗

A novel hypergraph partitioning method using tensor eigenvalue decomposition captures super-dyadic interactions.

problem Capturing super-dyadic interactions in k-uniform hypergraphs.
method Tensor-based representation and tensor eigenvalue decomposition for capturing interactions.
result Improved min-cut solution on 2-uniform hypergraphs (graphs) compared to standard spectral partitioning.

This paper proves SGD converges to global minimum for over-parameterized ReLU networks.

problem Theoretical understanding of implicit neural networks is limited.
method Gradient flow analysis of ReLU activated implicit neural networks.
result Randomly initialized gradient descent converges to global minimum at a linear rate for square loss function in over-parameterized ReLU networks.

This paper studies long term investing by an investor that maximizes either expected utility from terminal wealth or from consumption. We introduce the concepts of a generalized stochastic discount factor (SDF) and of the minimum price to attain target payouts. The paper finds that the dynamics of the SDF needs to be c…

2017-05-10abs ↗pdf ↗

The minimum number of colors is a challenging knot invariant since, by definition, its calculation requires taking the minimum over infinitely many minima. In this article we estimate and in some cases calculate the minimum number of colors for the Turk's head knots on three strands.

2010-02-25abs ↗pdf ↗

Minimum braids are a complete invariant of knots and links. This paper defines minimum braids, describes how they can be generated, presents tables for knots up to ten crossings and oriented links up to nine crossings, and uses minimum braids to study graph trees, amphicheirality, unknotting numbers, and periodic table…

2004-01-06abs ↗pdf ↗

The paper finds minimum Dehn colors for knots and defines useful graphs for coloring.

problem Finding the minimum number of colors for Dehn colorings of knots.
method Analyzes Dehn colorings for knots and defines R\R-palette graphs.
result For Dehn pp-colorable knots, the minimum number of colors is at least log2pfloor+2\lfloor \log_2 p floor +2.

Knots are commonly found in molecular chains such as DNA and proteins, and they have been considered to be useful models for structural analysis of these molecules. One interested quantity is the minimum number of monomers necessary to realize a molecular knot. The minimum lattice length $\mbox{Len}(K)$ of a knot KK i…

2014-11-07abs ↗pdf ↗

When an AI system interacts with multiple users, it frequently needs to make allocation decisions. For instance, a virtual agent decides whom to pay attention to in a group setting, or a factory robot selects a worker to deliver a part. Demonstrating fairness in decision making is essential for such systems to be broad…

2019-12-13abs ↗pdf ↗

Paper proposes neural network for efficient MIMO channel estimation and pilot reduction.

problem High overhead from pilot transmission in wideband MIMO systems.
method Neural network architecture for frequency-aware pilot design and channel estimation, with pruning technique.
result Neural network outperforms linear minimum mean square error (LMMSE) estimation.

Study introduces AMVP and AMRR for dynamic portfolio optimization in volatile markets.

problem Optimizing portfolios in volatile and nonstationary financial markets.
method Adaptive Minimum-Variance Portfolio (AMVP) framework with ARFIMA-FIGARCH processes and non-Gaussian innovations.
result Demonstrated superior performance in risk reduction and portfolio stability during market breaks.

We refine Expected Shortfall by controlling different tail portions, offering tailored risk assessments.

problem Risk assessment in financial positions, especially in tail regions.
method Introducing adjusted Expected Shortfall measures that control different tail portions.
result Adjusted Expected Shortfall measures ensure risk does not exceed specified thresholds for various probability levels.

This paper studies the geometry of minimum-volume confidence sets for multinomial parameters.

problem Determining if minimum-volume confidence sets for multinomial outcomes are disjoint.
method Enumerating and covering the continuous regions of the exact p-value function to study the geometry of minimum-volume confidence sets.
result The geometry of minimum-volume confidence sets for multinomial parameters is studied, providing insights into their structure and properties.

Building on previous results on the quadratic helicity in magnetohydrodynamics (MHD) we investigate particular minimum helicity states. Those are eigenfunctions of the curl operator and are shown to constitute solutions of the quasi-stationary incompressible ideal MHD equations. We then show that these states have inde…

2018-06-19abs ↗pdf ↗