Minimum attention improves reinforcement learning performance in high-dimensional dynamics.
problem Improving reinforcement learning performance in high-dimensional nonlinear dynamics.
method Applying minimum attention as a regularization technique in reinforcement learning, including model-based and model-free approaches.
result Minimum attention outperforms state-of-the-art algorithms in few-shot adaptation and variance reduction.
Sequence-to-sequence models, such as attention-based models in automatic speech recognition (ASR), are typically trained to optimize the cross-entropy criterion which corresponds to improving the log-likelihood of the data. However, system performance is usually measured in terms of word error rate (WER), not log-likel…
Attention models can overfit without harming test performance.
problem Understanding benign overfitting in single-head attention models.
method Analyzing conditions for benign overfitting in a single-head softmax attention model.
result A single-head attention model can overfit without harming test performance under certain conditions.
Inspired by the adaptation phenomenon of neuronal firing, we propose the regularity normalization (RN) as an unsupervised attention mechanism (UAM) which computes the statistical regularity in the implicit space of neural networks under the Minimum Description Length (MDL) principle. Treating the neural network optimiz…
MCSAE improves speaker embedding by focusing on both high- and low-level features.
problem Reduced effect of low-level features in speaker embedding encoding.
method Masked cross self-attentive encoding using ResNet with multi-layer aggregation and random masking regularization.
result Improved speaker embedding with equal error rate of 2.63% and minimum detection cost function of 0.1453.
Over the past decade, multivariate time series classification has received great attention. We propose transforming the existing univariate time series classification models, the Long Short Term Memory Fully Convolutional Network (LSTM-FCN) and Attention LSTM-FCN (ALSTM-FCN), into a multivariate time series classificat…
A-MMSE uses attention to learn efficient OFDM channel estimation.
problem Accurate OFDM channel estimation requires second-order statistics, which are hard to obtain in practice.
method A-MMSE is a model-based DNN framework that learns linear MMSE filters via Attention Transformer, reducing inference complexity.
result A-MMSE outperforms other methods in normalized MSE across various SNR conditions.
Approximations of loopy belief propagation, including expectation propagation and approximate message passing, have attracted considerable attention for probabilistic inference problems. This paper proposes and analyzes a generalization of Opper and Winther's expectation consistent (EC) approximate inference method. Th…
In links with two components there are three different types of crossings: self-crossings in the first component, self crossings in the second component, and crossings between components. In this paper we examine the minimum number of crossing changes needed to unlink without changing the crossings between components. …
Transformers learn linear models in-context without updates.
problem Understanding how transformers mimic linear models in-context.
method Gradient flow on linear regression tasks with random initialization.
result Transformers achieve prediction error competitive with best linear predictors.
CutMix training technique improves spatial locality in Vision Transformers.
problem Improving spatial locality in Vision Transformers trained from scratch.
method Comparison of Baseline and Modern training protocols on CIFAR-10, CIFAR-100, and Tiny-ImageNet.
result CutMix training component significantly reduces Mean Attention Distance (MAD) in early layers of Vision Transformers.
Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch methods tend to converge to sharp minimizers has received increasing attention. We theoretically justify this hypothesis by providing new pro…
A new optimizer DDC improves deep learning models by respecting symmetries.
problem Deep networks' loss is invariant to continuous symmetries, leading to optimization issues.
method DDC builds a Dead-Direction Conditioner that lifts a base optimizer into a G-equivariant one, preserving the quotient geometry.
result DDCAdam and DDCMuon outperform standard optimizers in various tasks, improving validation-train loss gaps and learning dynamics.
This article is written for the Proceedings of the Conference on Current Developments in Mathematics in Harvard University, November 16-17, 2007. It is an exposition of the analytic proof of the finite generation of the canonical ring for a compact complex algebraic manifold of general type. It lists and discusses the …
ESE-FN improves elderly activity recognition accuracy.
problem Recognizing individual actions and human-object interactions in elderly activities.
method Exploits multi-modal features from RGB videos and skeleton sequences using ESE attentions and a new Multi-modal Loss.
result ESE-FN achieves best accuracy on ETRI-Activity3D dataset.
Paper proposes MWDE for estimating finite location-scale mixtures.
problem Estimating finite location-scale mixtures using MLE is problematic.
method Investigates minimum Wasserstein distance estimators (MWDE).
result MWDE is consistent and provides a numerical solution.
This paper optimizes portfolios using HRP and CLA algorithms on NIFTY 50 stocks.
problem Designing an optimal stock portfolio with accurate forecasting of future returns and risks.
method Uses hierarchical risk parity and critical line algorithms on NIFTY 50 stocks.
result Hierarchical risk parity algorithm outperformed the critical line algorithm on test data.
Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures are comparable to stat…
Paper proposes a generalized precision matrix for t-Student distributions to improve portfolio optimization.
problem Limitations of inverse covariance matrix in non-Gaussian settings.
method Exploits local dependence function to define generalized precision matrix (GPM) for multivariate t-Student distribution.
result GPM leads to statistically significant lower out-of-sample variances in minimum-variance portfolios.
Reformulates binary classification on manifolds using Yang-Mills-Higgs theory.
problem Binary classification on non-contractible spaces.
method Formulates binary classification as a Yang-Mills-Higgs variational problem, encoding data as a functor.
result Reveals a geometric interpretation of binary classification and solves XOR on the torus.
Transfer learning improves MNI's performance in high-dimensional linear regression.
problem Improving model performance in high-dimensional linear regression with diverse data.
method Proposes a Transfer MNI approach, analyzing its excess risk and conditions for outperformance.
result Identifies free-lunch covariate shift regimes where knowledge transfer benefits.
In this paper we continue to study (`strong') Nielsen coincidence numbers (which were introduced recently for pairs of maps between manifolds of arbitrary dimensions) and the corresponding minimum numbers of coincidence points and pathcomponents. We explore compatibilities with fibrations and, more specifically, with c…
Paper uses algebraic signatures to identify probabilistic structures in empirical data.
problem Identifying probabilistic structure from observed binomials in empirical probability tensors.
method Treating vanishing binomials as algebraic signatures, matching signatures to identify models without parameter estimation.
result The method successfully identified rank-one structures in real language data, revealing interpretable sets of words.
Deep convolutional semantic segmentation (DCSS) learning doesn't converge to an optimal local minimum with random parameters initializations; a pre-trained model on the same domain becomes necessary to achieve convergence.In this work, we propose a joint cooperative end-to-end learning method for DCSS. It addresses man…
A novel hypergraph partitioning method using tensor eigenvalue decomposition captures super-dyadic interactions.
problem Capturing super-dyadic interactions in k-uniform hypergraphs.
method Tensor-based representation and tensor eigenvalue decomposition for capturing interactions.
result Improved min-cut solution on 2-uniform hypergraphs (graphs) compared to standard spectral partitioning.
Paper analyzes robustness of MDPDE under INH setups.
problem Global reliability and breakdown behavior of MDPDE under INH.
method Asymptotic breakdown point analysis of MDPDE.
result Derives a theoretical lower bound for the asymptotic breakdown point.
This paper proves SGD converges to global minimum for over-parameterized ReLU networks.
problem Theoretical understanding of implicit neural networks is limited.
method Gradient flow analysis of ReLU activated implicit neural networks.
result Randomly initialized gradient descent converges to global minimum at a linear rate for square loss function in over-parameterized ReLU networks.
This paper studies long term investing by an investor that maximizes either expected utility from terminal wealth or from consumption. We introduce the concepts of a generalized stochastic discount factor (SDF) and of the minimum price to attain target payouts. The paper finds that the dynamics of the SDF needs to be c…
This paper proves IRM minimizes o.o.d. risk under certain conditions.
problem Deep networks can fail to generalize to new domains with different distributions.
method Proves IRM minimizes o.o.d. risk through a bi-level optimization problem.
result IRM minimizes o.o.d. risk under specific conditions.
Minimum Description Length prevents overfitting in noisy data.
problem Learning from noisy data with overfitting risk.
method Minimum Description Length learning rule with tempered guarantees.
result Tempered agnostic finite sample learning guarantees and asymptotic behavior characterization.
Verifying robustness of neural network classifiers has attracted great interests and attention due to the success of deep neural networks and their unexpected vulnerability to adversarial perturbations. Although finding minimum adversarial distortion of neural networks (with ReLU activations) has been shown to be an NP…
The minimum number of colors is a challenging knot invariant since, by definition, its calculation requires taking the minimum over infinitely many minima. In this article we estimate and in some cases calculate the minimum number of colors for the Turk's head knots on three strands.
Minimum braids are a complete invariant of knots and links. This paper defines minimum braids, describes how they can be generated, presents tables for knots up to ten crossings and oriented links up to nine crossings, and uses minimum braids to study graph trees, amphicheirality, unknotting numbers, and periodic table…
Study tightens bounds for interpolating noisy data using minimum l1-norm.
problem Predicting noisy data with minimum l1-norm interpolation.
method Provided matching upper and lower bounds for prediction error.
result Tight consistency up to negligible terms for d≫n. A new classification method based on Minimum Spanning Trees
problem Improving classification in supervised learning
method Proposing a classification algorithm based on Minimum Spanning Trees
result The proposed method is effective and computationally efficient
Minimum-norm solutions generalize well in over-parametrized neural networks.
problem Generalization error in over-parametrized neural networks.
method Analyzing three models: random feature model, two-layer neural network, and residual network.
result Generalization error for minimum-norm solutions is comparable to Monte Carlo rate, up to logarithmic terms.
The paper finds minimum Dehn colors for knots and defines useful graphs for coloring.
problem Finding the minimum number of colors for Dehn colorings of knots.
method Analyzes Dehn colorings for knots and defines R-palette graphs. result For Dehn p-colorable knots, the minimum number of colors is at least ⌊log2pfloor+2. The paper calculates genus bounds for multibranched surfaces.
problem Finding genus bounds for multibranched surfaces.
method Using the first Betti number and boundary genus, the paper provides lower bounds for maximum and minimum genus.
result The maximum and minimum genus of GimesS1 equals twice that of G. Knots are commonly found in molecular chains such as DNA and proteins, and they have been considered to be useful models for structural analysis of these molecules. One interested quantity is the minimum number of monomers necessary to realize a molecular knot. The minimum lattice length $\mbox{Len}(K)$ of a knot K i…
Study shows how networks converge to minimum norm solutions with regularization.
problem Interpolating between known regions in shallow ReLU networks.
method Investigates empirical risk minimizers and weight decay regularizers.
result Empirical risk minimizers converge to minimum norm interpolants under specific conditions.
We find the minimum dilatation of pseudo-Anosov braids with many strands.
problem Finding the minimum dilatation of pseudo-Anosov braids with a large number of strands.
method Analyzing examples of Hironaka-Kin and Venzke to determine the minimum dilatation.
result The minimum dilatation is approximately 13.928 for large n. Paper proposes neural network for efficient MIMO channel estimation and pilot reduction.
problem High overhead from pilot transmission in wideband MIMO systems.
method Neural network architecture for frequency-aware pilot design and channel estimation, with pruning technique.
result Neural network outperforms linear minimum mean square error (LMMSE) estimation.
Study introduces AMVP and AMRR for dynamic portfolio optimization in volatile markets.
problem Optimizing portfolios in volatile and nonstationary financial markets.
method Adaptive Minimum-Variance Portfolio (AMVP) framework with ARFIMA-FIGARCH processes and non-Gaussian innovations.
result Demonstrated superior performance in risk reduction and portfolio stability during market breaks.
Minimum algebraic intersection found in hyperbolic surfaces, growing with genus.
problem Finding the minimum algebraic intersection form in hyperbolic surfaces.
method Analyzing algebraic intersection form in moduli space of hyperbolic surfaces.
result Minimum grows in the order of (logg)−2 with genus. We refine Expected Shortfall by controlling different tail portions, offering tailored risk assessments.
problem Risk assessment in financial positions, especially in tail regions.
method Introducing adjusted Expected Shortfall measures that control different tail portions.
result Adjusted Expected Shortfall measures ensure risk does not exceed specified thresholds for various probability levels.
Computed minimum crossing numbers for Turaev genus 2 links.
problem Verifying the Qazaqzeh-Chbili-Lowrance conjecture.
method Computed minimum crossing numbers for a specific family of links.
result Verified the Qazaqzeh-Chbili-Lowrance conjecture for the family.
This paper studies the geometry of minimum-volume confidence sets for multinomial parameters.
problem Determining if minimum-volume confidence sets for multinomial outcomes are disjoint.
method Enumerating and covering the continuous regions of the exact p-value function to study the geometry of minimum-volume confidence sets.
result The geometry of minimum-volume confidence sets for multinomial parameters is studied, providing insights into their structure and properties.
Building on previous results on the quadratic helicity in magnetohydrodynamics (MHD) we investigate particular minimum helicity states. Those are eigenfunctions of the curl operator and are shown to constitute solutions of the quasi-stationary incompressible ideal MHD equations. We then show that these states have inde…