Study optimizes resource allocation in noisy systems for better control.
problem Limited attention in stochastic systems with multiplicative noise.
method Analytical and numerical methods for optimal attention allocation.
result Effective resource allocation enhances noise estimation and control decisions.
ANCDEs improve time-series forecasting and classification using attention in NCDEs.
problem Improving time-series forecasting and classification using neural controlled differential equations.
method Integrating attention into neural controlled differential equations (ANCDEs).
result ANCDEs consistently show the best accuracy in time-series classification and forecasting.
Minimum attention improves reinforcement learning performance in high-dimensional dynamics.
problem Improving reinforcement learning performance in high-dimensional nonlinear dynamics.
method Applying minimum attention as a regularization technique in reinforcement learning, including model-based and model-free approaches.
result Minimum attention outperforms state-of-the-art algorithms in few-shot adaptation and variance reduction.
Simple attention model outperforms complex sEMG classifiers.
problem Improving myoelectric control for robotic prosthetics.
method Attention-based model for sEMG signal classification.
result Simple model achieves benchmark results on multiple datasets.
Paper proposes a method to control robots of different shapes efficiently.
problem Learning optimal control policies for robots of various shapes is challenging.
method Hierarchical architecture with hypernetworks and fixed attention mechanism.
result Method improves learning performance and generalizes to unseen morphologies.
Partially observable environments present an important open challenge in the domain of sequential control learning with delayed rewards. Despite numerous attempts during the two last decades, the majority of reinforcement learning algorithms and associated approximate models, applied to this context, still assume Marko…
This work explains the structural origins of attention sinks in LLMs.
problem Initial tokens disproportionately monopolize attention scores in LLMs.
method Traced to self-attention's value aggregation process and FFN layer activations.
result Attention sinks form due to variance discrepancy and dimension disparity.
Adaptive attention span extends Transformer's context size.
problem Limited context size in Transformers limits model performance.
method A self-attention mechanism that learns optimal attention span.
result State-of-the-art performance on character-level language modeling.
Typical reinforcement learning (RL) agents learn to complete tasks specified by reward functions tailored to their domain. As such, the policies they learn do not generalize even to similar domains. To address this issue, we develop a framework through which a deep RL agent learns to generalize policies from smaller, s…
ExGate improves neural network accuracy with feature-based attention.
problem Efficiency and accuracy in artificial neural networks for feature-based attention.
method Externally controlled neuron gating in artificial neural networks.
result 5% increase in classification accuracy on CIFAR-10 dataset.
FIGARO generates symbolic music with fine-grained control.
problem Minimal control over generated music sequences.
method Description-to-sequence task, learning conditional distribution of sequences given high-level descriptions.
result State-of-the-art controllable symbolic music generation.
Unified framework for sparse alternatives to softmax with control over sparsity.
problem Lack of understanding and explicit control over sparsity in probability mapping functions.
method Unified framework encompassing softmax, sum-normalization, spherical softmax, and sparsemax. Two novel sparse formulations (sparsegen-lin and sparsehourglass) and convex loss functions developed.
result Improved performance in multilabel classification and seq2seq tasks like neural machine translation and abstractive summarization.
HKT improves sequence processing with multi-scale attention and kernel analysis.
problem Processing sequences at multiple scales with efficient attention mechanisms.
method Trainable causal downsampling and convex weights for level-specific score matrices.
result HKT achieves consistent gains over standard attention across various tasks.
Transformers mimic Bayesian reasoning in controlled settings, revealing geometric mechanisms.
problem Verifying if transformers perform Bayesian reasoning rigorously in natural data.
method Constructing Bayesian wind tunnels with known posteriors and proving memorization impossibility.
result Transformers achieve 10−3-10−4 bit accuracy in Bayesian posteriors, while MLPs fail. Unified framework for convolution and attention models.
problem Complex neural network structures and their parameter control.
method Unified framework for convolution and attention models.
result Attention models are a special case of convolution with adaptive structure.
In this paper, we put the issue of dynamic equivalence of control systems in the context of pullbacks of coframings on infinite jet bundles over the state manifolds. While much attention has been given to differentially flat systems, i.e. systems dynamically equivalent to linear control systems, the advantage of this a…
Analyzing a comprehensive news dataset, we document that joint news coverage triggers attention contagion, causing temporarily inflated valuations for affected stocks. Tracing SEC EDGAR visits from unique IPs, we provide direct evidence of attention spillovers between stocks. Stocks with greater joint news coverage exh…
Investor selects portfolios based on news attention in a hidden Markov model.
problem Mean-variance portfolio selection in a dynamic attention context.
method Closed-loop equilibrium strategies via extended HJB equation and Markov chain approximation.
result Equilibrium strategies found through iterative algorithm and numerical examples.
New approach to understand recurrent policies as FSMs without minimization.
problem Minimization of FSMs obscures the semantics of policy decisions.
method Start with unminimized FSM, apply interpretable reductions, use attention tool.
result Reveals insights into policy decisions not previously noticed.
Quantum Transformer solves high-dimensional PDEs with improved accuracy.
problem Solving high-dimensional parabolic PDEs in engineering and physics.
method Quantum Transformer BSDE solver using FC-VQC with causal attention.
result Quantum Transformer consistently outperforms classical methods on PDE benchmarks.
We propose the Gaussian attention model for content-based neural memory access. With the proposed attention model, a neural network has the additional degree of freedom to control the focus of its attention from a laser sharp attention to a broad attention. It is applicable whenever we can assume that the distance in t…
Transformers can approximate Kalman Filtering in linear systems with small error.
problem Approximating Kalman Filtering using Transformers for linear dynamical systems.
method Two-step reduction: 1) Softmax self-attention block approximates Nadaraya-Watson kernel smoothing, 2) This estimator approximates Kalman Filter.
result Constructs a Transformer that implements the Kalman Filter with small additive error, uniformly bounded in time.
Automates decision-making for human operators managing multiple robots.
problem Limited human operator attention when controlling multiple robots.
method Learned model of user preferences from easy settings to automatically identify the most critical robot.
result Automated decision-making can assist human operators in managing more robots than their attention allows.
Graph neural networks benefit from attention under specific conditions.
problem Understanding and improving the effectiveness of attention in graph neural networks.
method Designing controlled graph reasoning tasks, analyzing performance under various conditions, proposing weakly-supervised training.
result Attention can provide significant gains in performance under certain conditions, but its effect is often negligible or harmful.
This work analyzes self-attention matrices using random matrix theory.
problem Understanding the theoretical behavior of self-attention layers in neural networks.
method Asymptotic spectral analysis of the attention matrix, Gaussian equivalence, and linearization.
result The singular value distribution of the attention matrix is asymptotically characterized by a linear model.
Study proposes a statistical test for Vision Transformer's attention mechanisms.
problem ViT's attention mechanisms may focus on irrelevant regions, leading to unreliable evidence.
method Selective inference framework to quantify statistical significance of attentions as p-values.
result Proposed method enables reliable quantification of false positive detection probability of attentions.
RB-Modulation trains free diffusion models without external adapters.
problem Training-free personalization of diffusion models with style and content control.
method Stochastic optimal control with a style descriptor and cross-attention aggregation.
result Precise content and style extraction and control without external adapters.
New model DWM mimics human memory, learning to retain, ignore or forget.
problem Lack of effective separation between episodic and working memory in neural networks.
method Designed Differentiable Working Memory (DWM) inspired by psychological studies.
result DWM learns psychology-inspired tasks faster and generalizes to longer sequences.
Adaptively sparse Transformers improve interpretability and diversity in NLP.
problem Standard Transformers use dense attention, limiting interpretability and diversity.
method Introduces adaptively sparse Transformers using α-entmax for context-dependent sparsity. result Improves interpretability and diversity in NLP tasks without sacrificing accuracy.
The p-Laplacian Transformer improves transformer models by assigning higher attention weights to tokens in close proximity.
problem The self-attention mechanism in transformers does not effectively distinguish attention weights between tokens in close and non-close proximity.
method Proposes a novel class of transformers, p-Laplacian Transformers, that use p-Laplacian regularization to assign higher attention weights to tokens in close proximity. result Empirically demonstrates that p-Laplacian Transformers outperform baseline transformers on various benchmark datasets.
Neural operators approximate Stackelberg game solutions.
problem Intractability of follower's best-response operator in dynamic Stackelberg games.
method Used attention-based neural operators to approximate the best-response operator.
result Approximate best-response operator yields close game value.
Attention temperature improves robustness of ICL in high-dimensional settings.
problem ICL robustness failure under distribution shift in high dimensions.
method Analyzed a Transformer with approximate softmax attention, derived a closed-form error expression, and showed optimal temperature minimizes error.
result Optimal attention temperature minimizes ICL generalization error under distribution shift.
A new convenient method of describing flat convex compact sets is proposed. It generalizes classical trigonometric functions sin and cos. Apparently, this method may be very useful for explicit description of solutions of optimal control problems with two-dimensional control. Using this method a series of sub-Fin…
Improved image synthesis with user scribbles and text prompts.
problem Inadequate details in generated images due to domain shift.
method Optimization problem formulation and cross-attention for control.
result Significant improvement in user satisfaction (85.32% higher).
Proposes RN for unsupervised attention in neural networks.
problem Limited, imbalanced, and non-stationary input distributions in various tasks.
method Inspired by neuronal adaptation, RN uses MDL principle and universal code length for incremental layer-wise computation.
result Outperforms existing normalization methods across diverse tasks.
A new method calibrates scientific models by adding randomness to their predictions.
problem Current scientific foundation models lack calibrated uncertainty.
method Stochastic Attention, which randomizes attention weights using multinomial samples.
result Stochastic Attention achieves the strongest native calibration and sharpest prediction intervals.
ChatGPT launch boosted AI-related crypto assets by 10.7% to 15.6%.
problem Investor perception of AI assets after ChatGPT launch.
method Synthetic difference-in-difference methodology.
result AI-related crypto assets experienced significant returns after ChatGPT launch.
The study compares feed-forward and attention layers in language models.
problem Understanding the role of feed-forward and attention layers in language models.
method Empirical and theoretical analysis in a synthetic setting.
result Feed-forward layers learn simple distributional associations, while attention layers focus on in-context reasoning.
The paper presents a model-free method for stabilizing unknown control systems.
problem Stabilizing unknown control systems in engineering.
method Solving discounted LQR problems with increasing discount factors.
result The method efficiently recovers a stabilizing controller for linear and smooth nonlinear systems.
A new framework uses deep reinforcement learning to improve aircraft separation in busy airspace.
problem Improving aircraft separation in high-density, dynamic airspace constrained by human controllers.
method Proximal Policy Optimization with an attention network for distributed vehicle autonomy.
result The framework significantly reduces offline training time and increases performance.
Unified thermodynamic approach to Transformer attention dynamics.
problem Understanding the statistical mechanics of Transformer attention.
method Constructing a Lagrangian on the information manifold to analyze attention dynamics.
result Establishes a formal correspondence between scaled dot-product attention and canonical ensemble statistics.
Study uses DRL with Lagrangian relaxation to solve temporal control tasks with STL constraints.
problem Optimal control problems with temporal logic constraints.
method Extended CMDP formulation, Lagrangian relaxation, two-phase constrained DRL algorithm.
result Demonstrated learning performance of the proposed algorithm through simulations.
Transformer models align words through attention weights, closely approximating Optimal Transport.
problem Understanding the internal mechanism of transformer models in language processing.
method Empirical evidence and theoretical analysis of attention weights and their relation to Optimal Transport.
result Transformer models can simulate gradient descent on the dual of entropy-regularized OT problem, providing a theoretical foundation for token alignment.
Two prediction models improve supply-demand forecasting for autonomous vehicles.
problem Improving accuracy and stability of supply-demand predictions for autonomous vehicles.
method Two prediction models based on residual network, LSTM, attention mechanism, and multi-attention mechanism.
result Our frameworks provide more accurate and stable prediction results than existing methods.
GPT learns a causal world model from token predictions, validated in game sequences.
problem Does GPT implicitly learn a causal world model from token predictions?
method Derived a causal interpretation of GPT's attention mechanism and proposed zero-shot causal structure learning.
result GPT can generate legal next moves with high confidence for sequences with encoded causal structures, but fails for illegal moves.
The aim of this paper is to generalize the PAC-Bayesian theorems proved by Catoni in the classification setting to more general problems of statistical inference. We show how to control the deviations of the risk of randomized estimators. A particular attention is paid to randomized estimators drawn in a small neighbor…
This work reviews left-invariant optimal control problems on Lie groups.
problem Optimal control problems on Lie groups with big symmetry.
method Review of main notions, methods, and results.
result Description of extremal trajectories and their optimality, cut time and cut locus, optimal synthesis.
The paper enhances a virtual assistant's humor to improve user satisfaction.
problem Improving a virtual assistant's ability to deliver humorous responses.
method Combines traditional NLP techniques with self-attentional networks and multi-task learning, using implicit feedback for labeling.
result Deep-learning models outperform heuristic methods in real-world user satisfaction.