Paper adds a restart mechanism to a drawdown control policy for better trading performance.
problem Missed profitable opportunities when drawdown limit is close to reality.
method Integrates a data-driven restart mechanism into the drawdown modulation trading system.
result The restart mechanism improves trading performance even with transaction costs.
Cross-Modulation Networks improve few-shot learning by combining information at multiple levels.
problem Few-shot learning challenges in deep feature extraction.
method Integrates support and query examples at various levels of abstraction using a feature-wise modulation mechanism.
result Encouraging initial results on miniImageNet, closing the gap with state-of-the-art.
The geometry of nonholonomic bundle gerbes, provided with nonlinear connection structure, and nonholonomic gerbe modules is elaborated as the theory of Clifford modules on nonholonomic manifolds which positively fail to be spin. We explore an approach to such nonholonomic Dirac operators and derive the related Atiyah-S…
R-SQAIR adds relational bias to sequential object attention models for better object interactions.
problem Traditional sequential multi-object attention models struggle with relational inferences.
method Proposes R-SQAIR, a relational extension of SQAIR with a parallel pairwise interaction module.
result Demonstrates gains in object relations and combinatorial generalization over sequential mechanisms.
NIPA aims to translate brain learning mechanisms into scalable Bayesian inference.
problem Scalable Bayesian inference for large-scale statistical machine learning problems.
method Neural-inspired algorithm combining model-based, model-free, and episodic-control modules.
result Advances Bayesian methods and facilitates their application to deep learning.
Study of cubic skein modules in 3-sphere and arbitrary 3-manifolds.
problem Lack of systematic study of higher degree skein modules.
method Investigation of cubic skein module structure and properties in 3-sphere and arbitrary 3-manifolds.
result Establishment of a foundational framework for higher skein modules.
Algebra Situs is a branch of mathematics which has its roots in Jones' construction of his polynomial invariant of links and Drinfeld's work on quantum groups. It encompasses the theory of quantum invariants of knots and 3-manifolds, algebraic topology based on knots, operads, planar algebras, q-deformations, quantum g…
Differentiable Window improves attention modules by enabling more focused attentions.
problem Improving attention mechanisms in neural networks.
method Proposes Differentiable Window, a neural module for dynamic window selection.
result Consistent and sizable improvements across various NLP tasks.
This paper proposes a new approach to Transformers by integrating hierarchical associative memory with MetaFormers.
problem Theoretical framework for Transformers and MLP-Mixers remains underdeveloped.
method Integrating hierarchical associative memory with MetaFormers to create a parallelized MLP-Mixer.
result Symmetry-breaking effects improve the performance of the MLP-Mixer in image recognition tasks.
This paper presents a Semantic Attribute Modulation (SAM) for language modeling and style variation. The semantic attribute modulation includes various document attributes, such as titles, authors, and document categories. We consider two types of attributes, (title attributes and category attributes), and a flexible a…
V-HMN integrates memory mechanisms for improved image recognition.
problem Limited interpretability and high data requirements of existing vision backbones.
method Brain-inspired hierarchical memory modules with iterative refinement.
result V-HMN achieves strong performance on image classification benchmarks.
AFS uses attention to select features efficiently.
problem Efficiently selecting features from high-dimensional data.
method AFS combines an attention module and a learning module to address feature selection challenges.
result AFS outperforms state-of-the-art feature selection algorithms in accuracy and stability.
Representation stability is a phenomenon whereby the structure of certain sequences Xn of spaces can be seen to stabilize when viewed through the lens of representation theory. In this paper I describe this phenomenon and sketch a framework, the theory of FI-modules, that explains the mechanism behind it.
Gradient Gating improves deep GNNs by modulating message passing updates.
problem Oversmoothing and performance degradation in deep GNNs.
method Gradient gating mechanism for multi-rate message passing.
result G2 framework alleviates oversmoothing and achieves state-of-the-art performance. Paper generalizes kernel mean embedding to von Neumann-algebra-valued measures.
problem Analyzing complex multivariate distributions and quantum mechanics.
method Generalizes kernel mean embedding to von Neumann-algebra-valued measures in reproducing kernel Hilbert modules.
result Injectivity and universality of the generalized KME are confirmed.
Study the algebraic action of torus on knot complement's skein module.
problem Understand the algebraic structure of knot complements and boundary tori.
method Analyze the Kauffman bracket skein algebra and module of the 3-twist knot complement.
result Determine the action of Kauffman bracket skein algebra on module of 3-twist knot complement.
Introduces a novel spatial attention module for convolutional networks.
problem Irregular boundaries in position-wise spatial attention maps hamper model generalization.
method Introduces a convolutional rectangular attention module with 5 parameters.
result Systematically outperforms position-wise counterparts in experiments.
Automatic modulation classification (AMC) is an important task for modern communication systems; however, it is a challenging problem when signal features and precise models for generating each modulation may be unknown. We present a new biologically-inspired AMC method without the need for models or manually specified…
Memory-Augmented Recurrent Networks improve dialogue coherence by expanding conversation history storage.
problem Fixed-size vectors limit dialogue coherence; attention mechanisms are computationally expensive.
method Introduce Neural Turing Machines (NTMs) to provide flexible and permanent storage for dialogue history.
result Improved perplexity performance compared to existing baselines.
Improved few-shot learning with LSSVM and transductive modules.
problem Few-shot learning with limited data and samples.
method Introducing LSSVM as a base learner and transductive modules to enhance classification accuracy.
result FSLSTM achieves state-of-the-art performance on miniImageNet and CIFAR-FS benchmarks.
A new network learns market conditions and predicts stock performance.
problem Optimizing stock portfolio performance in the US equities market.
method Residual Switching Network combining two ResNets: a switching module and a main module.
result The residual switching network strategy outperformed other models with an average annual Sharpe ratio of 2.22.
Statistical learning relies upon data sampled from a distribution, and we usually do not care what actually generated it in the first place. From the point of view of causal modeling, the structure of each distribution is induced by physical mechanisms that give rise to dependences between observables. Mechanisms, howe…
Kriformer uses graph transformers to estimate data in sparse sensor areas.
problem Sparse sensor deployment and unreliable data in spatiotemporal kriging tasks.
method Graph transformer model with positional encoding and attention mechanisms.
result Kriformer excels in representing unobserved locations in spatiotemporal kriging tasks.
TEAFormers preserve multi-dimensional time series structures for better forecasting.
problem Traditional Transformers flatten multi-dimensional time series data, losing critical multi-dimensional relationships.
method Tensor-Augmented Transformer (TEAFormer) with Tensor-Augmentation (TEA) module.
result Significant performance enhancements in time series forecasting across benchmarks.
Paper introduces methods to integrate external knowledge into RNNs using attention mechanisms.
problem Incorporating external knowledge into RNNs for improved performance.
method Proposes three methods: attentional concatenation, feature-based gating, and affine transformation.
result Attentional feature-based gating consistently improves performance across tasks.
Proposes GPCA module for channel attention in CNNs using Gaussian processes.
problem Improving performance in visual tasks through effective channel selection.
method Integrates Gaussian processes into channel attention mechanisms for probabilistic modeling of channel correlations.
result Demonstrates improved performance of GPCA module in end-to-end CNN training.
GAT-AGNN learns stock trends using graph and attention mechanisms.
problem Predicting dynamic stock trends in a complex market.
method Sequential graph structure with attention mechanisms.
result GAT-AGNN outperforms state-of-the-art methods in stock trend prediction.
Paper proposes a method to control robots of different shapes efficiently.
problem Learning optimal control policies for robots of various shapes is challenging.
method Hierarchical architecture with hypernetworks and fixed attention mechanism.
result Method improves learning performance and generalizes to unseen morphologies.
We describe a mechanism by which artificial neural networks can learn rapid adaptation - the ability to adapt on the fly, with little data, to new tasks - that we call conditionally shifted neurons. We apply this mechanism in the framework of metalearning, where the aim is to replicate some of the flexibility of human …
BiSHop tackles tabular data challenges with sparse Hopfield layers.
problem Non-rotationally invariant data structure and feature sparsity in tabular data.
method Sequential column-wise and row-wise processing through interconnected directional learning modules with generalized sparse modern Hopfield layers.
result BiSHop surpasses current SOTA methods with significantly less hyperparameter tuning.
Pyramid Attention Networks improve image restoration by leveraging self-similarities across scales.
problem Lack of full exploitation of self-similarities in image restoration by recent deep learning methods.
method Introduces a Pyramid Attention module that processes multi-scale feature correspondences to borrow clean signals from coarser levels.
result Pyramid Attention module achieves state-of-the-art results in various image restoration tasks.
Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose the use of such computational modules for extracting invariant and discriminativ…
SuTaT creates dialogue summaries for tete-a-tetes without labeled data.
problem Lack of high-quality paired dialogue-summary data.
method Unsupervised model for tete-a-tetes, modeling customer and agent roles separately.
result SuTaT outperforms on automatic and human evaluations.
AReLU uses attention-based rectification to improve neural network performance.
problem Improving neural network performance through better activation functions.
method Integrates attention mechanism with rectified linear unit (ReLU) to learn and scale feature maps.
result AReLU significantly boosts performance of most network architectures with minimal changes.
New technique improves imitation learning by preventing local minima and exploring states.
problem Behavioral cloning gets stuck in local minima and lacks effective exploration.
method Two-phase model with sampling mechanisms and self-attention modules.
result Significantly outperforms previous state-of-the-art in various environments.
CSML learns causal structures for few-shot learning.
problem Spurious correlations limit deep learning generalization.
method CSML combines perception, causal induction, and reasoning modules.
result CSML achieves superior few-shot learning across tasks.
Adversaries can fool deep learning modulator classifiers over wireless channels.
problem Vulnerability of deep learning modulator classifiers to adversarial attacks over wireless channels.
method Presented various adversarial attacks considering channel effects, including targeted and non-targeted attacks, and a universal adversarial perturbation attack.
result Modulation classification is vulnerable to adversarial attacks over wireless channels with realistic channel effects.
CRAUM-Net improves salient object detection with context and uncertainty modeling.
problem Accurate salient object detection with precise boundary delineation.
method Contextual Recursive Attention with Uncertainty Modeling, multi-scale context aggregation, attention mechanisms, edge-aware decoder, Monte Carlo Dropout.
result Superior performance in producing accurate and reliable saliency maps.
Paper introduces DNTs to clone black-box models efficiently.
problem Cloning functionality of black-box models.
method Deep Neural Trees (DNTs) trained with active learning.
result Trained DNT can clone task-specific behavior of black-box models.
Complex analysis techniques link Gaussian RBF kernels to quantum mechanics.
problem Understanding the Gaussian RBF kernel in machine learning and SVMs.
method Using Fock space and Segal-Bargmann theories in complex analysis.
result Proves connections between Gaussian RBF kernels and quantum mechanics operators.
MPSA-DenseNet improves accent classification accuracy.
problem Accurate English accent identification.
method Combines multi-task learning and PSA attention mechanism with DenseNet.
result MPSA-DenseNet outperforms other models in accent classification.
Extends driving model to control agent behavior in simulations.
problem Simulate realistic driving behavior for autonomous systems.
method Introduces Control-ITRA method to influence agent behavior through waypoint assignment and target speed modulation.
result Demonstrates controllable, infraction-free trajectories while preserving realism.
Variance-Calibrated Modulation (VCM) addresses the likelihood trap in LLMs by reshaping the probability distribution before truncation.
problem LLMs fall into the likelihood trap, leading to repetitive degeneration and vocabulary dullness.
method VCM reshapes the probability distribution before truncation through Contextual Searchlight and Adaptive Self-Debiasing.
result VCM mitigates the likelihood trap across open-ended generation, factual QA, and mathematical reasoning.
OLS predictions are shown to be similar to attention mechanisms in models.
problem OLS in traditional statistics and econometrics.
method Rewriting OLS as an attention mechanism in a transformed space.
result OLS can be understood as minimizing squared prediction errors via optimal embedding and decoding.
Representation stability is a theory describing a way in which a sequence of representations of different groups is related, and essentially contains a finite amount of information. Starting with Church-Ellenberg-Farb's theory of FI-modules describing sequences of representations of the symmetric groups, we now have …
Graph neural networks improve equipment health monitoring from multisensor data.
problem Leveraging complex machinery structure for condition-based maintenance.
method Captured machinery structure as a graph and used graph neural networks (GNNs) to model time-series data.
result GNN-based RUL estimation model outperforms RNNs and CNNs on turbofan engine benchmark.
A framework uses attention mechanisms to optimise financial portfolios by reducing noise and balancing returns.
problem Balancing investment returns and risks in noisy financial markets.
method Multi-agent framework with attention mechanisms and time series analysis.
result MASAAT framework produces more balanced portfolios with enhanced performance.
Unified framework for multi-objective curriculum learning in robotics.
problem Improving sample efficiency and final performance in robotic policy learning.
method Unified automatic curriculum learning framework with multi-task hyper-net and flexible memory mechanism.
result Superior performance compared to state-of-the-art methods in robotic manipulation tasks.