Improved Mamba model for long-range sequence tasks.
problem Mamba's poor performance on long-range sequential tasks.
method Proposed B2S6, combining block-wise selective dynamics and channel-specific bias.
result Empirically, B2S6 outperforms S4 and S4D on LRA tasks.
SOR-Mamba improves Mamba for robust time series forecasting by minimizing channel order bias.
problem Robust time series forecasting with Mamba's sequential order bias.
method SOR-Mamba incorporates regularization to minimize channel order discrepancy and introduces CCM for channel correlation preservation.
result SOR-Mamba enhances robustness to channel order and improves forecasting accuracy.
Mamba Hawkes Process improves modeling of event sequences with long-term dependencies.
problem Modeling mutual inhibition and nonlinearity in asynchronous event sequences.
method Introduces Mamba Hawkes Process using Mamba state space architecture.
result MHP outperforms existing models across various datasets.
Mamba improves time series forecasting by quantifying uncertainty.
problem Mamba forecasts have high mean errors in benchmarks.
method Dual-network framework for probabilistic forecasting.
result Predictive uncertainty reduced significantly for both synthetic and real-world data.
Bi-Mamba model predicts diffusion coefficients and exponents from short data.
problem Characterizing anomalous diffusion in complex systems.
method Bidirectional state-space deep learning architecture.
result Efficient inference of diffusion coefficient and exponent from short trajectories.
Mamba efficiently learns low-dimensional targets in-context via feature extraction.
problem Learning low-dimensional targets in context for computational efficiency.
method Test-time feature learning of a single-index model using Mamba's pretrained linear-time sequence model.
result Mamba achieves efficient in-context learning of low-dimensional targets via feature extraction.
MambaLRP enhances Mamba models' explainability and performance.
problem Lack of transparency in Mamba models for real-world applications.
method Layer-wise Relevance Propagation (LRP) with relevance conservation axioms.
result MambaLRP provides stable and reliable explanations for Mamba models.
SAMBA predicts stock returns efficiently using Mamba and graph neural networks.
problem Accurate stock price predictions for financial returns.
method SAMBA integrates Mamba architecture with graph neural networks to achieve near-linear computational complexity.
result SAMBA significantly outperforms state-of-the-art models in prediction accuracy.
State-space models improve dynamical system predictions efficiently and accurately.
problem Challenges in predicting dynamical systems, including long-time integration and long-range dependencies.
method State-space models implemented in Mamba, addressing limitations of existing architectures.
result Mamba outperforms other models in interpolation and challenging extrapolation tasks.
Mamba struggles with long context lengths, but spectrum scaling improves performance.
problem Mamba's performance degrades with increasing context length.
method Spectrum scaling applied to pre-trained Mamba models to improve long-context generalization.
result Spectrum scaling significantly improves performance in long-context settings.
Replicates and improves Uniswap V3 model using DDQN and Mamba.
problem Improving liquidity provision in Uniswap V3 with reinforcement learning.
method Combines DDQN with Mamba and introduces a new reward function.
result Shows stronger theoretical support and better performance than original model.
Study compares LSTM, Transformer, and Mamba for bladder cancer recurrence analysis.
problem Complex time-dependent data in bladder cancer recurrence analysis.
method Evaluation of LSTM, Transformer, and Mamba models using Cox proportional hazards model.
result LSTM-Cox model outperforms Transformer-Cox and Mamba-Cox models in prediction accuracy.
Study proposes a model to improve patient subtyping from EHR data.
problem Challenges in subtyping temporal EHR datasets.
method Self-supervised Mamba-based model for learning EHR representations.
result Model outperforms baseline models in EHR data subtyping.
MAMBA learns policies competitive with multiple conflicting oracles.
problem Learning policies from multiple conflicting oracles in reinforcement learning.
method MAMBA uses a gradient estimator in the style of GAE to optimize policies, leveraging demonstrations from multiple weak oracles.
result MAMBA outperforms the state-of-the-art in learning policies competitive with multiple conflicting oracles.
MambaStock predicts stock prices with high accuracy using a state space model.
problem Inaccurate stock price predictions due to nonlinearity in stock market data.
method Mamba-based state space model with selection mechanism and scan module.
result MambaStock outperforms previous methods in stock price prediction accuracy.
A new method predicts non-Markovian closure terms for complex systems.
problem Predicting the effect of unresolved variables on resolved dynamics in high-dimensional systems.
method Mamba-Assisted Closure (MAC) framework: sequence model trained to predict closure from resolved trajectory, coupled with reduced-order equations.
result Substantially outperforms existing methods in predictive accuracy and long-time stability.
KalMamba improves RL efficiency with probabilistic SSMs.
problem Efficiency in learning and inference for probabilistic SSMs in RL.
method Combines Mamba's scalability with Kalman filtering for efficient probabilistic SSMs.
result KalMamba outperforms state-of-the-art SSMs in RL, especially on longer sequences.
Mamba outperforms Reformer in minute-level stock prediction using LLM sentiment scores.
problem Improving minute-level stock market prediction accuracy in volatile markets.
method Combining sentiment scores from top LLMs with stock price data, training Mamba and Reformer models.
result Mamba achieved lower error rates across all tested LLMs, especially with LLaMA 3.3--70B.
Framework predicts Navier-Stokes solutions on 2D domains using graph neural networks.
problem Predicting stationary Navier-Stokes solutions in non-parametrized 2D geometries.
method Graph-based multi-fidelity learning framework combining reduced-order models, Transformers, and Mamba architectures.
result Mamba architecture reduces computational cost while maintaining performance.
Swift Hydra uses RL and generative AI to improve anomaly detection.
problem Generalization to unseen anomalies in critical systems.
method Generative AI and reinforcement learning (RL) for synthesizing diverse anomaly samples.
result Swift Hydra outperforms state-of-the-art models on ADBench benchmark.
Transformer-based multi-scale model outperforms traditional methods in solving PDEs on irregular domains.
problem Solving partial differential equations on irregular domains using deep learning.
method Introduces Multi-Scale Attention Transformer (\msat{}) for solving PDEs.
result Achieves state-of-the-art generalization on complex geometry problems with significant speedup.
Sharp stability threshold found for deep residual architectures.
problem Ensuring stable training and inference in deep residual networks.
method Sublinear-growth principle and optimal-control analysis.
result Stable training condition: input-magnitude exponent q ≤ 1.
FCOC framework improves financial volatility forecasting.
problem Tackles dual challenges of feature fidelity and model responsiveness in financial volatility forecasting.
method Synergizes fractal feature extraction and dynamic chaotic oscillation processing.
result Demonstrates profound and generalizable impact on S\&P 500 and DJI datasets.
Parallelizes autoregressive generation using VSSM.
problem Autoregressive models' inability to parallelize generation.
method Variational SSM (VSSM) with parallelizable sampling and decoding.
result Parallel generation possible with VSSM.
This study improves state estimation for nonlinear systems using conditional normalizing flows.
problem Performance degradation of traditional filtering algorithms in nonlinear systems with non-Gaussian uncertainty.
method Uses conditional normalizing flows with MLP, transformer, or state-space models for state and parameter estimation.
result Optimal-transport-inspired kinetic loss mitigates overparameterization in flows.
A new framework for efficient sequence maps using Bayesian filtering and covariance.
problem Designing efficient recurrent sequence maps from explicit memory assumptions.
method Design-model framework, exact Bayesian filtering, query-dependent readout, linear-Gaussian instantiation.
result Improved robustness and retrieval performance across various benchmarks.
CogScale benchmarks AI architectures for sequential processing.
problem Evaluating AI architectures' ability to process sequential information efficiently.
method 14 scalable synthetic tasks designed to isolate cognitive and memory abilities at different scales.
result Attention mechanisms and modern state-space models consistently maintain high performance as task difficulty scales.
Introduces alternators for modeling sequences, outperforming baselines.
problem Modeling complex sequential data with stability and efficiency.
method Two neural networks (OTN and FTN) alternate between outputting samples in observation and feature spaces, learned via cross-entropy criterion.
result Alternators outperform strong baselines in various domains (Lorenz equations, Neuroscience, Climate Science).
New model preserves symmetry in multivariate time series, improving performance.
problem Implicit ordering in MTS models violates inherent exchangeability.
method Permutation-equivariant 2D state space model with canonical architecture.
result Eliminates sequential dependency chains and simplifies stability analysis.
pLSTM tackles long-range language modeling and computer vision tasks with parallelizable linear source transition mark networks.
problem Challenges of existing recurrent architectures in handling sequences and multi-dimensional data.
method Introduces pLSTM, a parallelizable linear source transition mark network for linear graphs and DAGs, addressing vanishing/exploding activation/gradient issues.
result pLSTM outperforms Transformers in long-range tasks like arrow-pointing extrapolation and image size extrapolation.
Proposes a new neural network architecture inspired by biology to improve learning and information flow.
problem Improving artificial neural networks to match biological neuron properties like multidirectional propagation and probabilistic modeling.
method Extends KAN approach with joint distribution neurons that can propagate values and distributions, including variance and higher-order moments.
result Proposed architecture can predict and propagate distributions, including expected values and variances.
V-HMN integrates memory mechanisms for improved image recognition.
problem Limited interpretability and high data requirements of existing vision backbones.
method Brain-inspired hierarchical memory modules with iterative refinement.
result V-HMN achieves strong performance on image classification benchmarks.
ByteGen models LOB dynamics without tokenization, achieving realistic market metrics.
problem Modeling high-frequency LOB dynamics in finance.
method Autoregressive next-byte prediction on packed binary data, using H-Net architecture.
result Successfully reproduces stylized facts of financial markets.
The α-Alternator adapts to varying noise levels in sequences, improving robustness and performance.
problem Current models assume uniform noise levels, limiting performance on noisy temporal data.
method Introduces α-Alternator using Vendi Score to dynamically adjust noise sensitivity. result Outperforms Alternators and state-of-the-art models in trajectory prediction, imputation, and forecasting.