Survey of methods to train deep architectures without E2EBP.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Noise-driven neural networks emerge modular structures, improving robustness and generalization.
Training a Neural Network (NN) with lots of parameters or intricate architectures creates undesired phenomena that complicate the optimization process. To address this issue we propose a first modular approach to NN design, wherein the NN is decomposed into a control module and several functional modules, implementing …
NACs learn modular neural architectures without domain knowledge.
Neural networks learn modular arithmetic but not all, extending known solutions to generalize.
Modular RL modules solve complex 3D Sokoban tasks.
The study examines if ReLU activation function is optimal for modularity in neural networks.
RNNs solve modular addition tasks using low rank and sparse Fourier structures.
FinRL-X unifies trading components for AI and rule-based strategies.
New theory maps neural network weights to optimize faster and scale.
Neural architecture search methods are able to find high performance deep learning architectures with minimal effort from an expert. However, current systems focus on specific use-cases (e.g. convolutional image classifiers and recurrent language models), making them unsuitable for general use-cases that an expert migh…
BNNs enhance reservoir computing by acting as generalization filters.
Attention mechanism combines bottom-up and top-down signals in neural networks.
Neuro-inspired recurrent neural network algorithms, such as echo state networks, are computationally lightweight and thereby map well onto untethered devices. The baseline echo state network algorithms are shown to be efficient in solving small-scale spatio-temporal problems. However, they underperform for complex task…
A framework for modular training of robust generative models.
This paper introduces a new system for discovering patterns in morphogenetic systems using modular architecture and unsupervised learning.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
Improved recurrent neural networks learn long-term dependencies through multi-scale memory.
Contemporary sensorimotor learning approaches typically start with an existing complex agent (e.g., a robotic arm), which they learn to control. In contrast, this paper investigates a modular co-evolution strategy: a collection of primitive agents learns to dynamically self-assemble into composite bodies while also lea…
ModSSC unifies semi-supervised classification for various data types.
We explore efficient neural architecture search methods and show that a simple yet powerful evolutionary algorithm can discover new architectures with excellent performance. Our approach combines a novel hierarchical genetic representation scheme that imitates the modularized design pattern commonly adopted by human ex…
A core aspect of human intelligence is the ability to learn new tasks quickly and switch between them flexibly. Here, we describe a modular continual reinforcement learning paradigm inspired by these abilities. We first introduce a visual interaction environment that allows many types of tasks to be unified in a single…
We propose an actor-critic, model-free, and online Reinforcement Learning (RL) framework for continuous-state continuous-action Markov Decision Processes (MDPs) when the reward is highly sparse but encompasses a high-level temporal structure. We represent this temporal structure by a finite-state machine and construct …
Hydra boosts efficiency for long-context reasoning in resource-constrained settings.
NOMU improves neural network uncertainty estimation.
A new method recovers rewards from behavior policies using classification and regression.
Recursive Feature Machines show grokking in modular arithmetic without neural networks.
MOCA uses modular attention to estimate causal effects from complex data.
We focus on two supervised visual reasoning tasks whose labels encode a semantic relational rule between two or more objects in an image: the MNIST Parity task and the colorized Pentomino task. The objects in the images undergo random translation, scaling, rotation and coloring transformations. Thus these tasks involve…
There has been a rapid progress in the task of Visual Question Answering with improved model architectures. Unfortunately, these models are usually computationally intensive due to their sheer size which poses a serious challenge for deployment. We aim to tackle this issue for the specific task of Visual Question Answe…
Proposes efficient deep causal generative models for high-dimensional causal inference.
We present LumièreNet, a simple, modular, and completely deep-learning based architecture that synthesizes, high quality, full-pose headshot lecture videos from instructor's new audio narration of any length. Unlike prior works, LumièreNet is entirely composed of trainable neural network modules to learn mapping functi…
Deep reinforcement learning for high dimensional, hierarchical control tasks usually requires the use of complex neural networks as functional approximators, which can lead to inefficiency, instability and even divergence in the training process. Here, we introduce stacked deep Q learning (SDQL), a flexible modularized…
Understanding and modeling human driver behavior is crucial for advanced vehicle development. However, unique driving styles, inconsistent behavior, and complex decision processes render it a challenging task, and existing approaches often lack variability or robustness. To approach this problem, we propose Probabilist…
An artificial agent for financial risk and returns' prediction is built with a modular cognitive system comprised of interconnected recurrent neural networks, such that the agent learns to predict the financial returns, and learns to predict the squared deviation around these predicted returns. These two expectations a…
As deep learning applications continue to become more diverse, an interesting question arises: Can general problem solving arise from jointly learning several such diverse tasks? To approach this question, deep multi-task learning is extended in this paper to the setting where there is no obvious overlap between task a…
Convolutional architectures have recently been shown to be competitive on many sequence modelling tasks when compared to the de-facto standard of recurrent neural networks (RNNs), while providing computational and modeling advantages due to inherent parallelism. However, currently there remains a performance gap to mor…
CoE modularizes LLMs for scalable, cost-effective AI systems.
We find and propose an explanation for a large variety of modularity-related symmetries in problems of 3-manifold topology and physics of 3d theories where such structures a priori are not manifest. These modular structures include: mock modular forms, Weil representations, quantum mo…
Researchers found the global topology of the Eisenstein-Picard modular surface.
Study modular surfaces in Lorentz-Minkowski 3-space, classifying and analyzing their curvature and applications.
Modular neural networks generalize better with less data.
Study on NNs for forecasting time series with novel control variable combinations.
Study modular forms over Γ^0(2) and anomaly cancellation formulas.
Our aim is to introduce and advocate non- (non-symmetric) modular operads. While ordinary modular operads were inspired by the structure of the moduli space of stable complex curves, non- modular operads model surfaces with open strings outputs. An immediate application of our theory is a short proof that the mod…
The common pipeline in autonomous driving systems is highly modular and includes a perception component which extracts lists of surrounding objects and passes these lists to a high-level decision component. In this case, leveraging the benefits of deep reinforcement learning for high-level decision making requires spec…
Fuchsian groups with a modular embedding have the richest arithmetic properties among non-arithmetic Fuchsian groups. But they are very rare, all known examples being related either to triangle groups or to Teichmueller curves. In Part I of this paper we study the arithmetic properties of the modular embedding and deve…
Improved algorithm for modular links provides upper volume bounds.