SANAS adapts neural network architecture in real-time for efficient keyword spotting.
problem Efficiently identifying keywords in real-time audio streams with minimal computational cost.
method Stochastic Adaptive Neural Architecture Search (SANAS) that dynamically adjusts network size based on task difficulty.
result SANAS achieves high keyword recognition accuracy with significantly reduced computational cost.
MAML adapts faster with deeper architectures, especially in shallow tasks.
problem Understanding and improving MAML's fast adaptation.
method Empirical and theoretical studies on MAML's properties and optimization.
result MAML adapts better with deep architectures even for shallow tasks.
New RL algorithm learns state aggregation architecture adaptively.
problem Adapting reinforcement learning value function architectures.
method Adapts state aggregation architecture using state visit frequency feedback.
result Improves RL performance on various test problems.
New unsupervised speaker adaptation method for speech synthesis.
problem Adapting speech synthesis to new speakers with minimal data.
method Concatenating audio and text inputs, proposing new training schemes.
result Improves adaptation to unseen speakers and multi-speaker modeling.
ANTs integrate neural networks and decision trees for better performance.
problem Combining neural networks and decision trees for improved performance.
method Adaptive Neural Trees (ANTs) integrate representation learning into decision tree structures, using backpropagation for adaptive growth.
result ANTs achieve competitive performance on classification and regression datasets, with benefits in lightweight inference, feature separation, and adaptable architecture.
Paper outlines a system for ML models to learn continuously from evolving data.
problem Managing ML models in environments where data evolves.
method Describes a reference architecture for self-maintaining systems.
result Proposes a reference architecture for continual AutoML.
Adapts reinforcement learning architectures using state visit frequency.
problem Determining an optimal approximation architecture for reinforcement learning.
method Adapts state aggregation approximation architecture based on state visit frequency.
result Guarantees VF estimate arbitrarily close to zero with large S. A new autoencoder architecture captures multiscale data.
problem Multiscale spatio-temporal data representation.
method Integrates multigrid methods, convolutional autoencoders, and transfer learning.
result Adaptive, hierarchical architecture captures different scaled features dynamically.
Develops a robust, fast, and widely-applicable neural architecture search method.
problem Inability of current NAS methods to be easily applied to new problems.
method Adaptive stochastic natural gradient method for simultaneous optimization of weights and architecture.
result Near state-of-the-art performances with low computational budgets.
Learn to automatically plug domain-specific modules into a common network.
problem Learning inflexibility and computational intensiveness in multi-domain learning.
method Neural Architecture Search (NAS) for data-driven adapter plugging and structure design.
result NAS-driven MDL model achieves comparable performance to existing approaches.
Paper proposes a neural network for learning crossmodal stimuli.
problem Improving crossmodal processing in dynamic environments.
method Deep neural architecture trained by expectation learning.
result Self-adaptable deep learning model for crossmodal stimuli.
BASE framework learns task-agnostic architectures for faster NAS.
problem High computational cost in Neural Architecture Search.
method Bayesian Meta Architecture Search (BASE) framework.
result Found models achieving 25.7% top-1 error and 8.1% top-5 error in less than an hour.
NASIB adapts NAS to varying computation resources efficiently.
problem Constrained computation resources in enterprise environments.
method Adapts exploration vs. exploitation trade-off and uses Superkernels.
result Searches over a larger space with similar accuracy in less time.
AdaEnsemble learns adaptive feature interactions for CTR prediction.
problem Learning feature interactions for CTR prediction in recommender systems and Ads ranking.
method AdaEnsemble is a Sparsely-Gated Mixture-of-Experts (SparseMoE) architecture that dynamically selects feature interaction depth.
result AdaEnsemble achieves better prediction accuracy and inference efficiency compared to state-of-the-art models.
Adaptive convolution improves GANs performance on image generation.
problem GANs struggle with generating images of objects with diverse appearances.
method Proposes adaptive convolution to learn upsampling based on local context.
result Adaptive convolution models improve GANs performance on CIFAR-10 and STL-10 datasets.
Top-performing deep architectures are trained on massive amounts of labeled data. In the absence of labeled data for a certain task, domain adaptation often provides an attractive option given that labeled data of similar nature but from a different domain (e.g. synthetic images) are available. Here, we propose a new a…
Adapts deep learning with kernel methods for efficient learning.
problem Combining kernel methods and deep learning for efficient learning.
method Nyström approximation of kernel functions in neural networks.
result Performance comparable to standard architectures on datasets like SVHN and CIFAR100.
TAAN model learns optimal network architecture for MTL tasks.
problem Improving generalization performance of MTL by finding flexible and accurate shared architecture.
method TAAN model with flexible activation functions and functional regularization.
result TAAN and regularization methods improve MTL performance.
Generative adversarial networks benefit from optimal input dimension and adaptive generator architecture.
problem Minimizing generalization error in GANs through optimal input dimension.
method Introducing generalized GANs (G-GANs) with group penalty and architecture penalty for adaptive dimensionality reduction and network architecture identification.
result G-GANs achieve superior performance with 40%+ improvements in maximum mean discrepancy or Frechet inception distance compared to off-the-shelf methods.
MetaNAS improves few-shot learning by optimizing neural architectures with meta-learning.
problem Few-shot learning challenges due to limited data and compute time.
method MetaNAS integrates NAS with gradient-based meta-learning to adapt neural architectures to new tasks efficiently.
result MetaNAS achieves state-of-the-art results on few-shot classification benchmarks.
TINs use neural networks to interpret technical indicators for trading.
problem Lack of interpretable neural architectures for technical indicators in trading.
method Introduced TINs, a neural architecture that reformulates technical indicators into trainable modules.
result Improved risk-adjusted performance compared to traditional indicator-based strategies.
SpeqNets improve graph neural networks by scaling and adapting to graph sparsity.
problem Graph neural networks struggle with permutation-equivariant functions and scalability to large graphs.
method Introducing sparsity-aware, permutation-equivariant graph networks with heuristics for graph isomorphism.
result Significantly improved predictive performance and reduced computation times compared to existing methods.
Study compares neural networks for stock selection using fundamental ratios.
problem Predicting stock performance using fundamental financial ratios.
method Comparative study of feed-forward neural network (FNN) and adaptive neural fuzzy inference system (ANFIS).
result Both FNN and ANFIS can separate winners and losers, but FNN performs better.
New study finds best language model architecture and pretraining objective for zero-shot tasks.
problem Evaluating which language model architectures and pretraining objectives best enable zero-shot generalization.
method Compared three model architectures and two pretraining objectives across 170 billion tokens, with and without finetuning.
result Causal decoder-only models trained on autoregressive language modeling exhibit strongest zero-shot generalization.
This paper improves Adam's performance in machine learning tasks.
problem Improving the generalization ability of adaptive gradient methods.
method Develops a control theoretic framework to propose AdamSSM, a new variant of Adam.
result AdamSSM improves generalization accuracy and convergence compared to recent adaptive gradient methods.
AMF-VI uses adaptive mixtures of flows for robust VI across diverse distributions.
problem Inconsistent behavior of single-flow models across different distributions.
method Sequential expert training of individual flows and adaptive global weight estimation via likelihood-driven updates.
result AMF-VI achieves lower negative log-likelihood and stable gains in transport metrics across various posterior families.
TIME network simplifies complex physical processes with interpretable models.
problem Challenges in learning coupled dynamic processes from multiple observations.
method Fully convolutional architecture capturing invariant domain structure.
result Robust and transparent in capturing process kernels and anomalies.
A probabilistic framework for online test-time adaptation
problem Adapting models to new data under distributional shift
method State-space modelling architecture
result Characterizing parameter learning, time evolution, prior tuning, and prediction
AANets balance stability and plasticity in CIL.
problem Stability-plasticity dilemma in class-incremental learning.
method Adaptive Aggregation Networks (AANets) with stable and plastic residual blocks.
result AANets improve performance on CIL benchmarks.
Study improves stock price prediction using adaptive Mixture of Experts framework.
problem Tackles diverse volatility regimes in stock price prediction.
method Combines RNN for high-volatility stocks and linear regression for stable stocks with a gating mechanism.
result Achieves up to 33% improvement in MSE for volatile assets and 28% for stable assets.
Adapts LRP for LSTM to explain sequential data.
problem Lack of explainable AI for LSTM models.
method Extends LRP to LSTM, introduces new propagation scheme.
result Delivers faithful explanations for LSTM predictions.
ARM improves multivariate time series forecasting by better capturing series-wise relationships.
problem Challenges in handling complex temporal-contextual relationships in multivariate time series forecasting.
method ARM is an enhanced multivariate LTSF architecture that employs Adaptive Univariate Effect Learning, Random Dropping, and Multi-kernel Local Smoothing.
result ARM outperforms vanilla Transformers on multiple benchmarks without significantly increasing computational costs.
Adaptively preconditions SGLD for faster convergence and better generalization.
problem Pathological curvature in deep network loss landscapes.
method Adaptive estimation of noise parameters to precondition isotropic gradient noise.
result Adaptively preconditioned SGLD achieves faster convergence and generalization equivalent of SGD.
We introduce SADs to reveal how network architecture shapes score-based generative models.
problem Understanding and predicting the inductive biases of score-based generative models.
method Introducing Score Anisotropy Directions (SADs) to analyze network architecture.
result SADs reliably capture model behavior and correlate with performance.
Adaptive learning method improves classification in time series data analysis.
problem Improving classification capability in time series data analysis.
method Embedding adaptive learning method into recurrent temporal DBN.
result Higher classification capability than conventional methods.
Paper proposes a neural network for generating better questions from text.
problem Automatic generation of relevant questions from sentences and paragraphs.
method Adaptive copying recurrent neural network model with a copying mechanism added to a bidirectional LSTM architecture.
result The model outperforms state-of-the-art methods in question generation metrics.
PDNAS optimizes GNN architectures for diverse datasets.
problem Inadequate adaptability and combinatorial search space in GNNs.
method Dual architecture search (micro- and macro-architectures) with gradient-based optimization.
result PDNAS finds deeper GNNs with better performance on diverse datasets.
Adaptive networks improve model robustness through conditional normalization.
problem Limited robustness of adversarial-trained networks due to network capacity and training samples.
method Proposes a conditional normalization module to adapt networks during adversarial training.
result Adaptive networks outperform both clean validation accuracy and robustness compared to non-adaptive counterparts.
NACs learn modular neural architectures without domain knowledge.
problem Jointly learn module configuration and execution without domain knowledge.
method Jointly trains two systems: module configuration and execution.
result Improves low-shot adaptation and OOD robustness.
The paper optimizes dynamic scheduling for ring architectures in deep learning training.
problem Optimizing deep learning training times with ring architectures.
method Formulated a non-convex, non-linear, NP-hard integer programming problem and developed a doubling heuristic.
result Dynamic scheduling can significantly reduce job completion times in ring architectures.
New method trains deep networks robustly without adaptive methods.
problem Training deep networks with robustness and efficiency.
method Scale invariant architecture + SGD + weight decay + gradient clipping.
result SGD can achieve similar performance to adaptive methods like Adam.
New neural network learns adaptive behaviors inspired by neuromodulation.
problem Current AI lacks the ability to adapt to changing environments.
method Inspired by cellular neuromodulation, a new deep neural network architecture is designed.
result Neuromodulation-based networks improve adaptation in meta-reinforcement learning tasks.
XceptionTime improves hand gesture recognition accuracy using novel deep learning.
problem Improving hand gesture recognition from sparse sEMG signals.
method Depthwise separable convolutions, adaptive pooling, non-linear normalization.
result Significantly improved accuracy (5.71% improvement) in hand gesture recognition.
HS-MoE selects sparse experts using adaptive priors and data-adaptive gating.
problem Sparse expert selection in mixture-of-experts architectures.
method Combines horseshoe prior with input-dependent gating for data-adaptive sparsity.
result Data-adaptive sparsity in expert usage.
Improved software flaw detection using NAS on multimodal DL models.
problem Software flaw detection in multimodal deep learning models.
method Adapted NAS framework for multimodal learning, combined with multimodal deep learning models.
result Improved performance on the Juliet Test Suite.
Improved deep learning optimizers using adaptive stepsize.
problem Improving the performance of deep learning optimizers.
method Adapts stepsize directly with the loss function to make progress on loss.
result Enhanced optimizers outperform Adam and Momentum optimizers without increased computational cost.
Tabu Dropout improves performance of standard Dropout by generating more diverse neural network architectures.
problem Preventing co-adaptation of neurons in deep neural networks.
method Integrates a diversification strategy into dropout, marking units from the last forward propagation for re-selection in the current forward propagation.
result Improves performance of standard Dropout on MNIST and Fashion-MNIST datasets.
Paper introduces adaptive parameterization to improve neural network efficiency.
problem Neural networks' limited flexibility due to fixed activation functions.
method Adaptive parameterization of feed-forward layers that learn to adapt based on input.
result Adaptive LSTM achieves state-of-the-art performance with fewer parameters and faster convergence.