Enhances DenseNets with multi-scale convolutions and stochastic feature reuse.
problem Overfitting in DenseNets with dense feature reuse.
method Multi-scale Convolution Aggregation module and Stochastic Feature Reuse.
result Significant improvement in model accuracy with fewer parameters.
Improved scaling laws in linear regression using data reuse.
problem Sustainability of neural scaling laws when running out of new data.
method Data reuse in multi-pass stochastic gradient descent (multi-pass SGD) for M-dimensional linear models trained on N data with sketched features. result Multi-pass SGD achieves a test error of Θ(M1−b+L(1−b)/a) with L>N, improving scaling laws in data-constrained regimes. MAML's success is due to feature reuse, not rapid learning.
problem Understanding the effectiveness of MAML in few-shot learning.
method Ablation studies and analysis of latent representations.
result Feature reuse is the dominant factor in MAML's success.
Method predicts rarity of image features to support research integrity investigations.
problem Difficulty in determining if image reuse is by chance or intentional.
method Statistical estimation of ORB features' chance occurrence across PubMed Open Access Subset dataset.
result The method produces decreasingly smaller p-values for more complex imagery, supporting null hypothesis.
Transfer learning offers little benefit in medical imaging tasks.
problem Understanding transfer learning's impact on medical imaging tasks.
method Evaluation of transfer learning on two large-scale medical imaging tasks.
result Simple, lightweight models perform comparably to ImageNet architectures.
The paper analyzes MAML's representation using RSA, revealing that feature reuse is not the primary reason for its success.
problem Understanding why model-agnostic meta-learning (MAML) works well in few-shot learning tasks.
method Representation similarity analysis (RSA) applied to MAML's few-shot learning instantiation.
result Feature reuse is not the primary reason for MAML's success; instead, it is the learning task itself that increases representation similarity.
Transfer learning improves image classifier performance in data-starved regimes.
problem Deploying image classifiers in domains with limited labeled data.
method Transfer learning with deep neural networks, focusing on feature reuse and overparameterization.
result Transfer learning enhances CNN performance in data-starved regimes.
The paper investigates what enables successful transfer learning and separates feature reuse from data statistics.
problem Understanding what enables successful transfer learning and identifying the responsible parts of the network.
method Analyzes transfer learning on block-shuffled images to distinguish feature reuse from data statistics.
result Some benefit of transfer learning comes from learning low-level statistics of data, not just feature reuse.
A new method, Multi-Label Deep Forest, tackles multi-label learning problems.
problem Leveraging label correlations in multi-label learning models.
method Designs a deep forest framework with two mechanisms: measure-aware feature reuse and measure-aware layer growth.
result Outperforms compared methods on six measures across benchmark datasets.
Study on feature learning dynamics in infinite-depth neural networks, focusing on ResNets.
problem Understanding how features evolve during training in deep neural networks, especially in the large-depth limit.
method Conditional Gaussian representations and SDE system with decoupled backward weights.
result Depth-induced suppression of forward-backward coupling in infinite-depth networks, leading to a decoupled forward-backward SDE system.
Paper explores reusing and adapting training data for entity resolution.
problem Entity resolution with limited training data.
method Distributed representation and five algorithms for reuse scenarios.
result Significant performance improvements with reused training data.
Bayesian Experience Reuse improves learning from multiple experts.
problem Learning from multiple experts with conflicting goals.
method Bayesian neural networks with shared features to model uncertainty and derive a probability distribution over expert models.
result BERS method effectively samples demonstrations from the derived distribution to reuse them in new tasks.
The paper develops a theory linking pretraining and fine-tuning in neural networks.
problem Understanding how initialization choices impact feature learning and generalization in neural networks.
method Analytical theory of diagonal linear networks, deriving generalization error as a function of initialization parameters and task statistics.
result Different initialization choices place networks into four fine-tuning regimes with varying abilities to support feature learning and generalization.
VRER selectively reuses samples to improve policy optimization in complex systems.
problem Lack of effective reuse of historical samples in reinforcement learning.
method Variance reduction based experience replay (VRER) framework.
result VRER accelerates policy optimization and enhances performance.
Paper learns a versatile model from diverse networks without annotations.
problem Combining knowledge from different specialized networks without access to their training data.
method Transforms features of diverse networks into a common space and forces a student model to mimic them.
result The student model outperforms individual teacher models on various benchmarks.
Paper proposes PRR network for better experience reuse in reinforcement learning.
problem Efficient experience reuse in reinforcement learning across multiple granularities.
method Proposes PRR network trained on multi-level architecture to extract and store experience.
result PRR network leads to better experience reuse and improved performance.
Representations are fundamental to artificial intelligence. The performance of a learning system depends on the type of representation used for representing the data. Typically, these representations are hand-engineered using domain knowledge. More recently, the trend is to learn these representations through stochasti…
New method uses rank-conditioned Horvitz-Thompson estimation for unbiased sample reuse in Plackett-Luce best-of-K objective.
problem Estimating the expected maximum reward in Plackett-Luce draws without replacement.
method Rank-conditioned Horvitz-Thompson estimation with joint-score REINFORCE for unbiased sample reuse.
result Unbiased estimation of the Plackett-Luce best-of-K objective with finite second moment guarantees.
New insights into neural network feature learning through multi-step gradient descent.
problem Understanding feature learning in two-layer neural networks with limited width.
method Characterization of feature learning through two steps of gradient descent with specific step sizes.
result The second step of gradient descent reveals multiple learned directions, not limited to a single direction as in the first step.
DenseNets improve accuracy and efficiency in convolutional networks.
problem Improving accuracy and efficiency in deep convolutional networks.
method Introducing Dense Convolutional Networks (DenseNet) with direct connections between all layers.
result DenseNets achieve significant improvements over state-of-the-art networks on object recognition benchmarks.
Contrary to the situation with stochastic gradient descent, we argue that when using stochastic methods with variance reduction, such as SDCA, SAG or SVRG, as well as their variants, it could be beneficial to reuse previously used samples instead of fresh samples, even when fresh samples are available. We demonstrate t…
VRER selectively reuses past observations to reduce variance in policy optimization.
problem Lack of effective experience replay for accelerating policy optimization in complex systems.
method Variance Reduction Experience Replay (VRER) framework that selectively reuses informative samples.
result VRER reduces gradient variance and improves policy learning over state-of-the-art algorithms.
Study on reliability of latent reuse in diffusion models under distribution shift.
problem When can latent spaces from a source dataset be reused for a target dataset with different distributions?
method Considered a source-target setting with approximately low-dimensional datasets near different subspaces. Analyzed the target-domain score error due to principal-angle misalignment and target ambient noise.
result Latent reuse is reliable only if the source and target subspaces are close and the target ambient noise is not too amplified.
Pipeline-aware hyperparameter tuning speeds up machine learning pipelines by reusing intermediate computations.
problem High computational burden in hyperparameter tuning of multi-stage pipelines.
method Proposes a hybrid hyperparameter tuning method and a caching problem formulated as an ILP to maximize reuse.
result Pipeline-aware approach offers over an order-of-magnitude speedup over independent evaluations.
Bayesian optimization improves performance with common random numbers.
problem Optimizing expensive stochastic functions with common random numbers.
method Proposes a novel Gaussian process model and Knowledge Gradient for Common Random Numbers.
result Significant performance improvements with moderate computational cost.
DQN struggles to generalize, but regularization helps.
problem DQN's poor generalization in similar environments.
method Evaluation protocol on Atari games, dropout, and ℓ2 regularization. result Regularization improves DQN's generalization and sample efficiency.
Improved ResNets and DenseNets models for better feature reuse.
problem Diminishing feature reuse in ResNets and DenseNets.
method ResNEsts and DenseNEsts are block-based DNN models with improved representation guarantees.
result Wide ResNEsts with bottleneck blocks can guarantee desirable training properties.
New method speeds up diffusion models without sacrificing quality.
problem Slow inference in diffusion models.
method Adams-Bashforth method for caching and acceleration.
result Achieved nearly 3x speedup with maintained quality.
This paper uses antithetic sampling to reduce variance in stochastic gradient descent.
problem High variance in stochastic gradient descent slows down convergence.
method Antithetic sampling to make gradients negatively correlated.
result The proposed method accelerates convergence in machine learning applications.
We propose several ways of reusing subword embeddings and other weights in subword-aware neural language models. The proposed techniques do not benefit a competitive character-aware model, but some of them improve the performance of syllable- and morpheme-aware models while showing significant reductions in model sizes…
New RL algorithms improve control tasks with data reuse.
problem Real-world control requires performance guarantees and data efficiency.
method Generalized Policy Improvement combining on-policy guarantees and sample reuse.
result Extensive experimental analysis shows benefits of new algorithms.
ECPF improves classification accuracy and speed for evolving data streams.
problem Reusing classifiers trained on recurring concepts to maintain accuracy and speed.
method ECPF uses similarity of classifications to quickly identify the best classifier to reuse.
result ECPF significantly outperforms state-of-the-art frameworks on synthetic and real-world datasets.
Paper reduces communication in distributed machine learning.
problem Reduces burdensome communication in distributed machine learning.
method Introduces communication-censoring technique to reduce transmissions of variables.
result CSGD algorithm achieves same convergence rate as SGD but with significant communication reduction.
Paper proposes an online speech recognition model using Transformer.
problem Challenges in deploying Transformer-based E2E ASR for online speech recognition.
method Chunk self-attention encoder (chunk-SAE) and monotonic truncated attention (MTA) based self-attention decoder (SAD).
result Achieved 23.66% CER with 320 ms latency, significant improvement over offline models.
A key question in Reinforcement Learning is which representation an agent can learn to efficiently reuse knowledge between different tasks. Recently the Successor Representation was shown to have empirical benefits for transferring knowledge between tasks with shared transition dynamics. This paper presents Model Featu…
In this work a new way to calculate the multivariate joint entropy is presented. This measure is the basis for a fast information-theoretic based evaluation of gene relevance in a Microarray Gene Expression data context. Its low complexity is based on the reuse of previous computations to calculate current feature rele…
In many real-world applications, data are often collected in the form of stream, and thus the distribution usually changes in nature, which is referred as concept drift in literature. We propose a novel and effective approach to handle concept drift via model reuse, leveraging previous knowledge by reusing models. Each…
Improves reinforcement learning stability and efficiency.
problem Combining stability and efficiency in reinforcement learning.
method Combines on-policy stability with off-policy sample reuse.
result Demonstrates improved performance in both theory and practice.
SWNets optimize DL architectures for faster convergence.
problem Excessive training parameters in deep learning models.
method Transforms network topology to reach Small-World Network boundary.
result SWNets achieve faster convergence with fewer parameters.
LAZO reduces query complexity and variance in ZO methods.
problem High query complexity and variance in zeroth-order optimization.
method LAZO uses adaptive lazy queries to reduce variance and save queries.
result LAZO achieves lower regret and query complexity compared to existing methods.
Similar models predict similarly, reducing overfitting risk.
problem Excessive reuse of test data in machine learning.
method Proved model similarity mitigates overfitting and provided a generalization bound.
result Model similarity reduces the risk of overfitting, even when accuracy levels suggest otherwise.
We present a neurosymbolic framework for the lifelong learning of algorithmic tasks that mix perception and procedural reasoning. Reusing high-level concepts across domains and learning complex procedures are key challenges in lifelong learning. We show that a program synthesis approach that combines gradient descent w…
New bounds show multiclass problems reduce overfitting from test set reuse.
problem Overfitting from test set reuse in multiclass problems.
method Upper and lower bounds on bias of attacks, practical and inefficient attacks.
result Multiclass problems mitigate overfitting from test set reuse.
This paper proposes reusing CNN layers to speed up hyperparameters tuning.
problem Time-consuming hyperparameters tuning in CNNs.
method Reuse trained convolutional layers among different trainings.
result Reduces training time and increases accuracy of neural networks.
Paper proposes a method to reuse models without raw data.
problem Reuse existing models without accessing raw data.
method Two-phase framework: Upload and Deployment phases.
result Theoretical and experimental validation of approach effectiveness.
ImJoy simplifies deep learning for biomedical research.
problem Computational barriers limit deep learning adoption in biomedical research.
method Open-source browser-based platform for deep learning.
result Facilitates widespread reuse of deep learning solutions.
A method to reuse 98% of parameters for multi-task learning.
problem Improving efficiency in deep learning parameter usage.
method Learning model patches for each task, reusing pretrained network parameters.
result Significant improvement in transfer learning accuracy with fewer parameters.
Improves reinforcement learning by combining off-policy data and exploration.
problem Data inefficiency and local optima in policy gradient methods.
method Combines off-policy data reuse, exploration, and deterministic policies with stochastic optimization.
result Successfully learns solutions using fewer interactions than standard methods.