Proposes faster neural network learning by using subsets of training data.
problem Manual design and validation of network topologies is time-consuming.
method Exploits subsets of training data at each incremental training step and performs online hyperparameter selection.
result Significantly reduces overall training time while maintaining performance.
Proposes PSGAN for generating high-res anime images with structural consistency.
problem Lack of high-quality, structurally consistent full-body high-resolution anime images.
method Progressive Structure-conditional Generative Adversarial Networks (PSGAN) with progressive training.
result Demonstrates effectiveness through comparisons and diverse anime character generation.
The paper investigates how neural network weights evolve to monitor training progress.
problem Monitoring the training progress of neural networks in a cost-effective manner.
method Investigates the evolution of neural network weights in weight space.
result DNN models evolve on unique, smooth trajectories in weight space that can be used to track training progress.
PCRs compress data for deep learning, reducing training time.
problem Efficiently training deep learning models over large datasets.
method Combining progressive compression with an efficient storage layout.
result PCRs can tolerate up to 50% compression without significantly affecting training accuracy.
This paper explores the trade-off between spatial and adversarial robustness in neural networks.
problem Understanding the trade-off between spatial and adversarial robustness in neural networks.
method Quantitative analysis and empirical testing with curriculum learning.
result Spatial robustness and adversarial robustness are quantitatively related and can be improved simultaneously.
PaGoDA reduces diffusion model training costs by 64x.
problem Diffusion models are computationally expensive during training.
method Three-stage pipeline: downsampled training, distillation, progressive super-resolution.
result PaGoDA achieves state-of-the-art performance with reduced training costs.
RLVR training dynamics reveal an implicit curriculum that shapes learning progression.
problem Understanding how RLVR overcomes the long-horizon barrier.
method Developed a theory of training dynamics for RLVR on transformers, using Fourier analysis on finite groups.
result Mixed-difficulty training naturally follows an implicit curriculum, shaping the learning progression from easy to hard.
Paper introduces a new curriculum generation method for reinforcement learning.
problem Improving reinforcement learning performance and speed through curriculum learning.
method The paper proposes a novel curriculum generation paradigm based on progression and mapping functions.
result Empirical results show the new approach outperforms state-of-the-art algorithms.
Curriculum learning speeds up agent learning in Minecraft, a complex visual domain.
problem Training agents to learn multiple tasks in a complex, visual domain.
method Learning-progress based curriculum and dynamic exploration bonuses.
result Curriculum learning improves agent performance in a complex reinforcement learning problem.
Drop-Muon updates only some layers, speeding up training.
problem Conventional deep learning optimizers update all layers at once, which can be inefficient.
method Drop-Muon updates only a subset of layers per step, with randomized schedules.
result Drop-Muon achieves up to 1.4x faster training time with similar accuracy.
Extracting a curriculum from a teacher network improves distillation efficiency.
problem Efficiently training a small network using a large teacher network's output.
method Random projection of teacher network's hidden representations to progressively train the student network.
result Extracted curriculum significantly outperforms one-shot distillation and achieves similar performance to progressive distillation.
AI progress measured by reduced compute needed to reach past performance.
problem Quantifying algorithmic progress in AI.
method Analysis of floating-point operations required to train neural networks.
result Algorithmic efficiency doubled every 16 months over 7 years.
Gradient span algorithms show consistent progress in high dimensions.
problem Understanding consistent training progress in large machine learning models.
method Proving deterministic behavior of gradient span algorithms on Gaussian random functions.
result Gradient span algorithms have asymptotically deterministic behavior in high dimensions.
We describe a new training methodology for generative adversarial networks. The key idea is to grow both the generator and discriminator progressively: starting from a low resolution, we add new layers that model increasingly fine details as training progresses. This both speeds the training up and greatly stabilizes i…
New technique reduces complexity of deep networks.
problem High computational complexity and energy consumption of deep neural networks.
method Progressive Gradient Pruning (PGP) for iterative filter pruning during training.
result PGP achieves better trade-off between accuracy and network complexity.
Flat learning curves reveal no progress in ENAS controller.
problem Improving learning speed in neural architecture search.
method Evaluated learning progress of ENAS controller through architecture re-training.
result No observable progress in controller's generated architectures.
We present new intuitions and theoretical assessments of the emergence of disentangled representation in variational autoencoders. Taking a rate-distortion theory perspective, we show the circumstances under which representations aligned with the underlying generative factors of variation of data emerge when optimising…
System automates discovery and classification of training videos for career progression.
problem Difficulties in planning and navigating career paths due to changing job requirements and emerging sectors.
method Extracted educational videos, built a machine learning classifier, and optimized probability thresholds.
result Significant improvements in model performance by incorporating video attributes.
Improved generalization in abstract reasoning tasks using disentangled latent representations.
problem Improving generalization in unsupervised representation learning for abstract reasoning.
method Used disentangled VAEs to learn latent representations from relational reasoning problems.
result Disentangled latent representations outperform supervised learning in generalization.
We introduce Mix&Match (M&M) - a training framework designed to facilitate rapid and effective learning in RL agents, especially those that would be too slow or too challenging to train otherwise. The key innovation is a procedure that allows us to automatically form a curriculum over agents. Through such a curriculum …
Most approaches to machine learning from electronic health data can only predict a single endpoint. Here, we present an alternative that uses unsupervised deep learning to simulate detailed patient trajectories. We use data comprising 18-month trajectories of 44 clinical variables from 1908 patients with Mild Cognitive…
AL2 progressively penalizes network activations to prevent overfitting.
problem Avoiding overfitting in neural networks with limited training data.
method Progressive Activation Loss (PAL) method to regularize network representations.
result AL2 outperforms traditional regularization methods on benchmark datasets.
Improved NODEs for long-term time series forecasting.
problem Dealing with complex, multi-frequency data.
method Progressive learning paradigm with curriculum learning.
result Performance improved by over 64%.
CRBM generates digital twins for MS patients, aiding in disease progression analysis.
problem Characterizing and analyzing disease progression in MS patients.
method Unsupervised machine learning with Conditional Restricted Boltzmann Machines (CRBMs).
result Generated digital twins are statistically indistinguishable from actual subjects.
Efficiently selects best machine learning config from large dataset.
problem Finding the best machine learning configuration from a large dataset.
method CI-based progressive sampling and pruning strategy.
result Achieves more than two orders of magnitude speedup while maintaining similar accuracy.
New approach ties loss curvature to model performance in deep learning.
problem Understanding the relationship between loss curvature and model performance in deep learning.
method Empirical analysis of loss Hessians and theoretical results on input-output Jacobians.
result Novel generalization bound in terms of empirical Jacobian.
We use GG distributions to optimize LLMs, reducing size and training time.
problem Lack of understanding in the statistical structure of LLMs.
method GG-based initialization, ACT, GCT.
result Smaller, faster models with minimal communication overhead.
Despite the advancement of supervised image recognition algorithms, their dependence on the availability of labeled data and the rapid expansion of image categories raise the significant challenge of zero-shot learning. Zero-shot learning (ZSL) aims to transfer knowledge from labeled classes into unlabeled classes to r…
Deep nearest neighbors outperform self-supervised methods in anomaly detection.
problem Anomaly detection using self-supervised deep methods.
method Simple nearest-neighbor approach on Imagenet pretrained features.
result Nearest-neighbor method outperforms self-supervised methods in accuracy, few shot generalization, training time, and noise robustness.
We introduce a conceptually simple and scalable framework for continual learning domains where tasks are learned sequentially. Our method is constant in the number of parameters and is designed to preserve performance on previously encountered tasks while accelerating learning progress on subsequent problems. This is a…
Proposes a curriculum-based scheme to smooth CNN feature embeddings.
problem Distortion artifacts in early training stages of CNNs.
method Smoothes feature embedding using Gaussian kernels to control high-frequency information.
result Significant performance improvements on various vision tasks.
High volume of data, perceived as either challenge or opportunity. Deep learning architecture demands high volume of data to effectively back propagate and train the weights without bias. At the same time, large volume of data demands higher capacity of the machine where it could be executed seamlessly. Budding data sc…
Novel ECG classification method using deep time-frequency representation and progressive decision fusion.
problem Challenges in classifying abnormal ECG rhythms due to broad taxonomy, noises, and lack of annotated data.
method Transform ECG signal into time-frequency domain, train scale-specific deep CNNs, and fuse decisions progressively.
result Effective and efficient ECG classification method validated on synthetic and real-world datasets.
A new principle for extrapolating regression outside training data.
problem Regression extrapolation when predictions are outside training data range.
method Data-adaptive marginal transformation and simple relationship assumption.
result Progression method offers guarantees on approximation error beyond training data range.
New method uses game theory to rate generative models.
problem Evaluating generative models' performance.
method Tournaments between generators and discriminators to summarize outcomes.
result Tournament win rate and skill rating provide effective model evaluations.
Proposes a progressive label correction method for feature-dependent label noise.
problem Real-world large-scale datasets often suffer from heterogeneous, feature-dependent label noise.
method A progressive label correction algorithm that iteratively refines the model.
result A classifier trained with this strategy converges to be consistent with the Bayes classifier for various noise patterns.
In this paper we introduce Curriculum GANs, a curriculum learning strategy for training Generative Adversarial Networks that increases the strength of the discriminator over the course of training, thereby making the learning task progressively more difficult for the generator. We demonstrate that this strategy is key …
A new measure evaluates GANs without labels, capturing failures and tracking progress.
problem Evaluating GANs is challenging due to non-intuitive loss behavior and lack of suitable metrics.
method Proposes a duality gap measure from game theory for low computational cost.
result Effective at ranking GAN models and monitoring training progress.
Fastens diffusion model sampling by distillation.
problem Slow sampling time of diffusion models.
method New parameterizations and progressive distillation.
result Models can be distilled to take half as many sampling steps.
GANs simulate realistic galaxy images.
problem Simulate complex astronomical images efficiently.
method Progressive GANs with Wasserstein cost function.
result Generates naturalistic galaxy images.
Paper proposes a new PLL framework with a progressive identification algorithm.
problem Weakly supervised learning with partial labels.
method Flexible model and optimization algorithm for PLL, progressive identification algorithm.
result Established an estimation error bound and set new state of the art.
Bayesian meta-learning predicts Alzheimer's disease progression.
problem Predicting individual Alzheimer's disease progression from limited data.
method Bayesian meta-learning approach that dynamically predicts disease score distributions.
result Bayesian meta-learner outperforms single-task models and deterministic meta-learners, especially for long-term predictions.
PDA improves deep neural networks' robustness against adversarial and common corruptions.
problem Deep neural networks' lack of robustness against common corruptions and adversarial attacks.
method Progressive Data Augmentation (PDA) that injects diverse adversarial noises during training.
result PDA-trained networks are more robust against both adversarial and common corruptions.
EBM uses high-dimensional imaging biomarkers to improve dementia progression estimation.
problem Current EBMs only use scalar biomarkers, limiting accuracy from cross-sectional data.
method Proposes nDEBM, a novel method using semi-supervised SVM on voxel-wise imaging biomarkers.
result nDEBM outperforms state-of-the-art EBM methods using regional volume biomarkers.
In this paper, we investigate a new form of automated curriculum learning based on adaptive selection of accuracy requirements, called accuracy-based curriculum learning. Using a reinforcement learning agent based on the Deep Deterministic Policy Gradient algorithm and addressing the Reacher environment, we first show …
PTSD improves neural samplers by combining diffusion models and PT, enhancing efficiency.
problem Efficiency and correlation issues in neural samplers compared to PT.
method Sequential training of diffusion models across temperatures, combining high-temperature models for approximate lower-temperature samples.
result Significantly improved target evaluation efficiency, outperforming diffusion-based samplers.
Study examines APOE's impact on AD progression using a novel DEBM approach.
problem Understanding APOE's role in AD progression and developing targeted clinical trials.
method Developed a discriminative event-based model (DEBM) and proposed a stratified approach to improve model accuracy.
result Identified APOE carriers' impact on AD progression timeline, aiding clinical trial selection.
PS-KD distills a model's own knowledge to soften hard targets during training.
problem Improving generalization of deep neural networks by softening hard targets.
method Progressive self-knowledge distillation (PS-KD) that progressively distills a model's own knowledge to soften hard targets.
result PS-KD improves accuracy and provides high quality of confidence estimates in terms of calibration and ordinal ranking.