TorchGAN simplifies GAN training and evaluation with a flexible framework.
problem Training and evaluating GANs can be complex and time-consuming.
method Modular design, extensibility, built-in support for popular models and metrics.
result TorchGAN achieves near-zero overhead compared to vanilla PyTorch.
Solve-training trains neural nets to map physical solutions efficiently.
problem Representing complex physical solutions with neural networks.
method Variational training using loss functions from physical models.
result Effective neural network representation of solution maps without expensive labels.
Python framework for distributed Keras training on multiple GPUs/CPU.
problem Efficiently training neural networks on multiple GPUs/CPU.
method Built on Keras, uses MPI for coordination, suitable for supercomputing.
result Demonstrated performance on various system sizes.
Framework for few-shot relation classification with minimal training data.
problem Few-shot relation classification with limited training data.
method Meta-learning framework that combines instance and support knowledge.
result Framework outperforms state-of-the-art results and achieves competitive performance with large training data.
Unified framework for training neural networks converges for various types.
problem Mathematical tractability issues in training neural networks.
method Unified optimization framework for arbitrary loss, activation, and regularization functions.
result Framework generalizes and proves convergence of various training methods.
Unified framework for supervised classification with diverse training data.
problem Handling different types of training data for supervised classification.
method Generalized robust risk minimization (GRRM) with probabilistic transformations.
result GRRM can handle various training data types and new supervision schemes.
Low-bit training framework reduces energy consumption in CNNs.
problem Reducing energy consumption in convolutional neural networks.
method Low-bit training framework using MLS tensor format with dynamic quantization.
result Achieves superior trade-off between accuracy and bit-width.
A new multilevel framework speeds up ResNet training.
problem Training deep residual networks (ResNets) is time-consuming.
method Formulates ResNets as dynamical systems and uses time-dependent optimal control problems.
result Enhanced training of ResNets with multilevel auxiliary networks achieves significant speedup.
Unified framework for training SNNs using EP, faster convergence.
problem Training spiking neural networks with Expectation-Propagation.
method Message-passing framework for learning marginal distributions of SNN parameters.
result Faster convergence compared to gradient-based methods.
Unified framework for lifted training and inversion of neural networks.
problem Challenges in gradient-based training of deep neural networks.
method Unified framework encapsulating various lifted training strategies.
result Unified framework improves training landscape and stability.
Develops a robust training framework to detect backdoor attacks in DNNs.
problem Vulnerability of DNNs to backdoor attacks by poisoned training data.
method Collider framework selects prominent samples based on geometric structures and coreset selection objective.
result Significantly reduces backdoor success rate in various poisoned datasets.
New HPO frameworks needed for continual learning.
problem No standard HPO for continual learning.
method Comparative study of HPO frameworks.
result No HPO framework consistently outperforms others.
PaRoT simplifies robust training for deep neural networks.
problem Training deep neural networks to be robust to small input changes.
method Developed a practical framework on TensorFlow for robust training without code modifications.
result PaRoT's performance is comparable to existing methods and is easy to use on real-world models.
A framework for training deep networks in Apache Spark.
problem Expensive and time-consuming training of deep networks with large data and model parameters.
method Data and model-parallel, distributed training over Apache Spark clusters.
result Significant speedup and scalability for deep network training.
Paper proposes a new method for training nonconvex models.
problem Training nonconvex models like neural networks.
method Successive functional gradient optimization using mirror descent in a function space.
result The method leads to better performance than standard training techniques.
Paper develops a statistical framework for quantized training of deep neural networks.
problem Lack of theoretical understanding of gradient quantization in FQT.
method Presented a statistical framework for analyzing FQT algorithms, viewing quantized gradient as a stochastic estimator of QAT gradient.
result Developed two novel gradient quantizers with smaller variance than existing per-tensor quantizer.
Paper tackles model vulnerabilities by reconstructing training data.
problem Reconstructing training data from model parameters poses a security risk.
method Developed a mathematical framework and score matching method for both Bayesian and non-Bayesian models.
result First score matching framework for reconstructing data in Bayesian models.
This paper proposes a curriculum learning framework for NMT to reduce training time and improve performance.
problem Slow training and need for heuristics in NMT systems.
method A curriculum learning framework that decides training samples based on estimated difficulty and model competence.
result Up to 70% decrease in training time and up to 2.2 BLEU accuracy improvements.
Framework infers conservation laws from trained neural networks.
problem Building reduced models of complex systems from physical data.
method Derives conservation laws from symmetries of dynamics in trained DNNs using Noether's theorem.
result Consistent results with previous studies for metastable collective motion systems.
We propose a framework for training GANs on composed data, improving model modularity and interpretability.
problem Training GANs on complex, composed data.
method Composition/decomposition framework for adversarially training GANs on composed data.
result Improves modularity, extensibility, and interpretability of GANs.
ZipML framework trains models at low precision with provable guarantees and significant speedups.
problem Training machine learning models at low precision to achieve speedups and maintain accuracy.
method ZipML framework, double sampling, variance-optimal stochastic quantization, approximation to non-linear models.
result Training at low precision with ZipML framework achieves up to 6.5x speedups and maintains accuracy.
We propose a new framework for estimating generative models via an adversarial process, in which we simultaneously train two models: a generative model G that captures the data distribution, and a discriminative model D that estimates the probability that a sample came from the training data rather than G. The training…
New framework minimizes model complexity for improved few-shot learning.
problem Empirical benefits of pre-training scale with data size but lack theoretical explanation.
method Complexity Minimization framework for meta-representation learning.
result Theoretical analysis shows error rate improves with more meta-training data.
Pre-trains GNN for generic graph features, reducing data need.
problem Lack of labeled data and expressive features for GNN models.
method Synthetic graph tasks: link reconstruction, centrality ranking, cluster preservation.
result Pre-trained GNN achieves high performance with less labeled data.
Chicle tackles elastic machine learning training by avoiding micro-tasks.
problem Elasticity and load balancing in distributed machine learning training.
method Chicle is a new elastic distributed training framework that exploits machine learning algorithms to implement elasticity and load balancing without micro-tasks.
result Chicle achieves performance competitive with state-of-the-art rigid frameworks while enabling elastic execution and dynamic load balancing.
ENSURE framework trains deep image recon algorithms without clean data.
problem Lack of clean, fully sampled ground-truth data for deep learning image reconstruction.
method Introduces ENSURE framework, a generalization of SURE and GSURE to random sampling patterns.
result ENSURE loss function is an unbiased estimate for true mean-square error.
GP-TS optimizes TLM pre-training hyperparameters efficiently.
problem Resource inefficiency in TLM pre-training.
method Bayesian optimization with Thompson sampling and Gaussian process.
result GP-TS achieves lower MLM loss in fewer epochs.
New framework improves LLM performance by avoiding forgetting during sequential training stages.
problem Forgetting during sequential training stages of LLMs.
method Proposes a joint post-training framework with theoretical convergence guarantees.
result Empirically outperforms sequential post-training framework by up to 23%.
A new framework trains RBMs deterministically for unsupervised learning.
problem Training and evaluation of RBMs with weak interactions.
method TAP mean-field approximation for generalized latent-variable models.
result Effective deterministic training and interesting unsupervised learning features demonstrated.
This work proposes a new adversarial training method based on L2L framework.
problem Training robust neural networks against adversarial attacks.
method Generic learning-to-learn (L2L) framework to learn an optimizer and a robust classifier.
result L2L outperforms existing adversarial training methods in classification accuracy and computational efficiency.
POLAR framework interprets word embeddings using polar opposites.
problem Lack of interpretability in pre-trained word embeddings.
method Adopt semantic differentials and polar opposites to transform embeddings.
result Interpretable word embeddings maintain performance comparable to original embeddings.
We present a novel view that unifies two frameworks that aim to solve sequential prediction problems: learning to search (L2S) and recurrent neural networks (RNN). We point out equivalences between elements of the two frameworks. By complementing what is missing from one framework comparing to the other, we introduce a…
Neural ODEs provide a framework for studying the training dynamics of neural networks.
problem Training dynamics of neural networks
method Dynamical mean field theory
result Derive learning curves in the high-dimensional limit
Enhances robustness of AT frameworks to multiple perturbations without increasing training complexity.
problem Defending against the union of multiple perturbations in adversarial training.
method SNAP technique that augments a network with shaped noise to enhance robustness.
result 14%-to-20% improvement in adversarial accuracy for ResNet-18 on CIFAR-10.
Framework improves gradient estimation for faster training convergence.
problem Efficiently estimating noisy gradients in stochastic optimization.
method Dynamic adaptive importance sampling combining multiple distributions.
result Adaptively weighted multiple importance sampling yields superior gradient estimates.
A new framework explains why early pruning works well.
problem Understanding why early pruning of neural networks leads to good performance.
method Gradient flow framework to unify pruning measures.
result Magnitude-based pruning removes least contributing parameters, leading to faster convergence.
Framework trains safe agents avoiding deceptive behavior.
problem Training safe agents from unsafe incentives.
method Formal settings, causal influence analysis, maximizing non-mediated effects.
result Agents avoid manipulating delicate state for rewards.
Transfer learning improves sentiment classification using XR framework.
problem Lack of labeled data for deep learning.
method XR framework applied to transfer learning between related tasks, using expected label proportions.
result Improved performance on aspect-based sentiment classification.
A probabilistic framework for online test-time adaptation
problem Adapting models to new data under distributional shift
method State-space modelling architecture
result Characterizing parameter learning, time evolution, prior tuning, and prediction
Develops a framework for parallel and distributed neural network training.
problem Training neural networks in a distributed environment with sparse connectivity.
method Customizes a non-convex optimization framework over networks, including dynamic consensus and parallel optimization.
result Guarantees convergence to a stationary solution under mild assumptions.
Framework corrects noisy labels to improve DNN performance.
problem Performance degradation due to noisy labels in large-scale datasets.
method Joint optimization of DNN parameters and true labels estimation.
result Significantly outperforms state-of-the-art methods in experiments.
Unified view of generative models using GFlowNet framework.
problem Diverse deep generative models with varied training and inference methods.
method Integrates GFlowNet framework to unify training and inference.
result Unified training and inference algorithms for generative models.
GANs learn from incomplete data with missing data imputation.
problem Learning from incomplete data with GANs.
method Proposes a GAN framework with a complete data generator and a mask generator to model missing data.
result Demonstrates effective imputation of missing data using adversarial training.
Framework reuses pre-trained models for data-free transfer learning.
problem Challenges in retrieving source data for model training.
method Model Recycling Framework for parameter-efficient training.
result Makes multi-source data-free supervised transfer learning possible.
A Bayesian framework models adversarial uncertainty for robust machine learning.
problem Vulnerability of machine learning models to adversarial attacks.
method Formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating probabilistic assumptions.
result Explicitly modeling adversarial uncertainty leads to improved robustification strategies.
Paper develops a theory explaining contrastive pre-training for multimodal AI.
problem Limited theoretical understanding of contrastive pre-training for multi-modal AI.
method Introduces approximate sufficient statistics and Joint Generative Hierarchical Model.
result Near-minimizers of contrastive loss are approximately sufficient, enabling diverse downstream tasks.
Unsupervised pre-training improves model generalization, but lacks theoretical understanding.
problem Lack of theoretical understanding of unsupervised pre-training's impact on model generalization.
method Introduces a novel theoretical framework to analyze and enhance generalization.
result Enhances understanding of unsupervised pre-training and fine-tuning, proposing a new regularization method.
New federated learning framework reduces model complexity and improves performance.
problem Slow convergence in traditional federated learning due to non-i.i.d. data.
method Clients train personalized local models, server trains shared model, addressing heterogeneity.
result Substantial performance gains over baselines, robust to non-i.i.d. data.