Bayesian method pools spatial neural data to improve tuning function estimates.
problem Interpreting neural activity with spatially adjacent neurons.
method Block Gibbs sampler that pools information between neurons.
result The method de-noises tuning function estimates while preserving sharp discontinuities.
Bayesian Optimization improves Neural Network parameter tuning for XOR function.
problem Optimizing model parameters for neural networks, especially for large models.
method Bayesian Optimization with Gaussian Process Priors.
result Achieved higher prediction accuracy for the XOR function.
Fine-tuning neural networks to guarantee performance on specific examples can also introduce incorrect inputs.
problem Ensuring reliable performance of neural networks on specific examples.
method Using SMT solvers to fine-tune ReLU neural networks to guarantee outcomes on a finite set of particular examples.
result Fine-tuning can introduce incorrect inputs that trigger unexpected performance.
This paper introduces an efficient method for optimizing deep learning hyperparameters.
problem The high dependency of deep learning algorithms on hyper-parameters.
method Orthogonal Array Tuning Method (OATM) for deep learning hyper-parameter tuning.
result The proposed OATM method significantly saves tuning time compared to state-of-the-art methods.
Optimizes sparse fine-tuning for privacy in neural networks.
problem Performance gap between DP-SGD and non-private fine-tuning.
method Optimization-based approach using private gradient information for selecting trainable weights.
result Our selection method leads to better prediction accuracy compared to existing approaches.
Improved neural population modeling using shared features and ensemble detection.
problem Missing shared coding properties in neural latent variable models.
method Feature sharing across tuning curves and soft clustering of neurons.
result More interpretable and better-performing neural population models.
SMART-FAN-Lasso fine-tunes neural networks for high-dimensional nonparametric regression.
problem Fine-tuning neural networks for high-dimensional nonparametric regression with variable selection.
method Source-model-augmented residual tuning (SMART) framework for neural Lasso.
result SMART-FAN-Lasso achieves statistical acceleration over single-task learning under precise conditions.
Two retraining techniques outperform fine-tuning in neural network pruning.
problem Improving accuracy and compression in neural network pruning.
method Weight rewinding and learning rate rewinding compared to fine-tuning.
result Rewinding techniques outperform fine-tuning in accuracy and compression.
Transfer learning boosts chemically accurate neural network potentials for organic molecules.
problem Developing accurate interatomic potentials from ab-initio data.
method Discriminative fine-tuning of pre-trained neural networks.
result Fine-tuning with energy labels alone can achieve accurate atomic forces.
The paper introduces a Hessian-based method to improve generalization in fine-tuned deep neural networks.
problem Improving generalization in fine-tuned deep neural networks, especially in noisy conditions.
method PAC-Bayesian analysis to identify a Hessian-based distance measure, proving generalization bounds, and developing an algorithm with a generalization error guarantee.
result Hessian-based distance measure correlates well with observed generalization gaps and can match the scale of these gaps in practice.
New metric captures individual neuron tuning across neural networks.
problem Need a metric that respects individual neuron tuning across different neural networks.
method Derived a 'soft' permutation-based metric using optimal transport theory.
result Metric avoids counter-intuitive outcomes and captures geometric insights.
The paper develops a theory linking pretraining and fine-tuning in neural networks.
problem Understanding how initialization choices impact feature learning and generalization in neural networks.
method Analytical theory of diagonal linear networks, deriving generalization error as a function of initialization parameters and task statistics.
result Different initialization choices place networks into four fine-tuning regimes with varying abilities to support feature learning and generalization.
A new method constrains deep networks during fine-tuning to improve generalization.
problem Improving generalization of fine-tuned deep networks.
method A neural network generalisation bound based on distance from initial weights constrains the hypothesis class to a small sphere.
result Empirical evaluation shows superior generalization performance compared to existing methods.
New measure FTC quantifies how much a ReLU network can fine-tune.
problem Analyzing memorization capacity in fine-tuned neural networks.
method Defined Fine-Tuning Capacity (FTC) for additive fine-tuning of ReLU networks.
result Upper and lower bounds on FTC for 2 and 3-layer ReLU networks.
This paper studies learning rate policies for deep neural networks, offering metrics and tools for better tuning.
problem Effective tuning of learning rates for deep neural networks is challenging and crucial for achieving high accuracy.
method Comprehensive study of 13 learning rate functions, proposing metrics for evaluation, and developing LRBench for benchmarking and selection.
result Identification of good learning rate policies with effective ranges and step sizes for various LR update schedules.
CascadeML trains multi-label models without hyperparameter tuning.
problem Improving multi-label classification performance through hyperparameter optimization.
method CascadeML uses a two-phase incremental learning approach to train multi-label neural networks.
result CascadeML outperforms other multi-label classification algorithms without hyperparameter tuning.
ICE-Pruning accelerates deep neural network pruning by 9.61x.
problem Efficiently pruning deep neural networks while maintaining accuracy.
method Iterative pruning with automatic fine-tuning steps, freezing strategy, and custom learning rate scheduler.
result Significantly reduces pruning time by up to 9.61x.
Proposes a learning rate method for shallow nets based on gradient Lipschitz constant.
problem Finding optimal learning rates for shallow neural networks.
method Associates learning rate with gradient Lipschitz constant and proposes a search algorithm.
result The proposed method significantly outperforms existing tuning methods.
SpotTune adapts fine-tuning strategies per instance for improved transfer learning.
problem Improving transfer learning performance with deep neural networks.
method Adaptive fine-tuning approach using policy networks to decide whether to use pre-trained or fine-tuned layers.
result SpotTune outperforms traditional fine-tuning on 12 out of 14 standard datasets and achieves highest scores on Visual Decathlon.
DCTR uses neural networks to improve particle physics simulations and parameter tuning.
problem High computational cost limits precise scientific analysis in particle physics.
method Deep neural networks for reweighting and parameter tuning of simulations.
result DCTR enables precise simulations and parameter tuning, improving model accuracy.
Self-supervised fine-tuning corrects SR CNNs for unseen models and artifacts.
problem SR CNNs' lack of robustness to unseen image formation models and generation of artifacts.
method Iterative fine-tuning using a data fidelity loss at test time.
result Successfully corrects SR solutions for unseen models and GAN artifacts.
Multirate training speeds up neural network fine-tuning.
problem Efficiently fine-tuning deep neural networks.
method Partitioning neural network parameters into fast and slow parts, updating slowly over longer intervals.
result Significant computational speed-up for transfer learning tasks.
New benchmark protocol evaluates neural network optimizers for efficiency and data shift sensitivity.
problem Benchmarking neural network optimizers with hyperparameter complexity and data shift sensitivity.
method Proposed a new evaluation protocol combining end-to-end and data-addition training efficiency, using bandit hyperparameter tuning and human study validation.
result No clear winner across all tasks, highlighting the complexity of optimizer performance.
Paper proves multiplicative weight updates can train neural networks without learning rate tuning.
problem Vanishing and exploding gradients in gradient descent for compositional functions.
method Proves descent lemma for compositional functions using multiplicative weight updates and derives Madam optimizer.
result Madam optimizer trains state-of-the-art neural networks without learning rate tuning.
Proposes efficient multi-fidelity Bayesian optimization for deep neural network hyperparameter tuning.
problem Time-consuming validation error evaluation for hyperparameter tuning in deep neural networks.
method Introduces trace-aware knowledge-gradient acquisition function and a provably convergent optimization method.
result Outperforms state-of-the-art alternatives for hyperparameter tuning of deep neural networks.
Cold posteriors improve Bayesian neural networks by reducing overestimation of aleatoric uncertainty.
problem Overestimation of aleatoric uncertainty in Bayesian neural networks.
method Tuning the temperature of the posterior on a validation set.
result Reducing temperature leads to better reflection of true prior beliefs.
Sherpa automates hyperparameter tuning for machine learning models.
problem Hyperparameter optimization for computationally expensive models.
method Powerful and interchangeable algorithms, interactive dashboard.
result Automates tedious aspects of model tuning.
Fine-tunes GNNs by preserving generative patterns to improve transferability.
problem Vanilla fine-tuning fails due to structural divergence between pre-training and downstream graphs.
method G-Tuning, which reconstructs the generative patterns of the downstream graph using graphon bases.
result G-Tuning achieves an average improvement of 0.5% and 2.6% on in-domain and out-of-domain transfer learning experiments.
Reduces annotation costs in medical imaging by 50%.
problem Challenges in creating large annotated datasets for medical imaging.
method Integrates active learning and transfer learning into a single framework.
result Reduces annotation efforts by at least half.
Theoretical framework explains why few epochs are enough for LLM fine-tuning.
problem Understanding why few epochs are sufficient for LLM fine-tuning.
method Combining early stopping theory with attention-based Neural Tangent Kernel (NTK) for LLMs.
result Formalizes convergence rate of attention-based fine-tuning with respect to sample size.
HyperNOMAD automates deep neural network hyperparameter tuning.
problem Automating the calibration of hyperparameters in deep neural networks.
method Derivative-free optimization using MADS algorithm.
result Achieves comparable results to state-of-the-art methods.
Bayesian optimization improves hyperparameter tuning for deep learning.
problem Efficient tuning of hyperparameters in deep learning is expensive and time-consuming.
method Bayesian optimization exploiting iterative learning progress.
result Our algorithm identifies optimal hyperparameters faster than existing methods.
Bayesian optimization sped up with importance sampling.
problem Efficiently tuning hyperparameters for neural networks.
method Bayesian optimization with importance sampling.
result Significantly improved runtime and validation error.
Auptimizer simplifies hyperparameter tuning for machine learning models.
problem Difficulty and time-consuming hyperparameter tuning for machine learning models.
method General HPO framework that distributes computing resources and integrates various HPO techniques.
result Simplified model tuning and bookkeeping for data scientists.
Alternative neural network training using monotone variational inequality.
problem Training neural networks efficiently and with guarantees.
method Using monotone variational inequality to solve non-convex problems efficiently.
result Our approach leads to fast convergence and competitive performance compared to traditional methods.
Fine-tuning normalization layers can reconstruct smaller networks.
problem Understanding the expressive power of fine-tuning normalization layers.
method Random ReLU networks and sparsified networks were fine-tuned to reconstruct target networks.
result Fine-tuning normalization layers can reconstruct networks that are O ( e x t w i d t h ) O(\sqrt{ ext{width}}) O ( e x t w i d t h ) times smaller. PASHA optimizes model tuning for large datasets with limited resources.
problem Expensive HPO and NAS for large datasets.
method Dynamic resource allocation approach.
result Significantly reduces computational resources while maintaining performance.
Improved hypernetwork for efficient neural network hyperparameter tuning.
problem Efficiently optimizing hyperparameters in neural networks.
method Proposed Δ Δ Δ -STN architecture focusing on accurate best-response Jacobian approximation. result Significantly improved hyperparameter tuning accuracy and stability.
InverSynth automatically tunes synthesizer parameters from audio input.
problem Manual tuning of synthesizer parameters is time-consuming and requires expertise.
method Strided convolutional neural networks for inferring synthesizer parameters.
result InverSynth outperforms baselines in synthesizer parameter tuning.
Automated learning rate tuning for neural networks across datasets and models.
problem Finding optimal learning rates for neural network training, especially in adversarial settings.
method A parameterless, adaptive learning rate trajectory algorithm.
result Consistently achieves top-level accuracy in natural and adversarial training.
AlgoPerf competition evaluates neural network training speed-ups.
problem Improving neural network training speed using better algorithms.
method Compared 18 diverse submissions from 10 teams on multiple workloads.
result Schedule Free AdamW algorithm achieved best results in self-tuning ruleset.
We translate ML pipelines into neural networks to optimize multiple models together.
problem Isolated training of ML pipelines limits joint optimization of multiple models.
method Propose translating ML pipelines into neural networks and fine-tuning them jointly.
result Fine-tuning translated pipelines increases final accuracy.
New optimizer SF-NorMuon matches tuned AdamW across various horizons.
problem Fixed learning-rate schedules in neural network training lead to strong path dependence and costly re-tuning.
method Schedule-Free Spectral Optimization (SF-NorMuon)
result SF-NorMuon outperforms tuned AdamW on 125M and 772M parameter models across different horizons.
This paper optimizes neural network training by packing multiple models on a single GPU.
problem Efficiently sharing limited training resources among multiple neural network models.
method Proposes a primitive called 'pack' to jointly train multiple models on a single GPU.
result Significant performance improvements for hyperparameter tuning, up to 40% for two models.
Paper introduces robust method for training neural networks.
problem Neural network hyperparameter tuning, particularly learning rate sensitivity.
method Implicit Stochastic Gradient Descent (ISGD) layer-wise approximation for neural networks.
result Our method is more robust to high learning rates and generally outperforms standard backpropagation.
GOLS-I automatically determines learning rates for various neural network training algorithms.
problem Adapting learning rates in stochastic training algorithms for neural networks.
method Gradient-Only Line Search (GOLS-I) for automatically setting learning rates.
result GOLS-I learning rate schedules are competitive with manually tuned rates across multiple algorithms, architectures, datasets, and loss functions.
Automatically tunes learning rate and momentum for SGD methods.
problem Manual hyperparameter tuning is costly and lacks theoretical justification.
method Uses statistics of gradient estimator to automatically adjust learning rate and momentum.
result Matches performance of best manual settings for CNN training.
HyperFast instantly classifies tabular data in one pass.
problem Computational burden and time-consuming hyperparameter tuning for deep learning models.
method Meta-trained hypernetwork for instant classification.
result Significantly faster than competing methods with competitive performance.