Meta-strategy learns tuning parameters for online learning methods.
problem Difficulty in setting tuning parameters for online learning methods.
method Meta-learning approach to learn parameters from past tasks.
result Meta-strategy improves on learning each task in isolation.
SpotTune adapts fine-tuning strategies per instance for improved transfer learning.
problem Improving transfer learning performance with deep neural networks.
method Adaptive fine-tuning approach using policy networks to decide whether to use pre-trained or fine-tuned layers.
result SpotTune outperforms traditional fine-tuning on 12 out of 14 standard datasets and achieves highest scores on Visual Decathlon.
This paper optimizes hyperparameters for random forest models.
problem Improving prediction performance of random forest models.
method Literature review followed by model-based optimization (MBO) using the tuneRanger R package.
result Tuning hyperparameters can significantly improve random forest model performance.
Generative Adversarial Networks improve trading strategy performance.
problem Optimizing trading strategies in a competitive market.
method Conditional Generative Adversarial Networks (cGANs) for strategy calibration and combination.
result cGANs provide outperformance over traditional techniques in generating alpha.
OTSL improves structure learning accuracy with out-of-sample and resampling strategies.
problem Determining optimal hyperparameters for structure learning algorithms.
method Out-of-sample Tuning for Structure Learning (OTSL) using resampling strategies.
result Improves graphical accuracy of structure learning algorithms.
Evolutionary Strategies optimize hyper-parameters for off-policy learning.
problem Hyper-parameter sensitivity in off-policy learning.
method Application of Evolutionary Strategies for online hyper-parameter tuning.
result Our method outperforms state-of-the-art baselines.
This paper tackles hyperparameter tuning for large-scale kernel ridge regression.
problem Hyperparameter tuning is crucial but often left to users, hindering efficiency and usability.
method Proposes a complexity regularization criterion based on a data-dependent penalty for efficient optimization.
result Demonstrates the benefit of the proposed approach through extensive empirical evaluation.
Optimally estimates a functional using nuisance function tuning and sample splitting.
problem Estimating optimal rates for a doubly robust functional.
method Combines nuisance function tuning and sample splitting strategies.
result Shows optimal rates of convergence for various estimators.
PES method reduces bias in gradient estimation for unrolled graphs.
problem High variance and bias in gradient estimation for unrolled computation graphs.
method Divide graph into unrolls, apply ES update, accumulate correction terms.
result PES provides unbiased, low-variance gradient estimates.
Paper proposes a new algorithm for efficient hyper-parameter optimization.
problem Efficient hyper-parameter tuning for machine learning models.
method Information geometric optimization with stochastic natural gradient for discrete search domains.
result The proposed algorithm achieves faster optimization than existing methods without manual tuning.
Pairs trading strategy improved using Ornstein-Uhlenbeck process.
problem Improving pairs trading strategy effectiveness.
method Used Ornstein-Uhlenbeck process to model stock price spreads.
result OU model captures signals and trends effectively but underperforms compared to naive model.
Optimal tuning for estimating ECC in proportional asymptotics.
problem Estimating Expected Conditional Covariance (ECC) under proportional asymptotics.
method Debiased ridge regression estimators for nuisance functions, sample splitting strategies, and asymptotic variance analysis.
result Prediction-optimal tuning parameters may not minimize asymptotic variance of ECC estimator.
Jointly tuning ensemble models improves performance and uncertainty calibration.
problem Improving both predictive performance and uncertainty calibration in deep ensembles.
method Investigated the impact of jointly tuning weight decay, temperature scaling, and early stopping.
result Jointly tuning ensemble models generally matches or improves performance, with significant variation across tasks.
Fine-tuning with pre-training data improves performance.
problem Limited training data for tasks.
method Theoretical analysis of excess risk bound and selection of pre-training data subset.
result Improvement in generalization performance with pre-training data.
The paper analyzes and validates two step size schedules for SGD: exponential and cosine, proving their adaptivity and performance.
problem The variability of SGD performance due to step size choice.
method Analysis and empirical evaluation of exponential and cosine step sizes.
result Exponential and cosine step sizes are adaptive to noise and achieve optimal performance without tuning hyperparameters.
Fine-tuning LLMs improves capability but harms safety, study finds.
problem Balancing capability and safety in LLM fine-tuning.
method Theoretical framework and numerical experiments for two safety-aware fine-tuning strategies.
result Characterization of fundamental limits of safety-capability trade-off in LLM fine-tuning.
We introduce a means of automating machine learning (ML) for big data tasks, by performing scalable stochastic Bayesian optimisation of ML algorithm parameters and hyper-parameters. More often than not, the critical tuning of ML algorithm parameters has relied on domain expertise from experts, along with laborious hand…
This article reviews tuning parameter selection for high-dimensional regression.
problem Choosing the optimal tuning parameter for high-dimensional regression.
method The article discusses various strategies for selecting tuning parameters.
result The optimal tuning parameter depends on the design matrix and error distribution.
Mango automates hyperparameter tuning for large-scale ML training.
problem Manual hyperparameter tuning is tedious and inefficient for large-scale machine learning.
method Parallel hyperparameter tuning with intelligent search strategies and flexible abstractions.
result Mango achieves comparable performance to Hyperopt while supporting distributed computing.
Regularizes deep networks for k-shot learning with limited data.
problem Sub-optimal and overfitting issues in fine-tuning pre-trained deep networks for k-shot learning.
method Cluster model parameters, propagate gradients within clusters, use reinforcement learning for optimal group assignments.
result Improves k-shot learning performance by more than 10% compared to state-of-the-art methods.
Unified CLIP space manipulations improve GAN adaptation with a single target image.
problem Overfitting or underfitting in fine-tuning a pre-trained generator with a single target image.
method Two-step training strategy: latent optimization in CLIP space followed by generator fine-tuning with CLIP space consistency loss.
result Our model generates diverse outputs with the target texture and outperforms baseline models.
New algorithm tunes SGMCMC hyperparameters for scalable Bayesian inference.
problem Tuning hyperparameters for SGMCMC is challenging due to lack of principled methods.
method Proposes a bandit-based algorithm using Stein discrepancies to tune hyperparameters.
result The method effectively tunes SGMCMC hyperparameters for various applications.
This paper simplifies fine-tuning for small LLMs, reducing barriers for developers.
problem Limited resources for fine-tuning large language models (LLMs) by individual developers and small organizations.
method Instruction-tuning datasets, small-sized LLMs (3B to 7B parameters), various training configurations and strategies.
result Improved model performance on benchmarks with specific training configurations, and insights into early termination and hyperparameter simplifications.
Paper analyzes fine-tuning methods for machine unlearning, proposing a new strategy to improve forgetting accuracy.
problem Improving fine-tuning methods to effectively forget specific subsets of data in machine learning models.
method Theoretical analysis and a novel Retention-Based Masking (RBM) strategy are proposed.
result RBM significantly improves unlearning accuracy while preserving retaining accuracy.
Behavior Transfer improves reinforcement learning by leveraging pre-trained policies.
problem Efficient transfer of knowledge in reinforcement learning.
method Behavior Transfer (BT) that uses pre-trained policies for exploration.
result BT combined with pre-training leads to better solutions than without pre-training.
New adaptive step-size method for convex optimization without tuning.
problem Optimizing convex functions efficiently with stochastic gradients.
method Adapted Adaptive Gradient Descent Without Descent to stochastic setting.
result Stochastic gradient descent converges under various assumptions.
LP-FT improves personalized model training in FL by balancing generalization and personalization.
problem Federated Learning struggles with balancing global generalization and local personalization due to non-identical data distributions.
method Adapting Linear Probing followed by full Fine-Tuning (LP-FT) to the FL setting.
result LP-FT outperforms standard fine-tuning in balancing personalization and generalization across various datasets and PFT variants.
The paper formalizes hyperparameter tuning and benchmarks six algorithms.
problem Optimizing machine learning hyperparameters.
method Formalized problem, defined defaults, conducted benchmarking study.
result Default values and tunability measures for hyperparameters.
Self-play fine-tuning improves diffusion models for text-to-image generation.
problem Plateauing performance of diffusion models after data saturation.
method Self-play fine-tuning (SPIN-Diffusion) using competition among model versions.
result Significantly improved model performance and human preference alignment.
ATLAS adapts HMC step size and trajectory length for complex geometries.
problem Sampling complex geometries with constant step size HMC/NUTS.
method Adapts step size and trajectory length using local Hessian and no U-turn condition.
result ATLAS accurately samples complex geometries, outperforming NUTS.
Bayesian optimization tunes Kalman filters more efficiently.
problem Manual tuning of Kalman filters is time-consuming and prone to local minima.
method Developed a Bayesian optimization strategy to automatically tune Kalman filters.
result Bayesian optimization identifies multiple local minima and provides uncertainty quantification.
This paper proposes reusing CNN layers to speed up hyperparameters tuning.
problem Time-consuming hyperparameters tuning in CNNs.
method Reuse trained convolutional layers among different trainings.
result Reduces training time and increases accuracy of neural networks.
Optimizes sparse fine-tuning for privacy in neural networks.
problem Performance gap between DP-SGD and non-private fine-tuning.
method Optimization-based approach using private gradient information for selecting trainable weights.
result Our selection method leads to better prediction accuracy compared to existing approaches.
New algorithms eliminate stepsize tuning for bilevel optimization problems.
problem Bilevel optimization problems with unknown parameters and stepsizes.
method D-TFBO and S-TFBO algorithms with adaptive stepsizes.
result Achieve performance comparable to well-tuned approaches with theoretical guarantees.
This work improves NAT translation accuracy through fine-tuning with curriculum learning.
problem Improving NAT translation accuracy while maintaining inference speed.
method Curriculum learning applied to fine-tuning of a pre-trained AT model to a NAT model.
result Significant improvement in translation accuracy (over 1 BLEU score) and speedup in inference.
A meta-learning method learns adaptive robust loss functions for noisy labels.
problem Handling robust learning with noisy labels and optimizing hyperparameters.
method Adaptive learning of robust loss hyperparameters through mutual improvement with network parameters.
result Generalized and effective robust loss functions with good generalization capability.
New method preserves both benign and robust accuracy in deep neural networks.
problem Large neural networks have high computational and storage costs.
method Pruning with fine-tuning to preserve both benign and robust accuracy.
result Average 93% benign accuracy, 92.5% empirical robust accuracy, and 85.0% verifiable robust accuracy preserved while compressing by 10x.
By adopting the polynomial interpolation method, we propose an approach to hedge against the interest-rate risk of the default-free bonds by measuring the nonparallel movement of the yield-curve, such as the translation, the rotation and the twist. The empirical analysis shows that our hedging strategies are comparable…
Framework fine-tunes foundation models with semi-supervised learning for downstream tasks and latent spaces.
problem Training foundation models with limited labelled data.
method Mutual information decomposition for downstream and latent spaces, semi-supervised fine-tuning.
result Significant improvements in classification tasks under low-labelled conditions.
SMART-FAN-Lasso fine-tunes neural networks for high-dimensional nonparametric regression.
problem Fine-tuning neural networks for high-dimensional nonparametric regression with variable selection.
method Source-model-augmented residual tuning (SMART) framework for neural Lasso.
result SMART-FAN-Lasso achieves statistical acceleration over single-task learning under precise conditions.
A new method for hyperparameter tuning across multiple datasets and objectives.
problem Hyperparameter tuning for black-box functions, especially across different datasets and objectives.
method Quantile-based regression with Gaussian Copula distribution for semi-parametric modeling, combined with Thompson sampling and Gaussian Copula process.
result Significant improvements in hyperparameter optimization and neural architecture search over state-of-the-art methods.
HAMLET optimizes algorithm selection for machine learning tasks.
problem Limited time budgets and computational resources make traditional bandit approaches ineffective for automated algorithm selection.
method HAMLET incorporates learning curve extrapolation and time-awareness to select machine learning algorithms.
result HAMLET variants outperform other bandit-based strategies in experiments with recorded hyperparameter tuning traces.
Ortho-MADS optimizes SVM hyperparameters for better accuracy.
problem Optimizing hyperparameters for SVM with Gaussian kernel.
method Deterministic Mesh Adaptive Direct Search (MADS) with orthogonal directions (Ortho-MADS).
result Ortho-MADS consistently finds comparable or better solutions than other methods.
Study on how attention in prompt-tuning affects large language models.
problem Limited theoretical understanding of prompt-tuning and attention in LLMs.
method Exploration of prompt-tuning for one-layer attention architectures, contextual mixture-models, and self-contained prompt-attention model.
result Softmax-prompt-attention is more expressive than self-attention and linear-prompt-attention under contextual data model.
Pairs trading strategy fails to outperform market benchmarks, but performs well during bear markets.
problem The validity of pairs trading as a profitable strategy in modern markets.
method Used common distance and cointegration methods on US equities from 1990 to 2020, including the Covid-19 crisis.
result The pairs trading strategy does not consistently outperform market benchmarks, but performs well during bear markets.
Autotune optimizes Lasso tuning parameters efficiently and accurately.
problem Efficiently selecting Lasso tuning parameters in high-dimensional models.
method Alternates optimizing regression coefficients and noise standard deviation.
result Autotune outperforms existing methods in low signal-to-noise regimes.
EVA adapts LoRA for faster, more efficient fine-tuning.
problem Fast and efficient fine-tuning of large models for specific tasks.
method EVA uses directions capturing most activation variance for initialization, maximizing gradient signal and reducing parameters.
result EVA achieves faster convergence and higher average scores across tasks, reducing parameters.
Proposes a low-cost method to set hyperparameters using optimized default values.
problem Challenges of setting hyperparameters by trial and error, leading to subjective and inefficient results.
method Generates optimized default values using a small set of values that outperform existing defaults and tuned values.
result New default values deliver better predictive performance and are competitive with tuned values, making them easier to use.