Standard optimizers perform as well as LARS and LAMB at large batch sizes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Geodesic interpretation of global quasi-geostrophic equations on sphere.
Training large deep neural networks on massive datasets is computationally very challenging. There has been recent surge in interest in using large batch stochastic optimization methods to tackle this issue. The most prominent algorithm in this line of research is LARS, which by employing layerwise adaptive learning ra…
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
Accelerates BERT pretraining from 3 days to 54 minutes.
To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce c…
Deep learning estimates time-varying Markov model parameters.
Training-free model learns SDE dynamics without training, accelerating parameter studies.
Dual Bayesian Affine Estimators for Wiener-type state-space models
New method achieves optimal performance without needing problem parameters.
Study shows how numerical discretization affects reconstructions and parameter distributions in nano metrology.
We introduce two approaches for combining neural evolution strategy (NES) and proximal policy optimization (PPO): parameter transfer and parameter space noise. Parameter transfer is a PPO agent with parameters transferred from a NES agent. Parameter space noise is to directly add noise to the PPO agent`s parameters. We…
Develops a parameter-free SGD algorithm with optimal convergence rate.
Deep learning outperforms traditional methods in estimating OU process parameters.
Bayesian classification and regression with high order interactions is largely infeasible because Markov chain Monte Carlo (MCMC) would need to be applied with a great many parameters, whose number increases rapidly with the order. In this paper we show how to make it feasible by effectively reducing the number of para…
Paper proposes a reinforcement learning framework for efficient hyper-parameter tuning of stochastic optimization algorithms.
In the regression setting, given a set of hyper-parameters, a model-estimation procedure constructs a model from training data. The optimal hyper-parameters that minimize generalization error of the model are usually unknown. In practice they are often estimated using split-sample validation. Up to now, there is an ope…
Recent years, transfer learning has attracted much attention in the community of machine learning. In this paper, we mainly focus on the tasks of parameter transfer under the framework of extreme learning machine (ELM). Unlike the existing parameter transfer approaches, which incorporate the source model information in…
FNFs model parameter-dependent densities by combining a fixed flow with a polynomial parameter-dependent transformation.
ThriftyNet uses a single convolutional layer recursively to maximize parameter usage.
In this paper, we introduce a new parameter, the affine twist parameter for the affine deformation of a sphere with holes. We show that the affine deformation space can be parametrized by Margulis invariants and affine twist parameters. The affine twist parameter is canonically regarded as a correspondence to the Fench…
A new method estimates parameters of complex models using ordinary least squares.
Algorithms typically come with tunable parameters that have a considerable impact on the computational resources they consume. Too often, practitioners must hand-tune the parameters, a tedious and error-prone task. A recent line of research provides algorithms that return nearly-optimal parameters from within a finite …
This paper proposes an efficient autoHPO method based on data-to-hyper-parameter mapping.
Empirical study shows removing neural parameter symmetries impacts model performance.
Supervised learning is an active research area, with numerous applications in diverse fields such as data analytics, computer vision, speech and audio processing, and image understanding. In most cases, the loss functions used in machine learning assume symmetric noise models, and seek to estimate the unknown function …
The paper defines and studies canonical parameters on surfaces in 4D space.
New findings on Malgrange-Galois groupoid for Painlevé VI equation parameters.
Deep neural networks with specific parameter sets can approximate smooth functions efficiently.
As deep learning techniques advance more than ever, hyper-parameter optimization is the new major workload in deep learning clusters. Although hyper-parameter optimization is crucial in training deep learning models for high model performance, effectively executing such a computation-heavy workload still remains a chal…
NanoFlow reduces parameter complexity in normalizing flows.
We propose and demonstrate the use of a model-assisted generative adversarial network (GAN) to produce fake images that accurately match true images through the variation of the parameters of the model that describes the features of the images. The generator learns the model parameter values that produce fake images th…
We present two different approaches for parameter learning in several mixture models in one dimension. Our first approach uses complex-analytic methods and applies to Gaussian mixtures with shared variance, binomial mixtures with shared success probability, and Poisson mixtures, among others. An example result is that …
Study optimizes sensor placement for accurate parameter estimation in complex systems.
Deep learning used for parameter estimation in hard-to-infer models.
Bayesian active learning tackles nuisance parameters, leading to bias and dilemmas.
In order to quantize the gate parameters of the LSTM (Long Short-Term Memory) neural network model with almost no recognition performance degraded, a new quantization method named Quantization Loss Re-Learn Method is proposed in this paper. The method does lossy quantization on gate parameters during training iteration…
We study the Euler-Lagrange equations for a parameter dependent -invariant Lagrangian on a homogeneous -space. We consider the pullback of the parameter dependent Lagrangian to the Lie group , emphasizing the special invariance properties of the associated Euler-Poincaré equations with advected parameters.
We address the problem of parameter estimation in models of systems biology from noisy observations. The models we consider are characterized by simultaneous deterministic nonlinear differential equations whose parameters are either taken from in vitro experiments, or are hand-tuned during the model development process…
This work improves Bayesian Optimization for setting DNN hyper-parameters.
Hippo optimizes deep learning hyper-parameters by reducing redundant trials.
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP. However, in the presence of many downstream tasks, fine-tuning is parameter inefficient: an entire new model is required for every task. As an alternative, we propose transfer with adapter modules. Adapter modules yield a compact and extens…
Study on measuring vulnerability of neural network parameters via corruption.
We consider the bridge linear regression modeling, which can produce a sparse or non-sparse model. A crucial point in the model building process is the selection of adjusted parameters including a regularization parameter and a tuning parameter in bridge regression models. The choice of the adjusted parameters can be v…
This paper presents the asymptotic behavior of a linear instrumental variables (IV) estimator that uses a ridge regression penalty. The regularization tuning parameter is selected empirically by splitting the observed data into training and test samples. Conditional on the tuning parameter, the training sample creates …
Introduces a new stationary GE-process for gold price analysis.
Let and . We construct -parameters, -parameters, -parameters ancient solutions of the equation , , in for some . This equation arises in the study of Yamabe flow. We obtain various properties of the ancient so…
Deep networks can approximate functions with fewer learnable parameters than previously thought.