Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2595177761,034 · Jun 202019922001200920182026
48 results for Training Parameters

Training-free model learns SDE dynamics without training, accelerating parameter studies.

problem High computational cost of simulating parameter-dependent SDEs.
method Training-free conditional diffusion model with joint kernel-weighted Monte Carlo estimator.
result Accurate approximation of conditional distributions across varying parameter values.

Stage-based hyper-parameter optimization reduces GPU-hours and training time.

problem Efficiently executing hyper-parameter optimization for deep learning models.
method Stage-based execution strategy to remove redundant computations.
result Stage-based execution outperforms trial-based method by up to 6.60 times in GPU-hours and 4.13 times in training time.

Lapse improves parameter servers by dynamically allocating parameters, achieving near-linear scaling.

problem Efficiently managing distributed training with reduced communication overhead.
method Integrate dynamic parameter allocation into parameter servers, proposing Lapse.
result Lapse provides near-linear scaling and can be orders of magnitude faster than existing parameter servers.

We introduce a new parameter to measure the inhomogeneity of training datasets.

problem The need for non-stationary models in supervised learning.
method We introduce a new parameter, the inhomogeneity parameter, to measure the inhomogeneity of training datasets.
result A training set with a non-zero inhomogeneity parameter requires a non-stationary model for accurate predictions.

ABPS improves RL training efficiency by sharing policies and evolving hyper-params.

problem Data inefficiency in training deep RL models for real-world applications.
method ABPS: adaptive behavior policy sharing; ABPS-PBT: hybridizing ABPS with PBT for evolving hyper-params.
result ABPS achieves superior performance and reduced variance compared to conventional hyper-parameter tuning.

A new Federated Learning approach balances personalization and global training.

problem Breaking the curse of data heterogeneity in Federated Learning.
method Splitting variables into global and local parameters, using a simple algorithm.
result The approach allows each client to fit their data perfectly, breaking the curse of data heterogeneity.

ZeRO optimizes memory for training large models, scaling to trillions of parameters.

problem Training models with billions to trillions of parameters is challenging due to limited device memory.
method ZeRO eliminates memory redundancies in data- and model-parallel training, scaling model size proportional to the number of devices.
result ZeRO trains models of up to 13B parameters without model parallelism, achieving super-linear speedup and throughput of 15 Petaflops.

OPNP prunes parameters and neurons to improve OOD detection without training.

problem Detecting out-of-distribution samples in real-world machine learning models.
method OPNP approach that identifies and removes sensitive parameters and neurons.
result OPNP consistently outperforms existing methods on multiple OOD detection tasks.

Single neural network predicts ImageNet model parameters for faster training.

problem Training diverse ImageNet models requires significant resources and time.
method Trained a neural network to predict ImageNet model parameters and used them for initialization.
result Models initialized with predicted parameters converge faster and achieve competitive performance.

Stanza separates convolutional and fully connected layers for faster deep learning training.

problem Heavy data transfer between workers and servers in distributed deep learning.
method Layer separation: most nodes train convolutional layers, others train fully connected layers only.
result Significant acceleration of training time (1.34x--13.9x) over current systems.

Federated learning improves with adaptive hyper-parameters and representation matching.

problem Heterogeneous client data leads to divergent local models in federated learning.
method Representation matching and adaptive hyper-parameters.
result Significant performance and robustness improvements in federated learning.

Training only BatchNorm parameters achieves surprisingly high performance in deep networks.

problem Understanding the role and expressive power of affine parameters in BatchNorm.
method Investigating performance when training only BatchNorm parameters and freezing all other weights.
result BatchNorm achieves high performance by learning to disable around a third of random features.

Trains neural networks to efficiently solve Navier-Stokes equations across parameter space.

problem Efficiently solving Navier-Stokes equations in parameter space.
method Physics-informed neural networks, active learning algorithm.
result Neural networks can accurately interpolate and aggregate solutions to physical problems.

Hippo optimizes deep learning hyper-parameters by reducing redundant trials.

problem Redundant hyper-parameter trials in hyper-parameter optimization.
method Hippo breaks down hyper-parameter sequences into stages and executes them in a tree structure.
result Hippo reduces GPU-hours and training time significantly compared to existing methods.

ES for non-differentiable parameters scales to large models.

problem Learning non-differentiable parameters in large models.
method Hybrid approach combining ES for non-differentiable and gradient-based methods for differentiable parameters.
result Hybrid approach is competitive and allows training sparse models from the start.

This report proposes efficient ways to set hyper-parameters for neural networks.

problem Setting hyper-parameters remains a black art requiring years of experience.
method Examining training validation/test loss function, adjusting learning rate/momentum, balancing regularization.
result Significantly reduces training time and improves performance.

Syllable-aware models perform similarly to character-based ones but use fewer parameters and train faster.

problem Improving word-level language modeling performance with syllable-based models.
method Used a syllable-aware neural language model with fewer parameters and faster training.
result Achieved comparable performance to character-based models but with 18%-33% fewer parameters and 1.2-2.2 times faster training.

Study on rich regime training in deep learning, finding active parameters in bottom layers.

problem Understanding the practical success of deep learning models.
method Empirical study on rich regime training with benchmark datasets, re-initialization analysis, and probabilistic Layer-Wise Sparse SGD.
result Probabilistic Layer-Wise Sparse SGD matches vanilla SGD's generalization performance with improved efficiency.

TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.

problem Limited expressivity and generalization of standard LoRA.
method TensorGuide uses a unified tensor-train structure with controlled Gaussian noise to generate correlated low-rank matrices.
result TensorGuide achieves superior accuracy and scalability with fewer parameters compared to standard LoRA and TT-LoRA.

FURL improves model accuracy in FL by locally training user embeddings.

problem Improving prediction accuracy of neural-network-based models in Federated Learning.
method FURL divides model parameters into federated and private parameters, training private parameters locally.
result Significant performance improvement with 8% and 51% increases on two datasets.

Transformer learns to estimate negative binomial parameters efficiently.

problem Parameter estimation for over-dispersed count data in large screens.
method Pre-trained transformer trained on synthetic data generation to invert parameter to count transformation.
result Method of moments provides faster, more efficient, and better-calibrated estimates.

Study analyzes convergence of parameter estimation in contaminated mixture of experts.

problem Challenges in learning from prompts in large-scale models.
method Convergence analysis, distinguishability condition, partial differential equations.
result Comprehensive convergence rates and minimax lower bounds for parameter estimation.

New algorithm uses PSO to optimize DNN training parameters in distributed systems.

problem Reducing synchronization frequency in DNN training leads to poor convergence.
method Integrates PSO into distributed training to automatically compute new parameters.
result Proposed algorithm outperforms synchronous methods in distributed DNN training.

Autoencoder estimates parameters of noisy, multi-component damped signals.

problem Parameter estimation of damped sinusoidal signals under rapid decay and noise.
method Autoencoder-based approach using latent space for frequency, phase, decay, and amplitude estimation.
result High accuracy in parameter estimation, robustness to subdominant components and phase differences.

NPAS trains neural networks with a fixed parameter budget, improving performance and compactness.

problem Training neural networks requires memory, and existing methods struggle with arbitrary parameter budgets.
method NPAS learns to share parameters automatically, covering low and high budgets.
result NPAS and SSNs improve network performance and compactness across various tasks.

KD technique improves QDNN performance with reduced hyper-parameters.

problem Restoring performance loss in QDNNs due to quantization.
method Applied KD with reduced hyper-parameters, including a new coefficient reduction technique.
result Achieved 92.7% test accuracy on CIFAR-10 and 67.0% on CIFAR-100 with 2-bit weights.

Fine-tuning large language models requires minimal data, making them efficient.

problem Achieving state-of-the-art performance with large language models.
method Using BERT as an example, fine-tuning only the most critical layers of the pre-trained model.
result Fine-tuned models are close in parameter space to the pre-trained model, with many good solutions found in sparsified versions.

CoNNTrA trains DNNs with low-power, low-memory constraints.

problem Training deep neural networks on edge computing systems with low power and memory usage.
method Coordinate gradient descent-based approach for training DNNs with constrained learning parameters.
result CoNNTrA models use 32x less memory and have comparable errors to Backpropagation models.

PEP improves deep network performance and calibration by perturbing optimal parameters.

problem Improving deep network performance and calibration.
method Parameter Ensembling by Perturbation (PEP) constructs an ensemble of parameter values as random perturbations of the optimal set, maximizing log-likelihood on validation data.
result PEP provides a small to substantial improvement in calibration and log-likelihood, and in some cases, classification accuracy.

The quality of an induced model by a learning algorithm is dependent on the quality of the training data and the hyper-parameters supplied to the learning algorithm. Prior work has shown that improving the quality of the training data (i.e., by removing low quality instances) or tuning the learning algorithm hyper-para…

2014-03-13abs ↗pdf ↗

The paper provides a method to find optimal machine learning model parameters with confidence.

problem Finding optimal machine learning model parameters that generalize well to the entire population.
method Constructs valid confidence sets for the optimal parameter using only training data.
result Valid confidence sets for optimal machine learning model parameters can be generated using bootstrapping techniques.

Lazy training and mean field regimes studied for TD learning with nonlinear function approximation.

problem Approximating value function for MRP with TD learning and nonlinear functions.
method Lazy training and mean field scaling of parameters analyzed for convergence.
result Lazy training leads to exponential convergence to local/global minimizers, while mean field scaling results in all fixed points being minimizers.

This paper speeds up large-scale deep learning training.

problem Training large-scale deep architectures is slow and resource-intensive.
method Systematic approach to identify bottlenecks, develop guidelines, and derive lemmas.
result Developed procedures and lemmas for setting minibatch size, choosing algorithms, and determining component quantities.

Adversarial training improves linear regression solutions, revealing sparsity and abrupt interpolation.

problem Adversarial attacks on linear regression models.
method Formulated as a convex problem, adversarial training is used to find robust solutions that are sparse and interpolate data.
result Adversarial training with small disturbances gives the solution with the minimum-norm that interpolates the training data, revealing abrupt transition into interpolation.