Adaptive batch size schedules improve language model training efficiency and generalization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Firm size data usually do not show the normality that is often assumed in statistical analysis such as regression analysis. In this study we focus on two firm size data: the number of employees and sale. Those data deviate considerably from a normal distribution. To improve the normality of those data we transform them…
We determine the critical batch size for large language models and find it scales with data size, not model size.
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
A new scaling law predicts optimal batch size for training models.
Study finds optimal vocabulary size for neural machine translation.
Study examines mean estimation in high dimensions with small data.
We make policy optimization algorithms batch size-invariant by decoupling proximal and behavior policies.
This work studies scaling laws for low-precision training in high-dimensional linear regression.
We consider -dimensional linear stochastic approximation algorithms (LSAs) with a constant step-size and the so called Polyak-Ruppert (PR) averaging of iterates. LSAs are widely applied in machine learning and reinforcement learning (RL), where the aim is to compute an appropriate (that is a…
Muon optimizes training efficiency by improving data retention at large batch sizes.
Adaptive batch sizes improve local gradient methods in distributed training.
New kernel model scales to large datasets.
Sample size determination for a data set is an important statistical process for analyzing the data to an optimum level of accuracy and using minimum computational work. The applications of this process are credible in every domain which deals with large data sets and high computational work. This study uses Bayesian a…
An algorithm reduces breast cancer detection data complexity using effect sizes.
Paper presents a new way to estimate model changes without full model evaluation.
We introduce a new system of stochastic differential equations which models dependence of market beta and unsystematic risk upon size, measured by market capitalization. We fit our model using size deciles data from Kenneth French's data library. This model is somewhat similar to generalized volatility-stabilized model…
Adaptive SGD learns optimal batch size for strong convex functions.
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
We propose a novel method to train deep convolutional neural networks which learn from multiple data sets of varying input sizes through weight sharing. This is an advantage in chemometrics where individual measurements represent exact chemical compounds and thus signals cannot be translated or resized without disturbi…
Modeling functional data, this study uncovers the size-and-shape of functions under noisy observations.
Exact distribution of split conformal prediction coverage found.
Estimates unknown population sizes using the hypergeometric distribution.
We study the sample complexity of private synthetic data generation over an unbounded sized class of statistical queries, and show that any class that is privately proper PAC learnable admits a private synthetic data generator (perhaps non-efficient). Previous work on synthetic data generators focused on the case that …
A microscopic model of aggregation and fragmentation is introduced to investigate the size distribution of businesses. In the model, businesses are constrained to comply with the market price, as expected by the customers, while customers can only buy at the prices offered by the businesses. We show numerically and ana…
This paper finds a linear relationship between t-SNE perplexity and data set size.
Study finds that only a fraction of data is needed for accurate patient-level prediction models.
Conformal Prediction is a machine learning methodology that produces valid prediction regions under mild conditions. In this paper, we explore the application of making predictions over multiple data sources of different sizes without disclosing data between the sources. We propose that each data source applies a trans…
SPREV simplifies visualization of complex labeled datasets.
Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training. Among the simplest ways to harness next-generation hardware is to increase the batch size in standard mini-batch neural network training algorithms. In this work, we aim to experimentally charac…
In this paper we study a family of variance reduction methods with randomized batch size---at each step, the algorithm first randomly chooses the batch size and then selects a batch of samples to conduct a variance-reduced stochastic update. We give the linear convergence rate for this framework for composite functions…
Unified Bayesian model for multi-modal, small sample size biomedical data classification.
Solves a model for sudden problem-solving ability in deep learning.
Two methods estimate effect size for online experiments, improving accuracy and efficiency.
A new line search rule improves support recovery in high-dimensional data.
This research examines how the error rate of nearest neighbor classifiers varies with dataset size.
How does missing data affect our ability to learn signal structures? It has been shown that learning signal structure in terms of principal components is dependent on the ratio of sample size and dimensionality and that a critical number of observations is needed before learning starts (Biehl and Mietzner, 1993). Here …
Support vector regression (SVR) has been widely used to reduce the high computational cost of computer simulation. SVR assumes the input parameters have equal sample sizes, but unequal sample sizes are often encountered in engineering practices. To solve this issue, a new prediction approach based on SVR, namely as hig…
Paper develops an online learning algorithm for functional data models.
Deep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a deep random feature (DRF) expansion model, which makes inference tractable by app…
In this study, we consider classification problems based on neural networks in data-imbalanced environment. Learning from an imbalanced data set is one of the most important and practical problems in the field of machine learning. A weighted loss function based on cost-sensitive approach is a well-known effective metho…
The paper improves A/B testing for non-Gaussian data, ensuring reliable results with large sample sizes.
In this paper we address the question of the size distribution of firms. To this aim, we use the Bloomberg database comprising multinational firms within the years 1995-2003, and analyze the data of the sales and the total assets of the separate financial statement of the Japanese and the US companies, and make a compa…
Scaling laws in linear regression explain model performance improvements with size and data.
We address the issue of the distribution of firm size. To this end we propose a model of firms in a closed, conserved economy populated with zero-intelligence agents who continuously move from one firm to another. We then analyze the size distribution and related statistics obtained from the model. Our ultimate goal is…
In biospectroscopy, suitably annotated and statistically independent samples (e. g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training sample size and can help to determine the sample size needed to train good classi…
Paper reduces recommender system model size by 90%.
Estimates sample size for subgroup analysis in randomized experiments.