New phases identified in neural scaling laws with compute limits.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian model tackles high-dimensional inverse problems efficiently.
Investigates ways to train larger models with fewer resources, finding that test loss depends only on the actual number of trainable parameters.
Spectral clustering algorithms typically require a priori selection of input parameters such as the number of clusters, a scaling parameter for the affinity measure, or ranges of these values for parameter tuning. Despite efforts for automating the process of spectral clustering, the task of grouping data in multi-scal…
Paper studies MCCR models with scale parameters tending to zero, revealing optimal learning rate and comparing robustness.
Large models follow power laws in performance with dataset size or parameters.
Let be a group acting properly and by isometries on a metric space ; it follows that the quotient or orbit space is also a metric space. We study the Vietoris-Rips and Čech complexes of . Whereas (co)homology theories for metric spaces let the scale parameter of a Vietoris-Rips or Čech complex go to z…
Optimizing full likelihoods adapts loss scales and shapes for robust modeling.
Field theory explains optimal scaling in ResNets for signal propagation.
A new approach selects tuning parameters for embedding methods.
The Random Parameters model was proposed to explain the structure of the covariance matrix in problems where most, but not all, of the eigenvalues of the covariance matrix can be explained by Random Matrix Theory. In this article, we explore other properties of the model, like the scaling of its PDF as one take larger …
We determine the critical batch size for large language models and find it scales with data size, not model size.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
The analysis of temporal networks has a wide area of applications in a world of technological advances. An important aspect of temporal network analysis is the discovery of community structures. Real data networks are often very large and the communities are observed to have a hierarchical structure referred to as mult…
Metric-based meta-learning has attracted a lot of attention due to its effectiveness and efficiency in few-shot learning. Recent studies show that metric scaling plays a crucial role in the performance of metric-based meta-learning algorithms. However, there still lacks a principled method for learning the metric scali…
Wide neural networks with asymmetrical node scaling converge globally and learn features.
We present a new method for high-dimensional linear regression when a scale parameter of the additive errors is unknown. The proposed estimator is based on a penalized Huber -estimator, for which theoretical results on estimation error have recently been proposed in high-dimensional statistics literature. However, t…
Scale of data and scale of computation infrastructures together enable the current deep learning renaissance. However, training large-scale deep architectures demands both algorithmic improvement and careful system configuration. In this paper, we focus on employing the system approach to speed up large-scale training.…
Proposes a new hyperprior and predictive criterion for weakly informative hyperprior in relevance vector machine.
Using the procedure initiated in \cite{Ma2013}, we deform Lax-type equations though a scaling of the time parameter. This gives an equivalent (deformed) equation which is integrable in terms of power series of the scaling parameter. We then describe a regular Frölicher Lie group of symmetries of this deformed equation
HET-XL improves heteroscedastic classifiers for large-scale image classification.
A grand challenge of the 21st century cosmology is to accurately estimate the cosmological parameters of our Universe. A major approach to estimating the cosmological parameters is to use the large-scale matter distribution of the Universe. Galaxy surveys provide the means to map out cosmic large-scale structure in thr…
Data-driven decision-making is performed by solving a parameterized optimization problem, and the optimal decision is given by an optimal solution for unknown true parameters. We often need a solution that satisfies true constraints even though these are unknown. Robust optimization is employed to obtain such a solutio…
Scaling laws in linear regression explain model performance improvements with size and data.
Regularization is a popular technique in machine learning for model estimation and avoiding overfitting. Prior studies have found that modern ordered regularization can be more effective in handling highly correlated, high-dimensional data than traditional regularization. The reason stems from the fact that the ordered…
Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.
We consider the weighted belief-propagation (WBP) decoder recently proposed by Nachmani et al. where different weights are introduced for each Tanner graph edge and optimized using machine learning techniques. Our focus is on simple-scaling models that use the same weights across certain edges to reduce the storage and…
We propose a novel approach for nonlinear regression using a two-layer neural network (NN) model structure with sparsity-favoring hierarchical priors on the network weights. We present an expectation propagation (EP) approach for approximate integration over the posterior distribution of the weights, the hierarchical s…
TiAda adapts adaptive gradient methods for nonconvex minimax optimization.
Investigates multifractal scaling in critical dynamics of random surfaces.
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
In this paper, we perform registration of noisy curves. We provide an appropriate model in estimating the rotation and scaling parameters to adjust a set of curves through a M-estimation procedure. We prove the consistency and the asymptotic normality of our estimators. Numerical simulation and a real life aeronautic e…
Recently, deep neural networks (DNNs) have shown advantages in accelerating optimization algorithms. One approach is to unfold finite number of iterations of conventional optimization algorithms and to learn parameters in the algorithms. However, these are forward methods and are indeed neither iterative nor convergent…
Improves kernel ridge regression by optimizing scale and feature parameters.
New method improves parameter estimation in complex stochastic models.
This study reveals the critical role of scale vectors in large language models, improving optimization and expressivity.
Generatability in metric spaces studied with novel novelty parameters.
We discuss the approximation of the value function for infinite-horizon discounted Markov Reward Processes (MRP) with nonlinear functions trained with the Temporal-Difference (TD) learning algorithm. We first consider this problem under a certain scaling of the approximating function, leading to a regime called lazy tr…
Global minima found for multidimensional scaling with penalties.
In seeking for sparse and efficient neural network models, many previous works investigated on enforcing L1 or L0 regularizers to encourage weight sparsity during training. The L0 regularizer measures the parameter sparsity directly and is invariant to the scaling of parameter values, but it cannot provide useful gradi…
We construct a general stochastic process and prove weak convergence results. It is scaled in space and through the parameters of its distribution. We show that our simplified scaling is equivalent to time scaling used frequently. The process is constructed as an integral with respect to a Poisson random measure which …
New method improves likelihood-free parameter estimation in complex models.
A new method, tree-SNE, solves the scale problem in t-SNE.
The study improves Poincaré and log-Sobolev inequalities on hyperbolic spaces.
Ensembles of random-feature models can't outperform a single large model.
We investigate how the final parameters found by stochastic gradient descent are influenced by over-parameterization. We generate families of models by increasing the number of channels in a base network, and then perform a large hyper-parameter search to study how the test error depends on learning rate, batch size, a…
Deep learning (DL) training-as-a-service (TaaS) is an important emerging industrial workload. The unique challenge of TaaS is that it must satisfy a wide range of customers who have no experience and resources to tune DL hyper-parameters, and meticulous tuning for each user's dataset is prohibitively expensive. Therefo…
Large GNNs trained with Graph Parallelism improve atomic simulation accuracy.