Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
problem Improving loss functions for machine learning.
method Introduces Fitzpatrick losses based on the Fitzpatrick function.
result Fitzpatrick losses are tighter than Fenchel-Young losses.
Study finds object detection systems have higher error rates for darker-skinned pedestrians.
problem Predictive bias in object detection systems for pedestrians with different skin tones.
method Annotated a large dataset and compared performance between skin tone groups, investigating factors like time of day and occlusion.
result Predictive bias in object detection systems is not solely due to more difficult scenes for darker-skinned pedestrians.
We give a functional analytical proof of the equality between the Maslov index of a semi-Riemannian geodesic and the spectral flow of the path of self-adjoint Fredholm operators obtained from the index form. This fact, together with recent results on the bifurcation for critical points of strongly indefinite functional…
Introduces bounded scale measure and generalizes property A.
problem Defining property A for large scale spaces with bounded geometry.
method Introduces bounded scale measure, shows its coarse invariance, and generalizes property A.
result Definition of property A for large scale spaces with bounded scale measure is a coarse invariant.
Transformers can scale both context and task, but MLPs can only scale task.
problem Understanding and scaling In-Context Learning in transformers.
method Simplified transformer architecture, feature map, and MLP combination.
result Simplified transformer can perform ICL and context-scaling but not task-scaling.
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.
New principles needed for scaling large language models, challenging traditional regularization methods.
problem The shift from generalization to scaling in machine learning requires new guiding principles.
method Examining the effectiveness of traditional regularization methods in the scaling-centric era.
result Traditional principles of regularization may not generalize to larger scales, highlighting new phenomena like scaling law crossover.
Improves U-Net for scale equivariance in semantic segmentation.
problem Improving generalization in semantic segmentation tasks with varying scales.
method Introduces Scale Equivariant U-Net (SEU-Net) with carefully applied subsampling and upsampling layers and scale-equivariant layers.
result Significantly improved generalization to different scales compared to U-Net and scale-equivariant architecture without upsampling.
DSS networks use scale-equivariant cross-correlations to improve image recognition.
problem Improving image recognition by exploiting scale invariance.
method Constructing scale-equivariant cross-correlations based on scale-spaces and semigroups.
result Demonstrated utility on Patch Camelyon and Cityscapes datasets.
Scale-equivariant CNNs handle scale changes for improved performance.
problem Translation equivariance is not sufficient for handling scale changes in CNNs.
method Developed scale-equivariant convolutional networks with steerable filters.
result Demonstrated state-of-the-art results on MNIST-scale and STL-10 datasets.
Introduces resemblance structure for large scale geometry.
problem Defining similarity in large scale geometry.
method Axiomatizing the concept of resemblance for subsets of a set.
result Large scale resemblance structures can induce nearness and generalize large scale properties.
New scaling laws optimize model size, training, and inference for better performance.
problem Trade-off between model size and inference cost in modern LLMs.
method Train-to-Test (T2) scaling laws that jointly optimize model size, training tokens, and inference samples. result Optimal pretraining decisions shift into overtraining regime, leading to stronger performance.
Derives a family of hyperparameter scaling strategies for neural networks.
problem Optimizing hyperparameters for wide and deep neural networks.
method Introduces a one-parameter family of hyperparameter scaling strategies.
result Reveals proper scaling of depth with width for large-scale models.
To elucidate allometric scaling in complex systems, we investigated the underlying scaling relationships between typical three-scale indicators for approximately 500,000 Japanese firms; namely, annual sales, number of employees, and number of business partners. First, new scaling relations including the distributions o…
Study calculates tail risk for various mixture distributions.
problem Estimating tail risk for complex distribution mixtures.
method Analyzes tail conditional expectation for location-scale mixtures of elliptical distributions.
result Developed methods for calculating tail risk in various distributions.
A new robust scaling approach improves downstream metabolomics analysis.
problem Challenges in choosing scaling techniques for metabolomics data.
method Introduces a weighted scaling approach robust to outliers.
result The proposed method outperforms traditional scaling techniques in both outlier-free and outlier-present datasets.
Novel neural network solves PDEs with multi-scale resolution.
problem Solving time-dependent PDEs with varying spatial and temporal scales.
method Multi-scale message passing neural network with temporal and spatial gating modules.
result Outperforms baselines on PDEs with diverse scales.
New pruning method breaks power law scaling, potentially reducing error to exponential.
problem Improving neural network performance through scaling alone is costly.
method Developed a new data pruning metric to break power law scaling.
result Pruned datasets show better than power law scaling on various image datasets.
CrossAD detects anomalies in time series data by considering cross-scale associations and cross-window modeling.
problem Anomaly detection in time series data is challenging due to varying patterns at different scales and fixed window sizes.
method CrossAD incorporates cross-scale reconstruction and a query library to capture dynamic cross-scale associations and comprehensive context.
result CrossAD achieves state-of-the-art performance in anomaly detection across multiple real-world datasets.
ASRNN adapts scales dynamically for better sequence modeling.
problem Fixed scales in multiscale RNNs don't match temporal patterns.
method Adaptively learns and adjusts scales based on temporal contexts.
result ASRNNs yield better performances with dynamical scaling.
Develops active learning for scale-bridging simulations.
problem Quantitative predictions in nanoporous media and inertial confinement fusion.
method Active learning approach to optimize fine-scale simulations for coarse-scale hydrodynamics.
result Optimizes use of fine-scale simulations for coarse-scale predictions.
Theory explains neural network scaling with dataset and model size.
problem Neural network scaling laws with dataset and model size.
method Identified variance-limited and resolution-limited scaling behaviors.
result Four scaling regimes explained: infinite data, infinite width, resolution-limited, and large width.
We define the intrinsic scale at which a network begins to reveal its identity as the scale at which subgraphs in the network (created by a random walk) are distinguishable from similar sized subgraphs in a perturbed copy of the network. We conduct an extensive study of intrinsic scale for several networks, ranging fro…
Study shows how non-uniform scaling affects persistence diagrams.
problem Stability of persistence diagrams under non-uniform scaling.
method Explicit bounds on bottleneck distance derived for Euclidean scaling.
result Explicit bounds on the stability of persistence diagrams under non-uniform scaling.
The paper categorizes four types of scale-up: smart, dumb, forced, and fumbled.
problem Growing ventures in size and maintaining efficiency.
method Identifying modularity and speed as key factors, categorizing four types of scale-up.
result Modularity and speed are crucial for successful scale-up.
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weigh…
Measure-scaling quasi-isometries on graphs have specific scaling groups.
problem Understanding the scaling groups of graphs under quasi-isometries.
method Analyzing measure-scaling quasi-isometries on graphs and their properties.
result The scaling group of a graph is invariant under measure-scaling quasi-isometries.
Unified scaling laws reveal how model size and training time impact neural network performance.
problem Understanding how much performance improvement can be expected from scaling model size or data volume.
method Established scale-time equivalence and combined it with a linear model analysis of double descent.
result Unified theoretical scaling laws explain previously unexplained phenomena and offer a more accessible path to training large models.
TMSCD detects multi-scale communities in temporal networks automatically.
problem Discovering multi-scale communities in large, evolving networks.
method Spectral multilayer formulation of MM method with automatic parameter selection.
result Automatic detection of multi-scale communities without manual parameter selection.
This paper reviews some of the phenomenological models which have been introduced to incorporate the scaling properties of financial data. It also illustrates a microscopic model, based on heterogeneous interacting agents, which provides a possible explanation for the complex dynamics of markets' returns. Scaling and m…
Preprocessing data is an important step before any data analysis. In this paper, we focus on one particular aspect, namely scaling or normalization. We analyze various scaling methods in common use and study their effects on different statistical learning models. We will propose a new two-stage scaling method. First, w…
Optimizer choice affects neural scaling laws, changing the exponent α.
problem The exponent α in neural scaling laws L(N)∝N−α varies with the optimizer used. method Controlled random-feature regression experiments with five optimizer variants and six spectral conditions.
result Preconditioned optimizers yield steeper scaling (larger α), with the α-shift increasing across most of the tested spectral range. Model shows feature learning can improve neural scaling laws for hard tasks.
problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.
Analyzed US firm data 1970-2019, identifying scale effects and distributional forms.
problem Understanding differences between small and large firms over time.
method Examined all public US firms, used stylized facts and DLN distribution analysis.
result Small firms are systematically different from large firms, with scale-dependent heteroskedasticity.
Classical scaling is shown to be optimal under various noisy conditions.
problem Consistency of classical scaling under general noise conditions.
method Established using finite fourth moments of noise, derived convergence rates, and matching minimax lower bounds.
result Classical scaling achieves minimax optimality in recovering true configuration from noisy dissimilarities.
This work explores variably scaled kernels to improve non-stationary Gaussian processes.
problem Limited ability of stationary kernels to represent heterogeneous correlation structures.
method Introduces variably scaled kernels to modify correlation structures explicitly.
result Improved reconstruction accuracy and better uncertainty estimates for non-stationary data.
New learning dynamics achieve fast convergence in games without needing to know utility scales.
problem Fast convergence guarantees in learning games require prior knowledge of utility scales.
method Developed scale-free and scale-invariant learning dynamics using optimistic follow-the-regularized-leader with adaptive learning rates and clipping techniques.
result Achieved fast convergence rates to Nash and correlated equilibria without prior utility scale knowledge.
Paper introduces multi-scale methods to improve CATE estimation from EO data.
problem Challenges in balancing fine-grained and contextual information in EO-based causal inference.
method Multi-Scale Representation Concatenation, combining Vision Transformer and Causal Forests.
result Multi-scale approach captures effect heterogeneity better than single-scale models.
Abstract: Unifies small and large scale geometries using linear algebra concepts.
problem Tackles unification of small and large scale geometries.
method Uses analog of multilinear forms from Linear Algebra to compactify and unify various compactifications.
result Simple proofs of generalized theorems in coarse topology, including a new result about Higson coronas.
Adaptive loss scaling speeds up and improves deep learning training.
problem Numerical underflow in mixed precision training.
method Adaptive loss scaling that automatically computes layer-wise loss scale values during training.
result Adaptive loss scaling leads to shorter convergence time and improved accuracy.
New method selects diffusion scales for graph wavelets.
problem Choosing optimal diffusion scales for graph wavelets.
method Proposes an unsupervised method using information theory.
result Method selects diffusion scales for graph wavelets.
AutoScale improves LLM pre-training by adjusting data mixtures at different scales.
problem Data mixtures that work well at small scales may not perform as well at larger scales.
method AutoScale uses a two-stage approach: fitting a model to predict loss under different compositions and extrapolating optimal compositions to larger scales.
result AutoScale accelerates convergence and improves downstream performance.
This study presents a new lossy image compression method that utilizes the multi-scale features of natural images. Our model consists of two networks: multi-scale lossy autoencoder and parallel multi-scale lossless coder. The multi-scale lossy autoencoder extracts the multi-scale image features to quantized variables a…
Automated spectral clustering algorithm discovers clusters without parameter tuning.
problem Automatic spectral clustering for multi-scale data.
method Heuristic iterative eigengap search with global and local scaling.
result Discover different patterns with accuracy >90% in most cases.
SFM resolves small-scale physics challenges in weather data.
problem Challenges in super-resolving small-scale details in physical sciences like weather.
method Encoding inputs to a latent base distribution, flow matching for stochastic details, adaptive noise scaling.
result SFM framework significantly outperforms existing methods.
A new scaling calculus helps design and initialize ReLU networks more effectively.
problem Optimizing the design and initialization of ReLU neural networks.
method Proposes a scaling constant for neural network layers and weights, relating it to optimizability.
result A network with a uniform scaling constant is easier to train, and the geometric mean of fan-in and fan-out is a better initialization variance.
Theory predicts neural scaling exponents from language statistics.
problem No existing theory could quantitatively predict neural scaling exponents.
method Isolated two key statistical properties of language.
result Derives a simple formula predicting neural scaling exponents.
High-dimensional shrinkage risk depends on the default prior for the common scale.
problem Choosing the default prior for the common scale in high-dimensional shrinkage.
method Using radial-power benchmark to compare variance-flat and standard deviation-flat priors.
result The standard deviation-flat prior has a one-unit asymptotic risk advantage near the origin.