Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

78155233310 · Jun 202019922001200920172026
48 results for double scaling

We introduce a stochastic model to explain a double power-law distribution which exhibits two different Paretian behaviors in the upper and the lower tail and widely exists in social and economic systems. The model incorporates fitness consideration and noise fluctuation. We find that if the number of variables (e.g. t…

2011-03-10abs ↗pdf ↗

Improves early stopping in deep networks by adjusting stepsizes.

problem Epoch-wise double descent in deep networks.
method Analytical and empirical study of bias-variance tradeoffs in different network layers.
result Eliminating epoch-wise double descent through adjusting stepsizes of different layers improves early stopping performance.

Research provides explicit NPV expressions for double barrier strategies.

problem Calculating expected NPVs of double barrier strategies for regular diffusions.
method Explicit expression using bivariate q-scale function with perturbation technique.
result Explicit expressions for expected NPVs are derived for certain cases.

Graphs with nonnegative Bakry-Émery curvature have volume doubling and Poincaré inequalities.

problem Proving properties of graphs with specific curvature conditions.
method Graph-theoretic modified nonlinear heat-flow method, including point-mass consequences and diffusive exit-time control.
result Volume doubling and Poincaré inequalities for graphs with nonnegative Bakry-Émery curvature.

Unified scaling laws reveal how model size and training time impact neural network performance.

problem Understanding how much performance improvement can be expected from scaling model size or data volume.
method Established scale-time equivalence and combined it with a linear model analysis of double descent.
result Unified theoretical scaling laws explain previously unexplained phenomena and offer a more accessible path to training large models.

Double descent risk in L2-regularized models explained and mitigated.

problem Risk of overparameterized models in machine learning.
method Analysis of L2-regularized models, two-layer neural networks, and CNNs.
result Double descent risk in L2-regularized models can be explained and mitigated by adjusting regularization strengths.

Model shows loss curve with two distinct exponents due to sparse activations.

problem Sparse activations impact neural network scaling laws.
method Introduced a model for neural scaling laws under sparse activations, derived asymptotic population loss, and analyzed gradient-descent dynamics.
result Loss curve exhibits double-descent peak near interpolation threshold with two distinct scaling exponents.

Recent empirical and theoretical studies have shown that many learning algorithms -- from linear regression to neural networks -- can have test performance that is non-monotonic in quantities such the sample size and model size. This striking phenomenon, often referred to as "double descent", has raised questions of if…

2020-03-04abs ↗pdf ↗

Annealing Double-Head calibrates deep neural networks during training.

problem Overestimation or underestimation of predictive confidence in deep neural networks.
method An additional calibration head and Annealing technique to dynamically scale logits.
result State-of-the-art model calibration performance achieved without post-processing.

Double descent observed in tree-based models for genomic prediction.

problem Understanding the generalization behavior of tree-based models in machine learning.
method Systematic variation of model complexity in a genomic prediction task using whole-genome sequencing data.
result Double descent emerges only when complexity is scaled jointly across learner capacity and ensemble size.

We analyze the condition number of random feature matrices and prove their well-conditioned nature.

problem Understanding the condition number of random feature matrices and its impact on generalization error.
method Established concentration bounds and derived risk bounds for regression problems using random feature matrices.
result The risk associated with random feature matrices exhibits the double descent phenomenon, improving even with noise.

A new method uses deep neural networks for estimating individual treatment effects.

problem Estimating individual treatment effects in large models.
method Extended fiducial inference with Double Neural Network (Double-NN) method.
result The Double-NN method outperforms CQR in individual treatment effect estimation.

We propose a generalized double Pareto prior for Bayesian shrinkage estimation and inferences in linear models. The prior can be obtained via a scale mixture of Laplace or normal distributions, forming a bridge between the Laplace and Normal-Jeffreys' priors. While it has a spike at zero like the Laplace density, it al…

2011-04-05abs ↗pdf ↗

In recent years, a rich variety of shrinkage priors have been proposed that have great promise in addressing massive regression problems. In general, these new priors can be expressed as scale mixtures of normals, but have more complex forms and better properties than traditional Cauchy and double exponential priors. W…

2011-07-25abs ↗pdf ↗

Generative model tackles inconsistent attributes across datasets by enabling precise conditional generation.

problem Inconsistent attributes across merged datasets limit controllability in conditional generative modeling.
method Diffusion Model with Double Guidance, maintaining control over multiple conditions without joint annotations.
result Outperforms baselines in molecular and image generation tasks, aligning with target distributions and controlling missing conditions.

The study provides precise asymptotic theory for in-context learning by Transformers.

problem Understanding the sample complexity, pretraining task diversity, and context length for successful in-context learning.
method An exactly solvable model of linear regression task by linear attention, deriving sharp asymptotics.
result Double-descent learning curve with increasing pretraining examples, phase transition between low and high task diversity regimes.

In the framework of Lorentzian warped products, we study the Friedmann-Robertson-Walker cosmological model to investigate non-smooth curvatures associated with multiple discontinuities involved in the evolution of the universe. In particular we analyze non-smooth features of the spatially flat Friedmann-Robertson-Walke…

2003-08-16abs ↗pdf ↗

We study the problem of distribution to real-value regression, where one aims to regress a mapping ff that takes in a distribution input covariate PIP\in \mathcal{I} (for a non-parametric family of distributions I\mathcal{I}) and outputs a real-valued response Y=f(P)+εY=f(P) + ε. This setting was recently studied, and a "K…

2013-11-10abs ↗pdf ↗

Optimizer choice affects neural scaling laws, changing the exponent α\alpha.

problem The exponent α\alpha in neural scaling laws L(N)NαL(N) \propto N^{-\alpha} varies with the optimizer used.
method Controlled random-feature regression experiments with five optimizer variants and six spectral conditions.
result Preconditioned optimizers yield steeper scaling (larger α\alpha), with the α\alpha-shift increasing across most of the tested spectral range.

In this paper, we show the equivalence between the boundedness of the Riesz transform dΔ1/2dΔ^{-1/2} on LpL^p, p(2,p0)p\in (2,p_0), and the equality Hp=LpH^p=L^p, p(2,p0)p\in(2,p_0), in the class of manifold whose measure is doubling and for which the scaled Poincaré inequalities hold. Here, HpH^p is a Hardy space of exact 11-forms, …

2013-08-27abs ↗pdf ↗

Model shows feature learning can improve neural scaling laws for hard tasks.

problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.

Geometry of hypersurfaces defined by the relation which generalizes classical formula for free energy in terms of microstates is studied. Induced metric, Riemann curvature tensor, Gauss-Kronecker curvature and associated entropy are calculated. Special class of ideal statistical hypersurfaces is analyzed in details. No…

2016-02-25abs ↗pdf ↗

The paper analyzes the performance of random feature regression in high dimensions.

problem Understanding how model complexity and generalization depend on the number of parameters and sample size.
method Investigates random feature ridge regression (RFRR) and compares it to kernel ridge regression (KRR).
result RFRR exhibits a trade-off between approximation and generalization power, with a double descent phenomenon at a specific point.

A Klein surface is a surface with a dianalytic structure. A double of a Klein surface XX is a Klein surface YY such that there is a degree two morphism (of Klein surfaces) YXY\rightarrow X. There are many doubles of a given Klein surface and among them the so-called natural doubles which are: the complex double, the …

2014-04-03abs ↗pdf ↗

Study predicts doubling of U.S. maize insurance claims due to climate change.

problem Climate change increases U.S. maize loss probability, impacting insurance claims.
method Neural Network Monte Carlo simulations to predict crop loss metrics.
result Doubling of annual probability of maize Yield Protection insurance claims by mid-century.

We define a general notion of abstract double Lie algebroid. We show (1) that the double Lie algebroid of a double Lie groupoid is a double Lie algebroid in this sense; (2) that the double cotangent constructed from Lie algebroid structures on a vector bundle A and its dual A* is a double Lie algebroid if and only if (…

1998-08-17abs ↗pdf ↗

Neural networks generalize well despite overfitting due to high capacity.

problem Understanding why deep neural networks generalize well in overparameterized settings.
method High-dimensional asymptotic analysis of generalization under kernel regression with Neural Tangent Kernel.
result Test error exhibits non-monotonic behavior and can have additional peaks and descents in the overparameterized regime.

Study on pseudo-Riemannian metrics on Lie groups, finding new non-Einstein examples.

problem Characterizing and finding non-Einstein pseudo-Riemannian metrics on Lie groups.
method Analyzing left invariant metrics, using double extension process, and constructing examples.
result Construction of infinitely many new explicit examples of non-Einstein pseudo-Riemannian metrics on Lie groups.

We define double principal bundles (DPBs), for which the frame bundle of a double vector bundle, double Lie groups and double homogeneous spaces are basic examples. It is shown that a double vector bundle can be realized as the associated bundle of its frame bundle. Also dual structures, gauge transformations and conne…

2016-11-02abs ↗pdf ↗

This paper establishes an equivalence between transitive double Lie algebroids and core diagrams.

problem Understanding and characterizing transitive double Lie algebroids.
method Using core diagrams and equivalence of transitive core diagrams with transitive double Lie groupoids.
result Transitive double Lie algebroids are completely determined by their core diagrams.

We define an abstract notion of double Lie algebroid, which includes as particular cases: (1) the double Lie algebroid of a double Lie groupoid in the sense of the author, such as the iterated tangent bundle of an ordinary manifold, and various iterated tangent/cotangent constructions in symplectic and Poisson geometry…

2000-11-24abs ↗pdf ↗

Generatability in metric spaces studied with novel novelty parameters.

problem Understanding generatability in metric spaces with asymmetric novelty parameters.
method Introducing (ε,ε)(\varepsilon,\varepsilon')-closure dimension to characterize uniform and non-uniform generatability.
result Generatability is stable across novelty scales in doubling spaces but can be highly scale-sensitive in general metric spaces.

We develop new algebraic methods refining the Witt group of linking forms and Ranicki's torsion algebraic L-groups into double Witt groups and double L-groups. At each prime ideal of the underlying ring, our double Witt groups capture infinitely many more integral signatures of the linking form than the single Witt gro…

2015-03-24abs ↗pdf ↗