Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,878 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Jun 199319922001200920172026
48 results for generalisation error

Theory for RLHF generalization under reward shift and clipped KL.

problem Theoretical understanding of RLHF generalization, especially with reward shift and clipped KL.
method Developed generalization theory for RLHF, accounting for reward shift and clipped KL.
result Presented generalization bounds for RLHF, suggesting generalization error from sampling, reward shift, and KL clipping.

Graph neural networks generalize well under certain conditions, explained by learning theory.

problem Understanding why graph neural networks generalize well in transductive inference.
method Analysis of transductive Rademacher complexity to explain generalization properties of graph convolutional networks.
result Transductive Rademacher complexity can explain the generalization of graph convolutional networks for node classification in stochastic block models.

Study on SGD dynamics in neural networks, revealing generalisation patterns.

problem Understanding generalisation in over-parameterised neural networks.
method Analysis of SGD dynamics in a teacher-student setup using differential equations.
result Network size affects generalisation error differently depending on training layers and activation functions.

This work improves generalisation bounds using chaining and information theory.

problem Improving generalisation bounds for supervised learning algorithms.
method Developed a theoretical framework linking generalisation bounds to their chained counterparts, derived new bounds using Wasserstein distance.
result Chained generalisation bounds can be tighter than standard bounds, especially for concentrated hypothesis distributions.

We provide bounds on control learning error in stochastic systems.

problem Learning optimal controls in stochastic environments with uncontrolled parts.
method Dynamic programming and mean-field interpretation of neural networks.
result Non-asymptotic bounds on generalization error for stable overparametrised settings.

Study on infinitely-wide CNNs and their adaptability to function spatial scales.

problem Understanding how CNNs efficiently learn high-dimensional functions and their adaptability to function spatial scales.
method Study infinitely-wide deep CNNs in the kernel regime, characterizing their spectrum and using generalisation bounds to prove adaptability.
result Deep CNNs adapt to the spatial scale of the target function, with error decay controlled by the effective dimensionality of function subsets.

Study reveals phase transition in neural networks near interpolation.

problem Understanding generalization and learning transitions in neural networks.
method Effective theory for approximating Bayes-optimal generalisation error.
result Unveils a discontinuous phase transition between universal and specialisation phases.

Study on generalisation in random feature learning and hidden manifold models.

problem Generalisation in high-dimensional learning problems.
method Replica method from statistical physics for asymptotic generalisation performance.
result Closed-form expression for generalisation performance in various high-dimensional settings.

Improved prediction of soil parameters using Multi-target Stacked Generalisation on EDXRF spectra.

problem Challenges in predicting multiple soil parameters accurately from EDXRF spectra.
method Multi-target Stacked Generalisation (MTSG) method combining multiple regression models.
result MTSG significantly improved prediction accuracy for multiple soil parameters, reducing average error from 0.67 to 0.64.

Estimates and convergence of neural network approximations without structural assumptions.

problem Estimating and understanding the generalization error of neural networks.
method Introducing a new approach to estimate and analyze the convergence of neural network approximations.
result Estimates of the error without structural assumptions and convergence under mild regularity assumptions.

The paper improves transformer generalization bounds using rank-dependent covering number bounds.

problem Improving generalization bounds for transformers.
method Introducing rank-dependent covering number bounds for linear function classes and applying them to transformers.
result Generalization error bounds for transformers decay as O(1/n)O(1/\sqrt{n}) and O(logrw)O(\log r_w), improving existing bounds.

The study analyzes multi-class teacher-student perceptron performance and generalization errors.

problem Analyzing multi-class classification with the teacher-student perceptron.
method Deriving asymptotic expressions for Bayes-optimal and empirical risk minimization (ERM) generalization errors.
result Regularised cross-entropy minimization yields close-to-optimal accuracy for multi-class classification.

New approach to compute generalization performance using known risk distribution.

problem Computing generalization performance in machine learning.
method Assumes known risk distribution ρ(r)ρ(r), computes expected error using empirical risk minimization, and considers power-law behavior of ρ(r)ρ(r).
result Corrected typical behavior of generalization performance due to chance correlations in training set.

ARC algorithm optimizes dynamic pricing with correlated observations.

problem Optimizing dynamic pricing with correlated and generally distributed observations.
method Extends ARC algorithm to batched bandits with generalised linear model.
result ARC algorithm outperforms alternative approaches in dynamic pricing.

SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.

problem Understanding why overparameterized deep nets generalize well despite fitting training data perfectly.
method Empirical investigation of PSGD(fS)P_{SGD}(f\mid S) and PB(fS)P_B(f\mid S) for various architectures and datasets.
result The probability of SGD-converging on a function consistent with training data correlates well with the Bayesian posterior probability of expressing that function.

Paper connects neural networks to Gaussian processes for understanding double-descent.

problem Understanding the double-descent phenomenon in neural networks.
method Uses techniques from random matrix theory and Gaussian processes.
result Establishes a connection between NNGP and random matrix theory for neural networks.

Improved similarity search in embeddings using InfoNCE loss.

problem Improving similarity search in embedding models trained by contrastive learning.
method Introduced a new continuity bound for InfoNCE loss via Gâteaux differentiation, preserving the averaging effect of negative samples.
result Demonstrated that the averaging effect of kk negative samples in InfoNCE loss carries over to stabilisation of generalisation error as kk grows.

Double descent observed in tree-based models for genomic prediction.

problem Understanding the generalization behavior of tree-based models in machine learning.
method Systematic variation of model complexity in a genomic prediction task using whole-genome sequencing data.
result Double descent emerges only when complexity is scaled jointly across learner capacity and ensemble size.

The paper generalizes Cartan Geometry using Polacek and Siegel's approach.

problem Formulating sigma model dynamics in a covariant way.
method Using Polacek and Siegel's generalised curvature and torsion approach within the generalised metric formalism.
result Almost all higher generalised tensors correspond to covariant derivatives of the generalised Riemann tensor.

The paper analyzes fluctuations in ensemble models in high-dimensional settings.

problem Understanding statistical fluctuations in ensemble models in high-dimensional settings.
method Develops a rigorous theory for the study of fluctuations in ensemble of generalised linear models.
result Provides a complete description of the asymptotic joint distribution of the empirical risk minimizer for convex losses in high-dimensional settings.

The paper applies generalised geometry to semi-Riemannian immersions and hypersurfaces.

problem Analyzing semi-Riemannian immersions and hypersurfaces using generalised geometry.
method Develops the pullback of generalised metrics and divergence operators, introduces generalised exterior curvature, and derives Gauß-Codazzi equations.
result Establishes the constraint equations for the initial value formulation of the generalised Einstein equations.

We address the problem of speech enhancement generalisation to unseen environments by performing two manipulations. First, we embed an additional recording from the environment alone, and use this embedding to alter activations in the main enhancement subnetwork. Second, we scale the number of noise environments presen…

2018-10-26abs ↗pdf ↗

No real-world reward function is perfect. Sensory errors and software bugs may result in RL agents observing higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formal…

2017-05-23abs ↗pdf ↗

Stochastic RNNs classify biological neural network paths with robust error bounds.

problem Classifying biological neural network paths.
method Modelled as a continuous-time stochastic recurrent neural network (RNN) with identity activation function, analysed in the robust regime.
result Generalisation error bound holds with high probability, showing the empirical risk minimiser is the best-in-class hypothesis.

Defines T-duality and generalised Ricci flow relations using Courant algebroid relations.

problem Establishing compatibility between T-duality and generalised Ricci flow.
method Introducing Courant algebroid relations, invariant divergence operators, and generalised isometries.
result T-duality is compatible with generalised Ricci flow, and T-dual solutions are also solutions of generalised Ricci flow.

We define (p,q)(p,q) hermitian geometry as the target space geometry of the two dimensional (p,q)(p,q) supersymmetric sigma model. This includes generalised Kähler geometry for (2,2)(2,2), generalised hyperkähler geometry for (4,2)(4,2), strong Kähler with torsion geometry for (2,1)(2,1) and strong hyperkähler with torsion geometry f…

2018-10-15abs ↗pdf ↗

We present and analyse three online algorithms for learning in discrete Hidden Markov Models (HMMs) and compare them with the Baldi-Chauvin Algorithm. Using the Kullback-Leibler divergence as a measure of generalisation error we draw learning curves in simplified situations. The performance for learning drifting concep…

2007-08-17abs ↗pdf ↗