A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Graph neural networks generalize well under certain conditions, explained by learning theory.
problem Understanding why graph neural networks generalize well in transductive inference.
method Analysis of transductive Rademacher complexity to explain generalization properties of graph convolutional networks.
result Transductive Rademacher complexity can explain the generalization of graph convolutional networks for node classification in stochastic block models.
Deep neural networks achieve stellar generalisation on a variety of problems, despite often being large enough to easily fit all their training data. Here we study the generalisation dynamics of two-layer neural networks in a teacher-student setup, where one network, the student, is trained using stochastic gradient de…
SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.
problem Understanding why overparameterized deep nets generalize well despite fitting training data perfectly.
method Empirical investigation of PSGD(f∣S) and PB(f∣S) for various architectures and datasets.
result The probability of SGD-converging on a function consistent with training data correlates well with the Bayesian posterior probability of expressing that function.
For information retrieval and binary classification, we show that precision at the top (or precision at k) and recall at the top (or recall at k) are maximised by thresholding the posterior probability of the positive class. This finding is a consequence of a result on constrained minimisation of the cost-sensitive exp…
The paper analyzes fluctuations in ensemble models in high-dimensional settings.
problem Understanding statistical fluctuations in ensemble models in high-dimensional settings.
method Develops a rigorous theory for the study of fluctuations in ensemble of generalised linear models.
result Provides a complete description of the asymptotic joint distribution of the empirical risk minimizer for convex losses in high-dimensional settings.
The paper applies generalised geometry to semi-Riemannian immersions and hypersurfaces.
problem Analyzing semi-Riemannian immersions and hypersurfaces using generalised geometry.
method Develops the pullback of generalised metrics and divergence operators, introduces generalised exterior curvature, and derives Gauß-Codazzi equations.
result Establishes the constraint equations for the initial value formulation of the generalised Einstein equations.
Spectral embedding is a procedure which can be used to obtain vector representations of the nodes of a graph. This paper proposes a generalisation of the latent position network model known as the random dot product graph, to allow interpretation of those vector representations as latent position estimates. The general…
We address the problem of speech enhancement generalisation to unseen environments by performing two manipulations. First, we embed an additional recording from the environment alone, and use this embedding to alter activations in the main enhancement subnetwork. Second, we scale the number of noise environments presen…
No real-world reward function is perfect. Sensory errors and software bugs may result in RL agents observing higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formal…
We define (p,q) hermitian geometry as the target space geometry of the two dimensional (p,q) supersymmetric sigma model. This includes generalised Kähler geometry for (2,2), generalised hyperkähler geometry for (4,2), strong Kähler with torsion geometry for (2,1) and strong hyperkähler with torsion geometry f…
We present and analyse three online algorithms for learning in discrete Hidden Markov Models (HMMs) and compare them with the Baldi-Chauvin Algorithm. Using the Kullback-Leibler divergence as a measure of generalisation error we draw learning curves in simplified situations. The performance for learning drifting concep…