Batch normalisation doesn't affect variational inference but fails for larger batch sizes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Quantile normalisation is a popular normalisation method for data subject to unwanted variations such as images, speech, or genomic data. It applies a monotonic transformation to the feature values of each sample to ensure that after normalisation, they follow the same target distribution for each sample. Choosing a "g…
A new method normalizes EBM training by introducing a learnable parameter.
Proposes a method to apply conformal prediction to probabilistic time series forecasting models.
Bayesian methods promise to fix many shortcomings of deep learning, but they are impractical and rarely match the performance of standard methods, let alone improve them. In this paper, we demonstrate practical training of deep networks with natural-gradient variational inference. By applying techniques such as batch n…
Study of superintegrable systems linked to affine hypersurfaces.
Study of Coxeter diagrams and Artin-Tits groups, focusing on normalisers and wall intersections.
We show that normalising flows become pathological when used to model targets whose supports have complicated topologies. In this scenario, we prove that a flow must become arbitrarily numerically noninvertible in order to approximate the target closely. This result has implications for all flow-based models, and espec…
Generalisation of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named approximated orthonormal normalisation (AON), to improve the generalisation capacity of a DNN model. Considering a weight matrix W …
Kernelised flows improve density estimation and generation with fewer parameters.
In the present paper, we develop a picture formalism which gives rise to an invariant that dominates several known invariants of classical and virtual knots: the Jones polynomial, the Kuperberg bracket, and the normalised arrow polynomial.
Spectral clustering is a popular and versatile clustering method based on a relaxation of the normalised graph cut objective. Despite its popularity, however, there is no single agreed upon method for tuning the important scaling parameter, nor for determining automatically the number of clusters to extract. Popular he…
In recent times, many of the breakthroughs in various vision-related tasks have revolved around improving learning of deep models; these methods have ranged from network architectural improvements such as Residual Networks, to various forms of regularisation such as Batch Normalisation. In essence, many of these techni…
Improved normalising flows using Student's t-distribution for robust training.
Method estimates bivariate causal models using normalising flows and variational Gaussian process regression.
Squared families are a new model class derived from linear transformations, offering convenient properties and universal approximation.
Score-based methods fail with isolated components and incorrect mixing proportions.
Adapts linearised Laplace method for deep learning models.
We show uniqueness of classical solutions of the normalised two-dimensional Hamilton-Ricci flow on closed, smooth manifolds for smooth data among solutions satisfying (essentially) only a uniform bound for the Liouville energy and a natural space-time -bound for the time derivative of the solution. The result is s…
RotRNN uses rotations to simplify long sequence modelling.
Improved spectral convergence bounds for diffusion maps on tori.
In quantitative finance, we often model asset prices as a noisy Ito semimartingale. As this model is not identifiable, approximating by a time-changed Levy process can be useful for generative modelling. We give a new estimate of the normalised volatility or time change in this model, which obtains minimax convergence …
Bitcoin volatility shows multifractal structure, contradicting rough volatility models.
New algorithm STCV improves sparse model discovery from normalised data.
The paper derives inequalities for Riemannian submersions and their applications.
In 2004, Manning showed that the topological entropy of the geodesic flow for a surface of negative curvature decreases as the metric evolves under the normalised Ricci flow. It is an interesting open problem, also due to Manning, to determine to what extent such behaviour persists for higher dimensional manifolds. In …
In recent years there has been a growing interest in image generation through deep learning. While an important part of the evaluation of the generated images usually involves visual inspection, the inclusion of human perception as a factor in the training process is often overlooked. In this paper we propose an altern…
New proof of Schwarzschild stability using geometric gauge.
Convolutional DKMs improve kernel methods on MNIST, CIFAR-10, and CIFAR-100.
Gradient flow method solves isoperimetric inequality for maps.
Combines neural networks with splitting-up method for filtering equations.
Contrary to standard statistical models, unnormalised statistical models only specify the likelihood function up to a constant. While such models are natural and popular, the lack of normalisation makes inference much more difficult. Here we show that inferring the parameters of a unnormalised model on a space can …
This paper investigates methods for quantifying similarity between audio signals, specifically for the task of of cover song detection. We consider an information-theoretic approach, where we compute pairwise measures of predictability between time series. We compare discrete-valued approaches operating on quantised au…
A new score SiNNE improves OAM efficiency and accuracy.
We consider learning on graphs, guided by kernels that encode similarity between vertices. Our focus is on random walk kernels, the analogues of squared exponential kernels in Euclidean spaces. We show that on large, locally treelike, graphs these have some counter-intuitive properties, specifically in the limit of lar…
Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.
FlashIV solves Black-Scholes implied volatility efficiently and accurately.
Revisits zero modes of Dirac operator on Eguchi-Hanson space.
A new measure normalizes clustering accuracy to evaluate algorithms better.
We show how to construct unitary representations of the oriented Thompson group from oriented link invariants. In particular we show that the suitably normalised HOMFLYPT polynomial defines a positive definite function of .
NSFs learn SDE transition laws for efficient sampling.
New tractable density models from squaring neural networks.
Modeling complex conditional distributions is critical in a variety of settings. Despite a long tradition of research into conditional density estimation, current methods employ either simple parametric forms or are difficult to learn in practice. This paper employs normalising flows as a flexible likelihood model and …
Statistical physics approaches can be used to derive accurate predictions for the performance of inference methods learning from potentially noisy data, as quantified by the learning curve defined as the average error versus number of training examples. We analyse a challenging problem in the area of non-parametric inf…
Bayesian inference uses Stein discrepancy for robustness in intractable likelihoods.
We study the small-time behaviour of the rough Bergomi model, introduced by Bayer, Friz and Gatheral (2016), and prove a large deviations principle for a rescaled version of the normalised log stock price process, which then allows us to characterise the small-time behaviour of the implied volatility.
Kernel adaptive filters, a class of adaptive nonlinear time-series models, are known by their ability to learn expressive autoregressive patterns from sequential data. However, for trivial monotonic signals, they struggle to perform accurate predictions and at the same time keep computational complexity within desired …
New methods combine model predictions to avoid linear mixtures' limitations.