Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

52105157209 · Jun 202019922001200920172026
48 results for normalisation layers

Adapts linearised Laplace method for deep learning models.

problem Incompatibility of linearised Laplace method with modern deep learning tools.
method Examines and adapts linearised Laplace method for model selection in deep learning.
result Recommendations for better adapting linearised Laplace method to modern deep learning.

Generalisation of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named approximated orthonormal normalisation (AON), to improve the generalisation capacity of a DNN model. Considering a weight matrix W …

2019-11-21abs ↗pdf ↗

Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.

problem Training deep vanilla transformers without shortcuts and normalizations.
method Parameter initializations, bias matrices, and location-dependent rescaling.
result Deep vanilla transformers can train at similar speeds and performance to standard models.

Quantile normalisation is a popular normalisation method for data subject to unwanted variations such as images, speech, or genomic data. It applies a monotonic transformation to the feature values of each sample to ensure that after normalisation, they follow the same target distribution for each sample. Choosing a "g…

2017-06-01abs ↗pdf ↗

A new method normalizes EBM training by introducing a learnable parameter.

problem Training energy-based models with maximum likelihood is challenging due to intractable normalisation constants.
method Proposes a self-normalised log-likelihood (SNL) objective that introduces a learnable parameter representing the normalisation constant.
result The SNL objective is a lower bound of the log-likelihood and can be directly optimised using stochastic gradient techniques.

Batch normalisation doesn't affect variational inference but fails for larger batch sizes.

problem Failure of Monte Carlo Batch Normalisation (MCBN) for capturing epistemic uncertainty in larger batch sizes.
method Investigated MCBN as an approximate inference technique for Bayesian neural networks, showing its limitations and providing insights for improvement.
result For larger batch sizes, MCBN fails to capture epistemic uncertainty, requiring the batch size to be a variational parameter.

Proposes a method to apply conformal prediction to probabilistic time series forecasting models.

problem Obtaining accurate prediction regions for multi-step time series forecasting with probabilistic models.
method Conformalises conditional normalising flows to generate potentially disjoint prediction regions.
result Improves predictive efficiency in time series forecasting with multimodal distributions.

A novel method to propagate uncertainty through the soft-thresholding nonlinearity is proposed in this paper. At every layer the current distribution of the target vector is represented as a spike and slab distribution, which represents the probabilities of each variable being zero, or Gaussian-distributed. Using the p…

2018-11-29abs ↗pdf ↗

Study of superintegrable systems linked to affine hypersurfaces.

problem Understanding superintegrable systems through geometric structures.
method Established a correspondence between superintegrable systems and affine hypersurfaces, defining conformal equivalence.
result Identified conformal classes of abundant manifolds with abundant hypersurface immersions.

Study of Coxeter diagrams and Artin-Tits groups, focusing on normalisers and wall intersections.

problem Understanding normalisers of parabolic subgroups in Artin-Tits groups and their connections to Coxeter diagrams.
method Analyzing hyperplane arrangements, Coxeter groups, and wall-and-chamber structures.
result Complexified hyperplane complement is a K(π,1) space for normalisers of parabolic subgroups in finite-type Coxeter diagrams.

Proposes a method to model financial returns with extreme shocks using flexible tail transformations.

problem Capturing extreme shocks in financial return data.
method Introduces a transformation layer in normalizing flows to model heavy-tailed distributions.
result Trained models can generate synthetic sets of extreme returns.

Kernelised flows improve density estimation and generation with fewer parameters.

problem Limited expressiveness of flow-based models due to invertibility constraints.
method Integrates kernels into normalising flows to enhance expressiveness and efficiency.
result Kernelised flows outperform neural network-based flows in parameter efficiency and low-data scenarios.

Improved normalising flows using Student's t-distribution for robust training.

problem Training deep probabilistic models with robust statistics.
method Propose Student's t-distribution as a robust alternative to Gaussian in normalising flows.
result Improved robustness and reduced generalization gap with Student's t-distribution.

Method estimates bivariate causal models using normalising flows and variational Gaussian process regression.

problem Lack of explainability in AI models, especially in causal mechanisms.
method Combination of normalising flows for density estimation and variational Gaussian process regression for post-nonlinear models.
result Method better explains cause-effect pairs than simple additive noise models.

Squared families are a new model class derived from linear transformations, offering convenient properties and universal approximation.

problem Developing a new class of probability models that are easier to handle and have useful properties.
method Introducing squared families as families of probability densities obtained by squaring a linear transformation of a statistic, and showing their properties and applications.
result Squared families have convenient properties and can approximate target densities well.

Score-based methods fail with isolated components and incorrect mixing proportions.

problem Score-based methods struggle with distributions having isolated components and incorrect mixing proportions.
method Score-based methods, including score matching, are used but fail in the presence of isolated components and incorrect mixing proportions.
result Score-based methods cannot discover isolated components or identify correct mixing proportions.

In quantitative finance, we often model asset prices as a noisy Ito semimartingale. As this model is not identifiable, approximating by a time-changed Levy process can be useful for generative modelling. We give a new estimate of the normalised volatility or time change in this model, which obtains minimax convergence …

2013-12-20abs ↗pdf ↗

Bitcoin volatility shows multifractal structure, contradicting rough volatility models.

problem Applying rough volatility models to Bitcoin volatility data.
method Normalised p-variation framework, multifractal Detrended Fluctuation Analysis, log-log moment scaling, wavelet leaders.
result Bitcoin volatility exhibits multifractal structure, violating rough volatility model assumptions.

The paper derives inequalities for Riemannian submersions and their applications.

problem Characterizing Casorati inequalities for Riemannian submersions.
method Algebraic and geometric analysis of Casorati inequalities for normalised scalar and Casorati curvatures.
result Characterization of equality cases for Casorati inequalities in Riemannian submersions.

Unified theory explains two failure modes of deep transformers and provides initialisation guidelines.

problem Two failure modes (rank collapse and entropy collapse) of self-attention layers in deep transformers.
method Analytical theory of signal propagation through deep transformers, using the Random Energy Model analogy.
result Simple algorithm to compute trainability diagrams for correct initialisation hyper-parameters.

New proof of Schwarzschild stability using geometric gauge.

problem Linear stability of Schwarzschild spacetime under gravitational perturbations.
method Employing a new geometric gauge and exploiting the structure of transport equations.
result Established both orbital and asymptotic stability for linearised quantities.

Transformers without skip connections collapse token representations to a single direction.

problem Rapid convergence of token representations to a single direction in self-attention-only Transformers.
method Analysis of layer normalization, residual connections, and multi-head attention mechanisms.
result Residual connections prevent rank collapse in real Transformers, while MLPs generate new feature directions.

Combines neural networks with splitting-up method for filtering equations.

problem Approximating the solution of filtering equations for signal processes.
method Combines splitting-up method with neural networks.
result Produces an approximation of the unnormalised conditional distribution.

Contrary to standard statistical models, unnormalised statistical models only specify the likelihood function up to a constant. While such models are natural and popular, the lack of normalisation makes inference much more difficult. Here we show that inferring the parameters of a unnormalised model on a space ΩΩ can …

2014-06-11abs ↗pdf ↗

Attention mechanisms in deep learning become Gaussian process-like as the number of heads increases.

problem Understanding the behavior of attention mechanisms in deep learning models.
method Extending the equivalence between wide neural networks and Gaussian processes to attention architectures.
result Multi-head attention architectures behave as Gaussian processes as the number of heads tends to infinity.