Generative models converge to data distribution but not principal latent factors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Holomorphic networks on modular arithmetic show clear success or failure, no in-between.
DFA trains deep networks by aligning weights then memorizing data.
CDC-FM improves generative model quality-generalization tradeoff by regularizing with geometry-aware noise.
FMix improves model performance by distorting learned functions without distorting data distribution.
Numeracy is the ability to understand and work with numbers. It is a necessary skill for composing and understanding documents in clinical, scientific, and other technical domains. In this paper, we explore different strategies for modelling numerals with language models, such as memorisation and digit-by-digit composi…
Connections between nodes of fully connected neural networks are usually represented by weight matrices. In this article, functional transfer matrices are introduced as alternatives to the weight matrices: Instead of using real weights, a functional transfer matrix uses real functions with trainable parameters to repre…
The majority of ML research concerns slow, statistical learning of i.i.d. samples from large, labelled datasets. Animals do not learn this way. An enviable characteristic of animal learning is `episodic' learning - the ability to memorise a specific experience as a composition of existing concepts, after just one exper…
We introduce a framework for Continual Learning (CL) based on Bayesian inference over the function space rather than the parameters of a deep neural network. This method, referred to as functional regularisation for Continual Learning, avoids forgetting a previous task by constructing and memorising an approximate post…
Variational autoencoder (VAE) is a deep generative model for unsupervised learning, allowing to encode observations into the meaningful latent space. VAE is prone to catastrophic forgetting when tasks arrive sequentially, and only the data for the current one is available. We address this problem of continual learning …
SLT reveals how grokking occurs via basin selection in training.
This work uncovers algorithm-dependent regularisation in diffusion models.
Unified normative modeling for neuroimaging phenotypes using denoising diffusion models.
QRNN uses quantum neurons to learn sequences efficiently.