New metric measures dynamical richness without relying on accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on rich regime training in deep learning, finding active parameters in bottom layers.
The paper explains the richness scale of wide neural networks.
Study on how initialization scale affects neural network training regimes.
Grokking occurs when neural networks transition from lazy to rich training dynamics, fitting initial features before generalizing.
Transformers learn rich in-context dependencies efficiently.
A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effect of finding the minimum RKHS norm solution. This stands in contrast to other studies which demonstr…
In many situations, we need to build and deploy separate models in related environments with different data qualities. For example, an environment with strong observation equipments (e.g., intensive care units) often provides high-quality multi-modal data, which are acquired from multiple sensory devices and have rich-…
CHEER boosts poor models using rich model knowledge.
This paper shows how deep neural networks can learn rich, independent features that significantly deviate from initialization.
TF-GNN simplifies graph neural networks in TensorFlow.
We propose a novel learning method for multilayered neural networks which uses feedforward supervisory signal and associates classification of a new input with that of pre-trained input. The proposed method effectively uses rich input information in the earlier layer for robust leaning and revising internal representat…
Multi-Entity Dependence Learning (MEDL) explores conditional correlations among multiple entities. The availability of rich contextual information requires a nimble learning scheme that tightly integrates with deep neural networks and has the ability to capture correlation structures among exponentially many outcomes. …
New method uses adversarial training for structural model estimation.
Proposes a nonparametric approach for inferring spike train filters.
The limits of the nuclear landscape are determined by nuclear binding energies. Beyond the proton drip lines, where the separation energy becomes negative, there is not enough binding energy to prevent protons from escaping the nucleus. Predicting properties of unstable nuclear states in the vast territory of proton em…
The recent literature on deep learning offers new tools to learn a rich probability distribution over high dimensional data such as images or sounds. In this work we investigate the possibility of learning the prior distribution over neural network parameters using such tools. Our resulting variational Bayes algorithm …
SentenceMIM learns rich latent representations for variable-length language data.
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DN…
The study analyzes transfer learning in infinite-width neural networks, improving generalization on target tasks.
Generative adversarial networks (GANs) are a powerful framework for generative tasks. However, they are difficult to train and tend to miss modes of the true data generation process. Although GANs can learn a rich representation of the covered modes of the data in their latent space, the framework misses an inverse map…
Analysis of deep neural networks under various learning rules reveals dynamics of feature and prediction learning.
Generative Neuro-Symbolic model learns from raw data with rich conceptual representations.
Study reveals how initialization scale affects training accuracy in linear networks.
Consider a Riemannian metric on two-torus. We prove that the question of existence of polynomial first integrals leads naturally to a remarkable system of quasi-linear equations which turns out to be a Rich system of conservation laws. This reduces the question of integrability to the question of existence of smooth (q…
Study explores properties of bipartite knots.
In this paper we introduce a new neural architecture for sorting unordered sequences where the correct sequence order is not easily defined but must rather be inferred from training data. We refer to this architecture as OrderNet and describe how it was constructed to be naturally permutation equivariant while still al…
Study on Neural Collapse limits in deep learning.
Efficiently extracts linear dynamics from complex observations.
Ridge regression reveals surprising high-dimensional behaviors via random matrix theory.
We present Distributed Equivalent Substitution (DES) training, a novel distributed training framework for large-scale recommender systems with dynamic sparse features. DES introduces fully synchronous training to large-scale recommendation system for the first time by reducing communication, thus making the training of…
This paper studies an intelligent ultimate technique for health-monitoring and prognostic of common rotary machine components, particularly bearings. During a run-to-failure experiment, rich unsupervised features from vibration sensory data are extracted by a trained sparse auto-encoder. Then, the correlation of the ex…
This study reveals efficient finite-difference computation for gradient regularization in deep learning.
This study shows why training Neural ODEs is hard and proposes a new method.
Adam avoids simplicity bias in neural networks, leading to better generalization.
Study binary choice with asymmetric loss, offering simple solutions.
This paper proposes new get-rich-quick schemes that involve trading in a financial security with a non-degenerate price path. For simplicity the interest rate is assumed zero. If the price path is assumed continuous, the trader can become infinitely rich immediately after it becomes non-constant (if it ever does). If i…
New method detects corporate fraud in noisy financial networks.
Intelligent agents can cope with sensory-rich environments by learning task-agnostic state abstractions. In this paper, we propose an algorithm to approximate causal states, which are the coarsest partition of the joint history of actions and observations in partially-observable Markov decision processes (POMDP). Our m…
Robust risk minimisation has several advantages: it has been studied with regards to improving the generalisation properties of models and robustness to adversarial perturbation. We bound the distributionally robust risk for a model class rich enough to include deep neural networks by a regularised empirical risk invol…
Particle filtering is a powerful approach to sequential state estimation and finds application in many domains, including robot localization, object tracking, etc. To apply particle filtering in practice, a critical challenge is to construct probabilistic system models, especially for systems with complex dynamics or r…
Loss-guided training accelerates node embedding methods on graphs.
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameter…
The study examines the balancedness of random partition models and finds the rich-get-richer characteristic is a result of model assumptions.
Compositional structures between parts and objects are inherent in natural scenes. Modeling such compositional hierarchies via unsupervised learning can bring various benefits such as interpretability and transferability, which are important in many downstream tasks. In this paper, we propose the first deep latent vari…
Pretraining models improves text classification accuracy, but diminishing returns are observed with large datasets.
We present an end-to-end trained memory system that quickly adapts to new data and generates samples like them. Inspired by Kanerva's sparse distributed memory, it has a robust distributed reading and writing mechanism. The memory is analytically tractable, which enables optimal on-line compression via a Bayesian updat…
The optimization problem behind neural networks is highly non-convex. Training with stochastic gradient descent and variants requires careful parameter tuning and provides no guarantee to achieve the global optimum. In contrast we show under quite weak assumptions on the data that a particular class of feedforward neur…