A method to analyze neural network performance by measuring layer saturation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a metric, Layer Saturation, defined as the proportion of the number of eigenvalues needed to explain 99% of the variance of the latent representations, for analyzing the learned representations of neural network layers. Saturation is based on spectral analysis and can be computed efficiently, making live ana…
Entrocraft addresses RL performance saturation in LLMs by customizing entropy curves.
We extend the adaptive regression spline model by incorporating saturation, the natural requirement that a function extend as a constant outside a certain range. We fit saturating splines to data using a convex optimization problem over a space of measures, which we solve using an efficient algorithm based on the condi…
Water saturation is an important property in reservoir engineering domain. Thus, satisfactory classification of water saturation from seismic attributes is beneficial for reservoir characterization. However, diverse and non-linear nature of subsurface attributes makes the classification task difficult. In this context,…
Study shows scaling up models doesn't always improve downstream tasks.
Modelling long-term dependencies is a challenge for recurrent neural networks. This is primarily due to the fact that gradients vanish during training, as the sequence length increases. Gradients can be attenuated by transition operators and are attenuated or dropped by activation functions. Canonical architectures lik…
Paper proves KRR saturation effect for smooth functions.
As one of standard approaches to train deep neural networks, dropout has been applied to regularize large models to avoid overfitting, and the improvement in performance by dropout has been explained as avoiding co-adaptation between nodes. However, when correlations between nodes are compared after training the networ…
Neural marked point processes show saturation with complexity, leading to new simple architectures.
Sign information is the key to overcoming the inevitable saturation error in compressive sensing systems, which causes information loss and results in bias. For sparse signal recovery from saturation, we propose to use a linear loss to improve the effectiveness from existing methods that utilize hard constraints/hinge …
New method resolves density ratio estimation saturation issues.
This paper presents the development of a hybrid learning system based on Support Vector Machines (SVM), Adaptive Neuro-Fuzzy Inference System (ANFIS) and domain knowledge to solve prediction problem. The proposed two-stage Domain Knowledge based Fuzzy Information System (DKFIS) improves the prediction accuracy attained…
The paper normalizes Poisson saturation of coregular submanifolds.
The paper accelerates regression algorithms by identifying saturated coordinates.
A homogeneously saturated equation for the time development of the price of a financial asset is presented and investigated for the pricing of European call options using noise that is distributed as a Student's t-distribution. In the limit that the saturation parameter of the equation equals zero, the standard model o…
33 curves on a 3-genus surface, all intersecting at most once.
This paper explores saturation effects in spectral algorithms over large dimensions.
A recent paper suggests that Deep Neural Networks can be protected from gradient-based adversarial perturbations by driving the network activations into a highly saturated regime. Here we analyse such saturated networks and show that the attacks fail due to numerical limitations in the gradient computations. A simple s…
NS-GAN mode collapse due to sample weighting inversion, solved with MM-nsat.
New insights explain speedup saturation in distributed learning with large batches and delays.
The time development of the price of a financial asset is considered by constructing and solving Langevin equations for a homogeneously saturated model, and for comparison, for a standard model and for a logistic model. The homogeneously saturated model uses coupled rate equations for the money supply and for the price…
To improve how neural networks function it is crucial to understand their learning process. The information bottleneck theory of deep learning proposes that neural networks achieve good generalization by compressing their representations to disregard information that is not relevant to the task. However, empirical evid…
New theory explains GAN's high quality but low diversity.
A general nonlinear logistic equation has been proposed to model long-time saturation in industrial growth. An integral solution of this equation has been derived for any arbitrary degree of nonlinearity. A time scale for the onset of nonlinear saturation in industrial growth can be estimated from an equipartition cond…
DANCE improves saliency maps by adding subtle input variations.
This work optimizes mean estimation under varying privacy constraints.
We give a method to compute presentations of saturated cluster modular groups. Using this, we obtain finite presentations of the saturated cluster modular groups of finite mutation type and . We verify that the cluster modular groups of finite mutation type , , $\widetilde{E…
Common nonlinear activation functions used in neural networks can cause training difficulties due to the saturation behavior of the activation function, which may hide dependencies that are not visible to vanilla-SGD (using first order gradients only). Gating mechanisms that use softly saturating activation functions t…
New examples show deletion type admissible pairs can be rigid under rational saturation.
Much effort has been devoted to understanding the decisions of deep neural networks in recent years. A number of model-aware saliency methods were proposed to explain individual classification decisions by creating saliency maps. However, they are not applicable when the parameters and the gradients of the underlying m…
Soft-Radial Projection solves gradient saturation in constrained deep learning.
The paper extends kernel ridge regression to product kernels and reveals new convergence behaviors.
New formulations for Ricci flows without smoothness.
This work provides a thorough study on how reward scaling can affect performance of deep reinforcement learning agents. In particular, we would like to answer the question that how does reward scaling affect non-saturating ReLU networks in RL? This question matters because ReLU is one of the most effective activation f…
Surrogate strategies are used widely for uncertainty quantification of groundwater models in order to improve computational efficiency. However, their application to dynamic multiphase flow problems is hindered by the curse of dimensionality, the saturation discontinuity due to capillarity effects, and the time-depende…
Open, connected, saturated sets W without holonomy in codimension one foliations play key roles as fundamental building blocks. Here, for the case of foliated 3-manifolds, we produce a finite system of closed, convex, non-overlapping polyhedral cones in the first cohomology of W with real coefficients such that the iso…
Artificial neural network has achieved unprecedented success in the medical domain. This success depends on the availability of massive and representative datasets. However, data collection is often prevented by privacy concerns and people want to take control over their sensitive information during both training and u…
Noiseless KRR achieves optimal rates and exhibits saturation effects.
Efficient NTF algorithm for large sparse tensors.
New proof classifies orbit closures in Hodge bundle.
Investigates optimal parameter allocation in Transformers for efficiency and expressivity.
In a previous article, we defined a very flexible notion of suborbifold and characterized those suborbifolds which can arise as the images of orbifold embeddings. In particular, suborbifolds are images of orbifold embeddings precisely when they are saturated and split. This article addresses the problem of orbifold str…
We propose a novel high-performance and interpretable canonical deep tabular data learning architecture, TabNet. TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient fea…
Extends return extrapolation to nonlinear, asymmetric functions under stochastic volatility.
Estimates model performance based on compute budget and evaluates stability over time.
We consider the reconstruction problem in compressed sensing in which the observations are recorded in a finite number of bits. They may thus contain quantization errors (from being rounded to the nearest representable value) and saturation errors (from being outside the range of representable values). Our formulation …
We extend return extrapolation to incorporate asymmetry and saturation, finding that asymmetric nonlinear extrapolation leads to lower welfare loss.