Exact minimization of saturated loss functions for robust regression and subspace estimation.
problem Minimizing saturated loss functions for robust regression and subspace estimation.
method Developed an exact algorithm with polynomial time-complexity for robust regression and subspace estimation, relating the problems to linear model approximation.
result Exact minimization of saturated loss functions for robust regression and subspace estimation is possible with polynomial time-complexity.
GANs can generate realistic data without minimizing a divergence, contrary to current theory.
problem Current theory suggests GANs minimize a divergence to generate realistic data.
method Discussed various loss functions for G, showing they are not divergences and do not have the same equilibrium.
result GANs can use a wide range of loss functions, not just divergences, to generate realistic data.
New method recovers signals from saturated data using linear loss and nonconvex penalties.
problem Signal recovery from saturated measurements with sign information loss.
method Linear loss and nonconvex penalties (e.g., minimax concave penalty, sorted ℓ1 norm).
result Estimation error is bounded and recovery performance improved.
Extends return extrapolation to nonlinear, asymmetric functions under stochastic volatility.
problem Behavioral anomalies in portfolio choice under stochastic volatility.
method Smooth, nonlinear, asymmetric extrapolation function; CRRA investor; Heston stochastic volatility; Hamilton-Jacobi-Bellman equation; Numerical solutions (finite-difference ADI, deep learning-driven iterative).
result Saturation acts as an endogenous correction mechanism, reducing welfare loss.
We extend return extrapolation to incorporate asymmetry and saturation, finding that asymmetric nonlinear extrapolation leads to lower welfare loss.
problem Optimal portfolio choice under stochastic volatility
method Smooth, nonlinear extrapolation function with sentiment and variance hedging
result Lower welfare loss with asymmetric nonlinear extrapolation
Study network equilibria in saturated systems, revealing how small shocks can trigger major losses.
problem Understanding how small shocks can lead to major losses in financial networks and games.
method Derived explicit expressions for network equilibria, proved conditions for their uniqueness, and analyzed discontinuities.
result Bifurcation phenomenon in network equilibria, showing sensitivity to small shocks.
Neural marked point processes show saturation with complexity, leading to new simple architectures.
problem Performance saturation in neural marked point processes with complex architectures.
method Proposed GCHP with graph convolutional layers and likelihood ratio loss.
result GCHP reduces training time and improves model performance.
A method to analyze neural network performance by measuring layer saturation.
problem Understanding which layers contribute to network performance.
method Layer saturation method: restricts layer output to eigenspace of variance matrix.
result Layer saturation indicates which layers contribute to network performance.
Paper proves KRR saturation effect for smooth functions.
problem Kernel ridge regression fails to reach theoretical limits for smooth functions.
method Proof of conjectured saturation lower bound for KRR.
result Proved the conjectured saturation lower bound for KRR.
We extend the adaptive regression spline model by incorporating saturation, the natural requirement that a function extend as a constant outside a certain range. We fit saturating splines to data using a convex optimization problem over a space of measures, which we solve using an efficient algorithm based on the condi…
A new recurrent unit alleviates vanishing gradients for long-term dependencies.
problem Vanishing gradients in recurrent neural networks make long-term dependencies hard to model.
method Proposes a new NRU architecture that avoids saturating activation functions and gates.
result Demonstrates superior performance across various tasks with and without long-term dependencies.
Deep neural network models for efficient uncertainty quantification in multiphase flow.
problem Uncertainty quantification of dynamic multiphase flow in heterogeneous media due to high dimensionality and discontinuities.
method Convolutional encoder-decoder neural network for image-to-image regression, incorporating time as an input.
result Accurate surrogate model capable of characterizing spatio-temporal pressure and saturation fields with limited training data.
Common nonlinear activation functions used in neural networks can cause training difficulties due to the saturation behavior of the activation function, which may hide dependencies that are not visible to vanilla-SGD (using first order gradients only). Gating mechanisms that use softly saturating activation functions t…
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
problem Escaping local minima in loss landscapes.
method Hill-ADAM alternates between minimizing and maximizing error to explore the loss space.
result Hill-ADAM finds the global minimum state in loss landscapes.
New methods show deep networks can compress info without saturating activations.
problem Understanding how neural networks generalize and compress information.
method Adaptive mutual information estimation techniques for neural networks.
result Compression occurs in networks with non-saturating activation functions.
Dropout improves neural networks by accelerating gradient flow.
problem Understanding why dropout works and improving neural network performance.
method Proposed an optimization technique to push input towards saturation area of activation functions.
result Gradient acceleration in activation function (GAAF) improves image classification performance.
DeepCausalMMM models marketing impacts using deep learning and causal inference.
problem Traditional MMM approaches struggle with non-linear dynamics and temporal patterns.
method Combines deep learning, causal inference, and marketing science. Uses GRUs for temporal patterns and DAG structure for channel dependencies.
result Captures non-linear dynamics and temporal patterns in marketing impacts.
EoS selectively shapes learning, affecting some groups more than others.
problem EoS affects learning differently across the data distribution.
method Branching intervention to enter or exit EoS regime, controlled perturbation to isolate mechanisms.
result EoS redistributes learning, amplifying progress on some groups and suppressing others.
This paper explores saturation effects in spectral algorithms over large dimensions.
problem Saturation effects in spectral algorithms over large dimensions.
method Improved minimax lower bound and gradient flow with early stopping strategy.
result Exact convergence rates of spectral algorithms in large dimensional settings.
A new metric measures saturation of neural network layers.
problem Analyzing the quality of latent representations in neural networks.
method Layer Saturation metric based on spectral analysis.
result Saturation is related to generalization and predictive performance.
Noiseless KRR achieves optimal rates and exhibits saturation effects.
problem Understanding optimal rates and saturation phenomena in noiseless kernel ridge regression.
method Comprehensive study of noiseless KRR, establishing minimax optimal rates and uncovering phenomena of extra-smoothness and saturation.
result Noiseless KRR achieves minimax optimal rates and exhibits saturation effects.
Proposes a new method for generating random parameters in neural networks.
problem Improving randomized learning of feedforward neural networks.
method Randomly selects slope angles, rotates activation functions, and distributes them across the input space.
result The method gives better results than the common approach, especially for complex target functions.
New insights into neural network training efficiency.
problem Understanding the optimal initialization for deep neural networks.
method Exploring the edge of chaos and saturation of tanh activation function.
result The line of uniformity in phase space intersects the edge of chaos, indicating saturation begins to hinder training efficiency.
New method resolves density ratio estimation saturation issues.
problem Error saturation in density ratio estimation methods.
method Iterated regularization to improve kernel methods.
result Achieves fast error rates on regular learning problems.
We stabilize MI-based losses by adding a regularization term, improving their performance and stability.
problem Instability of MI-based losses in machine learning.
method Added a novel regularization term to stabilize MI-based losses.
result Regularization stabilizes training and improves the performance of MI-based losses.
Develops a control framework for systemic risk under uncertainty.
problem Systemic risk under model uncertainty.
method Linear-quadratic mean-field control framework with viscosity solutions and verification theorems.
result Explicit feedback controls derived from a coupled Riccati system, preserving analytical tractability.
A new loss function improves uncertainty estimation in neural networks.
problem Uncertainty quantification in neural networks, especially for regression tasks.
method Second-moment loss (SML) to optimize model variance alongside mean prediction.
result SML leads to comparable prediction accuracies and uncertainty estimates with a single model.
Develops ADMM for deep neural networks with sigmoid activations to avoid saturation and improve approximation.
problem Gradient saturation in deep neural networks with sigmoid activations.
method Introduces sigmoid-ADMM pair for training deep sigmoid nets and proves its convergence.
result ADMM avoids saturation and improves approximation of deep sigmoid nets compared to ReLU nets.
This work shows MLPs can approximate monotonic functions without bounded activations.
problem Optimizing MLPs with monotonic constraints and bounded activations.
method Generalized theoretical results showing MLPs with non-negative weights and saturating activations are universal approximators.
result MLPs with non-negative weights and saturating activations are universal approximators for monotonic functions.
The paper normalizes Poisson saturation of coregular submanifolds.
problem Normalizing the Poisson saturation of coregular submanifolds.
method Normal form construction and Poisson geometry analysis.
result Local Poisson saturation of coregular submanifolds is an embedded Poisson submanifold with a normal form.
This work optimizes mean estimation under varying privacy constraints.
problem Mean estimation with heterogeneous privacy constraints.
method Proposes an algorithm for mean estimation under different privacy levels for users.
result Shows a saturation phenomenon in performance as privacy levels are relaxed.
The paper accelerates regression algorithms by identifying saturated coordinates.
problem Non-negative and bounded-variable linear regression problems.
method Safe screening technique to identify saturated coordinates.
result The approach provides theoretical guarantees for identifying saturated coordinates.
The study analyzes spectral algorithms for kernel methods and derives generalization error.
problem Estimating generalization error of spectral algorithms for kernel methods.
method Considered spectral algorithms including KRR and GD, derived generalization error as a functional of learning profile.
result Showed the loss localizes on certain spectral scales and conjectured universality of the loss for noisy observations.
A homogeneously saturated equation for the time development of the price of a financial asset is presented and investigated for the pricing of European call options using noise that is distributed as a Student's t-distribution. In the limit that the saturation parameter of the equation equals zero, the standard model o…
Contextual PDA improves explanation of image classifications for saturated models.
problem Difficulty in explaining decisions of saturated classifiers.
method Proposes Contextual PDA, a faster method for explaining image classifications.
result Contextual PDA outperforms PDA in explaining image classifications of state-of-the-art deep networks.
33 curves on a 3-genus surface, all intersecting at most once.
problem Finding a saturated system of curves on a surface of genus 3.
method Constructing 33 essential curves pairwise non-homotopic and intersecting at most once.
result The constructed system is saturated, not properly contained in any other system.
Learning capacity measures model complexity, correlating with test loss and sample size.
problem Understanding model complexity and its relation to test performance.
method Formal correspondence between thermodynamics and inference; learning capacity as a measure of effective dimensionality.
result Learning capacity correlates with test loss and is a small fraction of model parameters.
This study examines how reward scaling impacts non-saturating ReLU networks in reinforcement learning.
problem The impact of reward scaling on non-saturating ReLU networks in reinforcement learning.
method Proposes an Adaptive Network Scaling framework to find a suitable reward scale during learning.
result Empirical studies justify the effectiveness of the Adaptive Network Scaling framework.
Water saturation is an important property in reservoir engineering domain. Thus, satisfactory classification of water saturation from seismic attributes is beneficial for reservoir characterization. However, diverse and non-linear nature of subsurface attributes makes the classification task difficult. In this context,…
A recent paper suggests that Deep Neural Networks can be protected from gradient-based adversarial perturbations by driving the network activations into a highly saturated regime. Here we analyse such saturated networks and show that the attacks fail due to numerical limitations in the gradient computations. A simple s…
NS-GAN mode collapse due to sample weighting inversion, solved with MM-nsat.
problem Mode collapse in GANs due to sample weighting inversion.
method Preserves MM-GAN sample weighting while avoiding saturation by rescaling gradients.
result MM-nsat improves mode coverage, stability, and FID on MNIST and CIFAR-10.
The paper computes presentations of cluster modular groups and verifies their generation by Dehn twists.
problem Computing presentations and verifying generation of cluster modular groups.
method A method to compute presentations of saturated cluster modular groups and verification of generation by cluster Dehn twists.
result The cluster modular groups of specified types are virtually generated by cluster Dehn twists.
A new method detects and compacts saturated entries in antisparse coding.
problem Efficiently solving antisparse coding problems with ℓ∞-norm penalties. method Safe squeezing methodology to detect and compact saturated entries, reducing problem dimensionality.
result The method accelerates the computation of antisparse representation by detecting and compacting saturated entries.
MonoFlow rethinks GANs using Wasserstein gradient flows.
problem Inconsistencies between GAN theory and practice.
method Unified generative modeling framework based on Wasserstein gradient flows.
result Adversarial training can be seen as particle flow optimization.
Entrocraft addresses RL performance saturation in LLMs by customizing entropy curves.
problem Performance saturation in RL algorithms for LLMs.
method Entrocraft uses rejection sampling to bias advantage distributions for customized entropy schedules.
result Entrocraft significantly improves generalization, output diversity, and long-term training in 4B models.
Learning shrinks hard tail, improving inference performance.
problem Improving inference performance in neural networks.
method Latent Instance Difficulty (LID) model analyzing fine-tuning of neural networks.
result Training-dependent inference scaling, with βexteff growing with sample size before saturating. Proposes a method to improve pWCET estimation for heavy-tailed distributions.
problem Improving pWCET estimation for heavy-tailed distributions in real-time systems.
method Incorporates saturating functions into Chebyshev's inequality to mitigate the influence of large outliers.
result Achieves safe and tighter bounds for heavy-tailed distributions.
New insights explain speedup saturation in distributed learning with large batches and delays.
problem Understanding and optimizing speedup in distributed learning with large batches and delays.
method Theoretical analysis of strongly convex, convex, and non-convex settings, considering data sparsity.
result Identification of a data-dependent parameter explaining speedup saturation in both batch size and gradient staleness.