Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,236 papers · 148 categories

Trend · papers per month

199398597796 · Jun 202019922001200920182026
48 results for saturated loss function

Exact minimization of saturated loss functions for robust regression and subspace estimation.

problem Minimizing saturated loss functions for robust regression and subspace estimation.
method Developed an exact algorithm with polynomial time-complexity for robust regression and subspace estimation, relating the problems to linear model approximation.
result Exact minimization of saturated loss functions for robust regression and subspace estimation is possible with polynomial time-complexity.

GANs can generate realistic data without minimizing a divergence, contrary to current theory.

problem Current theory suggests GANs minimize a divergence to generate realistic data.
method Discussed various loss functions for G, showing they are not divergences and do not have the same equilibrium.
result GANs can use a wide range of loss functions, not just divergences, to generate realistic data.

New method recovers signals from saturated data using linear loss and nonconvex penalties.

problem Signal recovery from saturated measurements with sign information loss.
method Linear loss and nonconvex penalties (e.g., minimax concave penalty, sorted ℓ1 norm).
result Estimation error is bounded and recovery performance improved.

Extends return extrapolation to nonlinear, asymmetric functions under stochastic volatility.

problem Behavioral anomalies in portfolio choice under stochastic volatility.
method Smooth, nonlinear, asymmetric extrapolation function; CRRA investor; Heston stochastic volatility; Hamilton-Jacobi-Bellman equation; Numerical solutions (finite-difference ADI, deep learning-driven iterative).
result Saturation acts as an endogenous correction mechanism, reducing welfare loss.

We extend return extrapolation to incorporate asymmetry and saturation, finding that asymmetric nonlinear extrapolation leads to lower welfare loss.

problem Optimal portfolio choice under stochastic volatility
method Smooth, nonlinear extrapolation function with sentiment and variance hedging
result Lower welfare loss with asymmetric nonlinear extrapolation

Study network equilibria in saturated systems, revealing how small shocks can trigger major losses.

problem Understanding how small shocks can lead to major losses in financial networks and games.
method Derived explicit expressions for network equilibria, proved conditions for their uniqueness, and analyzed discontinuities.
result Bifurcation phenomenon in network equilibria, showing sensitivity to small shocks.

Neural marked point processes show saturation with complexity, leading to new simple architectures.

problem Performance saturation in neural marked point processes with complex architectures.
method Proposed GCHP with graph convolutional layers and likelihood ratio loss.
result GCHP reduces training time and improves model performance.

We extend the adaptive regression spline model by incorporating saturation, the natural requirement that a function extend as a constant outside a certain range. We fit saturating splines to data using a convex optimization problem over a space of measures, which we solve using an efficient algorithm based on the condi…

2016-09-21abs ↗pdf ↗

A new recurrent unit alleviates vanishing gradients for long-term dependencies.

problem Vanishing gradients in recurrent neural networks make long-term dependencies hard to model.
method Proposes a new NRU architecture that avoids saturating activation functions and gates.
result Demonstrates superior performance across various tasks with and without long-term dependencies.

Deep neural network models for efficient uncertainty quantification in multiphase flow.

problem Uncertainty quantification of dynamic multiphase flow in heterogeneous media due to high dimensionality and discontinuities.
method Convolutional encoder-decoder neural network for image-to-image regression, incorporating time as an input.
result Accurate surrogate model capable of characterizing spatio-temporal pressure and saturation fields with limited training data.

Common nonlinear activation functions used in neural networks can cause training difficulties due to the saturation behavior of the activation function, which may hide dependencies that are not visible to vanilla-SGD (using first order gradients only). Gating mechanisms that use softly saturating activation functions t…

2016-03-01abs ↗pdf ↗

DeepCausalMMM models marketing impacts using deep learning and causal inference.

problem Traditional MMM approaches struggle with non-linear dynamics and temporal patterns.
method Combines deep learning, causal inference, and marketing science. Uses GRUs for temporal patterns and DAG structure for channel dependencies.
result Captures non-linear dynamics and temporal patterns in marketing impacts.

EoS selectively shapes learning, affecting some groups more than others.

problem EoS affects learning differently across the data distribution.
method Branching intervention to enter or exit EoS regime, controlled perturbation to isolate mechanisms.
result EoS redistributes learning, amplifying progress on some groups and suppressing others.

This paper explores saturation effects in spectral algorithms over large dimensions.

problem Saturation effects in spectral algorithms over large dimensions.
method Improved minimax lower bound and gradient flow with early stopping strategy.
result Exact convergence rates of spectral algorithms in large dimensional settings.

Noiseless KRR achieves optimal rates and exhibits saturation effects.

problem Understanding optimal rates and saturation phenomena in noiseless kernel ridge regression.
method Comprehensive study of noiseless KRR, establishing minimax optimal rates and uncovering phenomena of extra-smoothness and saturation.
result Noiseless KRR achieves minimax optimal rates and exhibits saturation effects.

Proposes a new method for generating random parameters in neural networks.

problem Improving randomized learning of feedforward neural networks.
method Randomly selects slope angles, rotates activation functions, and distributes them across the input space.
result The method gives better results than the common approach, especially for complex target functions.

A new loss function improves uncertainty estimation in neural networks.

problem Uncertainty quantification in neural networks, especially for regression tasks.
method Second-moment loss (SML) to optimize model variance alongside mean prediction.
result SML leads to comparable prediction accuracies and uncertainty estimates with a single model.

Develops ADMM for deep neural networks with sigmoid activations to avoid saturation and improve approximation.

problem Gradient saturation in deep neural networks with sigmoid activations.
method Introduces sigmoid-ADMM pair for training deep sigmoid nets and proves its convergence.
result ADMM avoids saturation and improves approximation of deep sigmoid nets compared to ReLU nets.

This work shows MLPs can approximate monotonic functions without bounded activations.

problem Optimizing MLPs with monotonic constraints and bounded activations.
method Generalized theoretical results showing MLPs with non-negative weights and saturating activations are universal approximators.
result MLPs with non-negative weights and saturating activations are universal approximators for monotonic functions.

The study analyzes spectral algorithms for kernel methods and derives generalization error.

problem Estimating generalization error of spectral algorithms for kernel methods.
method Considered spectral algorithms including KRR and GD, derived generalization error as a functional of learning profile.
result Showed the loss localizes on certain spectral scales and conjectured universality of the loss for noisy observations.

A homogeneously saturated equation for the time development of the price of a financial asset is presented and investigated for the pricing of European call options using noise that is distributed as a Student's t-distribution. In the limit that the saturation parameter of the equation equals zero, the standard model o…

2013-01-24abs ↗pdf ↗

Contextual PDA improves explanation of image classifications for saturated models.

problem Difficulty in explaining decisions of saturated classifiers.
method Proposes Contextual PDA, a faster method for explaining image classifications.
result Contextual PDA outperforms PDA in explaining image classifications of state-of-the-art deep networks.

Learning capacity measures model complexity, correlating with test loss and sample size.

problem Understanding model complexity and its relation to test performance.
method Formal correspondence between thermodynamics and inference; learning capacity as a measure of effective dimensionality.
result Learning capacity correlates with test loss and is a small fraction of model parameters.

This study examines how reward scaling impacts non-saturating ReLU networks in reinforcement learning.

problem The impact of reward scaling on non-saturating ReLU networks in reinforcement learning.
method Proposes an Adaptive Network Scaling framework to find a suitable reward scale during learning.
result Empirical studies justify the effectiveness of the Adaptive Network Scaling framework.

The paper computes presentations of cluster modular groups and verifies their generation by Dehn twists.

problem Computing presentations and verifying generation of cluster modular groups.
method A method to compute presentations of saturated cluster modular groups and verification of generation by cluster Dehn twists.
result The cluster modular groups of specified types are virtually generated by cluster Dehn twists.

A new method detects and compacts saturated entries in antisparse coding.

problem Efficiently solving antisparse coding problems with \ell_\infty-norm penalties.
method Safe squeezing methodology to detect and compact saturated entries, reducing problem dimensionality.
result The method accelerates the computation of antisparse representation by detecting and compacting saturated entries.

Entrocraft addresses RL performance saturation in LLMs by customizing entropy curves.

problem Performance saturation in RL algorithms for LLMs.
method Entrocraft uses rejection sampling to bias advantage distributions for customized entropy schedules.
result Entrocraft significantly improves generalization, output diversity, and long-term training in 4B models.

Learning shrinks hard tail, improving inference performance.

problem Improving inference performance in neural networks.
method Latent Instance Difficulty (LID) model analyzing fine-tuning of neural networks.
result Training-dependent inference scaling, with βexteffβ_ ext{eff} growing with sample size before saturating.

Proposes a method to improve pWCET estimation for heavy-tailed distributions.

problem Improving pWCET estimation for heavy-tailed distributions in real-time systems.
method Incorporates saturating functions into Chebyshev's inequality to mitigate the influence of large outliers.
result Achieves safe and tighter bounds for heavy-tailed distributions.

New insights explain speedup saturation in distributed learning with large batches and delays.

problem Understanding and optimizing speedup in distributed learning with large batches and delays.
method Theoretical analysis of strongly convex, convex, and non-convex settings, considering data sparsity.
result Identification of a data-dependent parameter explaining speedup saturation in both batch size and gradient staleness.