This work interprets GELU and related activations via a first-order loss function.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study assesses neural nets for optimization problems, highlighting SiLU's effectiveness.
Study learns a neuron with non-monotonic activation functions.
Deep networks can approximate various activation functions with modest adjustments.
This paper studies activation sparsity in large language models, finding key trends and implications.
AGGLIO optimizes non-convex functions with local convexity guarantees.
Attention-based GNNs can't prevent oversmoothing, leading to homogeneous node representations.
Improved robustness for deep neural networks with tighter bounds and attacks.
Temporal Functional Circuits explain KAN forecasts with interpretable edge functions.
Quantized neural networks can represent all fixed-point functions under certain conditions.
This paper explores adversarial training limits and improves model robustness against norm-bounded perturbations.
Unified framework proves neural networks' ability to mimic complex tasks.