Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

13263851 · Jun 202019922001200920172026
48 results for dying ReLU

The paper investigates how target normalization and momentum affect dying ReLUs in neural networks.

problem Understanding and mitigating the dying ReLU problem in neural networks.
method Empirical analysis and theoretical modeling of a discrete-time linear autonomous system.
result Target variance plays a crucial role in the dying ReLU phenomenon, and momentum exacerbates this issue.

The dying ReLU refers to the problem when ReLU neurons become inactive and only output 0 for any input. There are many empirical and heuristic explanations of why ReLU neurons die. However, little is known about its theoretical analysis. In this paper, we rigorously prove that a deep ReLU network will eventually die in…

2019-03-15abs ↗pdf ↗

New activation function BrownianReLU improves LSTM network performance on financial time series.

problem Gradient instability in noisy financial time series data.
method Introduces BrownianReLU, a stochastic activation function based on Brownian motion.
result Significantly improved predictive accuracy and generalization on financial datasets.

DyS model improves survival analysis accuracy and interpretability.

problem Accurate and interpretable survival analysis models for healthcare.
method Feature-sparse Generalized Additive Model combining feature selection and interpretable prediction.
result DyS model outperforms other survival analysis models in interpretability and accuracy.

In this note, we study the integral of the 1-form logxdyylogydxx\log x\frac{dy}{y}-\log y\frac{dx}{x} over certain plane curves defined by A-polynomials of knots. It is quite surprising that a Chern-Simons type invariant of 3-manifolds, which can be geometrically computed, may be used to get the exact values of those integrals. Th…

2008-11-17abs ↗pdf ↗

We study a variant of decision-theoretic online learning in which the set of experts that are available to Learner can shrink over time. This is a restricted version of the well-studied sleeping experts problem, itself a generalization of the fundamental game of prediction with expert advice. Similar to many works in t…

2019-10-29abs ↗pdf ↗

We prove that, starting at an initial metric g(0)=e2u0(dx2+dy2)g(0)=e^{2u_0}(dx^2+dy^2) on R2\mathbb{R}^2 with bounded scalar curvature and bounded u0u_0, the Ricci flow tg(t)=Rg(t)g(t)\partial_t g(t)=-R_{g(t)}g(t) converges to a flat metric on R2\mathbb{R}^2.

2009-08-16abs ↗pdf ↗

The isotropic 3-space \mathbb{I}^{3} is a real affine 3-space endowed with the metric dx^{2}+dy^{2}. In this paper we describe Weingarten and linear Weingarten affine translation surfaces in \mathbb{I}^{3}. Further we classify the affine translation surfaces in \mathbb{I}^{3} that satisfy certain equations in terms of …

2016-11-07abs ↗pdf ↗

This study connects ReLU neural networks to toric geometry to analyze function realization.

problem Determining which continuous piecewise linear functions can be realized by ReLU neural networks.
method Established a connection between toric geometry and ReLU neural networks, defining key structures like the ReLU fan, toric variety, and Cartier divisor.
result Proved a criterion for functions realizable by unbiased shallow ReLU networks using intersection numbers.

We introduce a new Self-Organized Criticality (SOC) model for simulating price evolution in an artificial financial market, based on a multilayer network of traders. The model also implements, in a quite realistic way with respect to previous studies, the order book dy- namics, by considering two assets with variable f…

2016-06-29abs ↗pdf ↗

Optimal transport theory characterizes convex order between probability measures.

problem Characterizing convex order between probability measures using optimal transport.
method Quantitative bounds on optimal transport, infimum of functionals over 1-Lipschitz functions.
result Two measures are in convex order if and only if a specific cost functional inequality holds.

In this paper we characterize the degenerate elliptic equations F(D^2u)=0 whose viscosity subsolutions, (F(D^2u) \geq 0), satisfy the strong maximum principle. We introduce an easily computed function f(t) for t > 0, determined by F, and we show that the strong maximum principle holds depending on whether the integral …

2013-09-06abs ↗pdf ↗

Large deviation principle for deep neural networks with ReLU activation.

problem Understanding the behavior of deep neural networks with ReLU activation.
method Proving a large deviation principle for networks with Gaussian weights and ReLU activation functions.
result Simplified expressions and power-series expansions for the ReLU case.

Study shows efficient neural network approach for stochastic bandits.

problem Optimizing decisions in uncertain environments with neural network models.
method OFU-ReLU algorithm that balances exploration and exploitation, using a transformed feature space.
result Achieves ildeO(T) ilde{O}(\sqrt{T}) regret guarantee for stochastic bandits with ReLU neural networks.

The paper analyzes deep ReLU CNNs' approximation properties in 2D space.

problem Establishing L2L^2 approximation properties for deep ReLU CNNs.
method Analysis based on decomposition theorem for convolutional kernels, properties of ReLU activation, and connections with one-hidden-layer ReLU NNs.
result Universal approximation theorem for deep ReLU CNNs with classic structure.

Proves existence of optimal shallow neural networks with ReLU activation.

problem Proving the existence of optimal shallow feedforward networks with ReLU activation.
method Proves existence of global minima in the loss landscape for continuous target functions using shallow feedforward neural networks with ReLU activation.
result Existence of global minima in the loss landscape for shallow feedforward networks with ReLU activation.

We present an approach to cohomological dimension theory based on infinite symmetric products and on the general theory of dimension called the extension dimension. The notion of the extension dimension $\ExD(X)$ was introduced by A.N.Dranishnikov \cite {D5_5} in the context of compact spaces and CW complexes. This pa…

2004-04-19abs ↗pdf ↗

Bayesian ReLU nets fix asymptotic overconfidence with infinite features.

problem Bayesian ReLU nets can be asymptotically overconfident far from training data.
method Extend finite ReLU BNNs with infinite ReLU features via a Gaussian process.
result The resulting model is asymptotically maximally uncertain far from the data.

Let g=e2u(dx2+dy2)g=e^{2u}(dx^2+dy^2) be a conformal metric defined on the unit disk of C\mathbf{C}. We give an estimate of uL2,(D12)\|\nabla u\|_{L^{2,\infty}(D_\frac{1}{2})} when K(g)L1\|K(g)\|_{L^1} is small and μ(Brg(z),g)πr2<Λ\frac{μ(B_r^g(z),g)}{πr^2}<Λ for any rr and zD34z\in D_\frac{3}{4}. Then we will use this estimate to study the Gromov-Hausdor…

2019-11-07abs ↗pdf ↗

Despite their prevalence in neural networks we still lack a thorough theoretical characterization of ReLU layers. This paper aims to further our understanding of ReLU layers by studying how the activation function ReLU interacts with the linear component of the layer and what role this interaction plays in the success …

2018-12-06abs ↗pdf ↗

Gradient descent converges to minimum Bayes risk for two-layer ReLU networks in mean field regime.

problem Training two-layer ReLU networks using gradient descent in the mean field regime.
method Describes a condition for convergence to minimum Bayes risk, extending previous results to ReLU-activated networks.
result The condition for convergence does not depend on initialization and concerns weak convergence of network realization.

PHP connects to ReLU neural networks for scalable Bayesian inference.

problem Scalability and Bayesian inference in two-layer ReLU neural networks.
method PHP with Gaussian prior, decomposition propositions, annealed sequential Monte Carlo.
result PHP provides an alternative scalable representation for two-layer ReLU neural networks.

We study the approximation properties of random ReLU features through their reproducing kernel Hilbert space (RKHS). We first prove a universality theorem for the RKHS induced by random features whose feature maps are of the form of nodes in neural networks. The universality result implies that the random ReLU features…

2018-10-10abs ↗pdf ↗

Paper analyzes GLM-tron for high-dimensional ReLU regression, providing upper and lower bounds.

problem Learning a single ReLU neuron in high-dimensional settings with overparameterization.
method Perceptron-type algorithm GLM-tron, with finite-sample analysis.
result Sharp characterization of high-dimensional ReLU regression problems via GLM-tron, contrasting with SGD.

The paper explores how ReLU DNNs can represent MPC policies and vice versa.

problem Representing MPC policies as ReLU DNNs and vice versa.
method Developed an approximate method for identifying input-space in ReLU nets resulting in PWA functions over polyhedral regions. Studied inverse multiparametric linear or quadratic programs for reconstruction of constraints and cost functions given a PWA function.
result Identification and representation of MPC policies as ReLU DNNs and vice versa.

Study measures risk spillovers between US and China's agricultural futures markets.

problem Interconnectedness and risk transmission in agricultural futures markets.
method TVP-VAR-DY model with quantile method.
result CBOT corn, soybean, and wheat are primary risk transmitters; DCE corn and soybean are main receivers.

We study algebraic varieties of ReLU networks to understand their representable functions.

problem Understanding the functions that ReLU neural networks can represent.
method We introduce algebraic varieties associated with ReLU networks and derive polynomial equations to characterize representable functions.
result Conditions under which ReLU networks attain their expected dimension, providing insight into their structural properties.

Using the large deviation principle (LDP) for a re-scaled fractional Brownian motion BtHB^H_t where the rate function is defined via the reproducing kernel Hilbert space, we compute small-time asymptotics for a correlated fractional stochastic volatility model of the form $dS_t=S_tσ(Y_t) (\barρ dW_t +ρdB_t), \,dY_t=dB^H…

2016-10-27abs ↗pdf ↗

Gradient descent with logistic loss can interpolate deep networks with smoothed ReLU activations under certain conditions.

problem Conditions for gradient descent to drive logistic loss to zero in deep networks with smoothed ReLU activations.
method Gradient descent applied to fixed-width deep networks with smoothed ReLU approximations (e.g., Swish, Huberized ReLU).
result Gradient descent can drive logistic loss to zero under specific conditions, providing bounds on convergence rate.

The paper calculates upper bounds on ReLU network Lipschitz constants.

problem Determining the maximum perturbation size for robustness of neural networks.
method Analyzing ReLU, affine-ReLU, and max pooling functions; combining results; tracking zero elements; using a computational approach.
result The method produces the largest known bounds on minimum adversarial perturbations for large networks.

Study approximates nonlinear functionals using deep ReLU networks.

problem Approximating nonlinear continuous functionals with neural networks.
method Constructs continuous piecewise linear interpolation under simple triangulation, analyzes rates of approximation.
result Established rates of approximation for functional deep ReLU networks.