Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

1223 · Sep 202319922001200920172026
48 results for tanh

We propose K-TanH, a novel, highly accurate, hardware efficient approximation of popular activation function TanH for Deep Learning. K-TanH consists of parameterized low-precision integer operations, such as, shift and add/subtract (no floating point operation needed) where parameters are stored in very small look-up t…

2019-09-17abs ↗pdf ↗

Finite-precision learning of anh anh networks is limited by the Monte Carlo rate.

problem Learning anh anh neural networks under finite precision
method Using iterated anh anh activations to construct localized bump functions
result No adaptive randomized algorithm can achieve higher convergence rate than Monte Carlo rate in finite precision

In this paper we show all possible ramps where an object can move with constant speed under the effect of gravity and friction. The planar ramp are very easy to describe, just rotate a curve with velocity vector (tanh(as),sech(as)). Recall that tanh(as)^2+sech^2(as) = 1. Therefore, the solution of the planar constant s…

2013-05-02abs ↗pdf ↗

The Rectified Linear Unit (ReLU) is a foundational activation function in artficial neural networks. Recent literature frequently misattributes its origin to the 2018 (initial) version of this paper, which exclusively investigated ReLU at the classification layer. This paper formally corrects the citation record by tra…

2018-03-22abs ↗pdf ↗

Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.

problem Improving the out-of-sample performance of Bayesian neural networks.
method Numerical sampling of Bayesian posterior, ensembling over architectures, analysis of evidence vs. model size.
result Good correlation between out-of-sample performance and Bayesian evidence; ensembling improves performance.

New theory explains signal propagation in normalization-free transformers.

problem Understanding signal propagation in normalization-free transformers.
method Deriving recurrence relations for activation statistics and APJNs across layers.
result Transformers with elementwise tanh-like nonlinearities exhibit subcritical signal propagation.

New findings on neural network identifiability using affine symmetries.

problem Identifying all neural networks that produce a given function.
method Examined affine symmetries of nonlinearities and their impact on neural network identifiability.
result Symmetries can be used to find a rich set of networks giving rise to the same function, except for a special case.

The paper analyzes the score field of diffusion models using Burgers dynamics.

problem Understanding the evolution of score fields in diffusion models.
method Analyzes the score field through Burgers-type evolution law for diffusion models.
result Identifies a universal \( anh\) interfacial term in the score field.

New method improves training of PINNs for PDEs by adding noisy supervision terms.

problem Slow or failed convergence of PINNs on challenging PDEs.
method Operator preconditioning using Feynman-Kac supervision and non-asymptotic error bounds.
result Non-asymptotic error bounds for FK-PINNs, showing improved performance over standard PINNs.

New PINNs method improves accuracy in computing Mean Escape Time from bounded domains.

problem Computing Mean Escape Time from bounded domains with high accuracy.
method Boundary-adapted Physics-Informed Neural Networks (PINNs) with exact Dirichlet boundary enforcement.
result Derivation of H2(Ω)H^2(Ω) a priori error bounds for PINNs with normalized distance approximations.

We present a Statistical Mechanics (SM) model of deep neural networks, connecting the energy-based and the feed forward networks (FFN) approach. We infer that FFN can be understood as performing three basic steps: encoding, representation validation and propagation. From the meanfield solution of the model, we obtain a…

2018-05-22abs ↗pdf ↗

We seek to improve the data efficiency of neural networks and present novel implementations of parameterized piece-wise polynomial activation functions. The parameters are the y-coordinates of n+1 Chebyshev nodes per hidden unit and Lagrangian interpolation between the nodes produces the polynomial on [-1, 1]. We show …

2019-06-24abs ↗pdf ↗

SGD converges globally to logistic loss minima for two-layer nets.

problem Global convergence of SGD for logistic loss on two-layer neural nets.
method Demonstrates existence of Frobenius norm regularized logistic loss functions as Villani functions, proving convergence and exponential rate.
result SGD converges globally to the global minima of appropriately regularized logistic empirical risk of depth 2 nets.

Randomly initialized wide neural networks with zero-mean activations are nearly independent, potentially solving AI interpretability limits.

problem Measuring the limits of AI interpretability.
method Randomly initialized neural networks with large width and zero-mean activation functions.
result Neural networks with zero-mean activations are nearly independent, solving the computational no-coincidence conjecture.

Let H denote the standard one-point completion of a real Hilbert space. Given any non-trivial proper sub-set U of H one may define the so-called `Apollonian' metric d_U on U. When U \subset V \subset H are nested proper subsets we show that their associated Apollonian metrics satisfy the following uniform contraction p…

2011-02-21abs ↗pdf ↗

Previous work has questioned the conditions under which the decision regions of a neural network are connected and further showed the implications of the corresponding theory to the problem of adversarial manipulation of classifiers. It has been proven that for a class of activation functions including leaky ReLU, neur…

2019-01-25abs ↗pdf ↗

Deep Learning is applied to energy markets to predict extreme loads observed in energy grids. Forecasting energy loads and prices is challenging due to sharp peaks and troughs that arise due to supply and demand fluctuations from intraday system constraints. We propose deep spatio-temporal models and extreme value theo…

2018-08-16abs ↗pdf ↗

A celebrated conjecture due to De Giorgi states that any bounded solution of the equation Δu+(1u2)u=0inRNΔu + (1-u^2) u = 0 \hbox{in} \R^N with $\pp_{y_N}u >0$ must be such that its level sets $\{u=\la\}$ are all hyperplanes, {\em \bf at least} for dimension N8N\le 8. A counterexample for N9N\ge 9 has long been believed to exist. …

2008-06-19abs ↗pdf ↗

The paper identifies a new geometric and spectral phenomenon in the critical hyperbolic catenoid family.

problem The study investigates the critical hyperbolic catenoid family and its geometric and spectral properties.
method The approach involves analyzing the critical hyperbolic catenoid family, identifying parameter-criticality, and studying the Robin spectrum.
result The paper proves that at a parameter-critical value aa^\sharp, the Robin nullity of ΣaΣ_{a^\sharp} is at least 3, with an additional kernel element in mode k=0k=0.

In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pai…

2015-02-09abs ↗pdf ↗

Using back-propagation and its variants to train deep networks is often problematic for new users. Issues such as exploding gradients, vanishing gradients, and high sensitivity to weight initialization strategies often make networks difficult to train, especially when users are experimenting with new architectures. Her…

2018-03-05abs ↗pdf ↗

In this paper, we consider parameter recovery for non-overlapping convolutional neural networks (CNNs) with multiple kernels. We show that when the inputs follow Gaussian distribution and the sample size is sufficiently large, the squared loss of such CNNs is  locally strongly convex\mathit{~locally~strongly~convex} in a basin of attraction…

2017-11-08abs ↗pdf ↗

Exact bounds derived for neural network outputs with noisy inputs.

problem Bounding the output distribution of neural networks with random inputs.
method Applying ReLU NNs to derive bounds for general NNs, then using these to find exact error guarantees.
result Exact upper and lower bounds for the output distribution of neural networks with random inputs.

Paper establishes bounds for RNN-TPPs, showing four-layer networks can achieve vanishing errors.

problem Understanding theoretical limits of RNN-TPPs.
method Characterized RNN complexity, constructed neural approximations, applied truncation technique.
result Four-layer RNN-TPPs can achieve vanishing generalization errors.

Study shows DNNs can recover functions with fewer samples than model parameters at overparameterization.

problem Determining reliable function recovery in overparameterized deep neural networks.
method Introducing 'local linear recovery' (LLR) and proving upper bounds on sample sizes for recovery.
result Upper bounds on optimistic sample sizes for function recovery in overparameterized DNNs are achieved.

We propose Mish\textit{Mish}, a novel self-regularized non-monotonic activation function which can be mathematically defined as: f(x)=xtanh(softplus(x))f(x)=x\tanh(softplus(x)). As activation functions play a crucial role in the performance and training dynamics in neural networks, we validated experimentally on several well-known benchmarks…

2019-08-23abs ↗pdf ↗

A complex-valued convolutional network (convnet) implements the repeated application of the following composition of three operations, recursively applying the composition to an input vector of nonnegative real numbers: (1) convolution with complex-valued vectors followed by (2) taking the absolute value of every entry…

2015-03-11abs ↗pdf ↗

Efficiently prices American options with multiple assets using sparse grids.

problem Pricing American options with multiple underlying assets efficiently.
method Dynamic programming formulation followed by sparse grid interpolation.
result Sparse grids reduce the number of interpolation points and maintain function smoothness.