Sigmoid autoencoders can implement associative memory with certain conditions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes HTAF for stable training of binary neural networks.
Training feedforward neural networks with standard logistic activations is considered difficult because of the intrinsic properties of these sigmoidal functions. This work aims at showing that these networks can be trained to achieve generalization performance comparable to those based on hyperbolic tangent activations…
We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understanding the rôle played by piecewise affine and convex nonlinearities like the ReLU and absolute value …
The Rectified Linear Unit (ReLU) is a foundational activation function in artficial neural networks. Recent literature frequently misattributes its origin to the 2018 (initial) version of this paper, which exclusively investigated ReLU at the classification layer. This paper formally corrects the citation record by tra…
ReLU is widely seen as the default choice for activation functions in neural networks. However, there are cases where more complicated functions are required. In particular, recurrent neural networks (such as LSTMs) make extensive use of both hyperbolic tangent and sigmoid functions. These functions are expensive to co…
Study shows how certain foliations in unit tangent bundles behave.
We equip the whole tangent space to a hyperbolic manifold (of constant sectional curvature -1) with a natural metric in an intrinsic way, so that the isometries of extend to isometries of by holomorphic continuation. The image to the tangent space to a geodesic is equivalent to a hyperbolic disk. In t…
Researchers create metrics on hyperbolic space's tangent bundle.
We show, using two different approaches, that there exists a family of Riemannian metrics on the tangent bundle of a two-sphere, which induces metrics of constant curvature on its unit tangent bundle. In other words, given such a metric on the tangent bundle of a two-sphere, the Hopf map is identified with a Riemannian…
Improved logistic MoE with sigmoid gate shows better sample efficiency.
In this paper, we develop an alternating direction method of multipliers (ADMM) for deep neural networks training with sigmoid-type activation functions (called \textit{sigmoid-ADMM pair}), mainly motivated by the gradient-free nature of ADMM in avoiding the saturation of sigmoid-type activations and the advantages of …
Abstract Neural Networks (ANNs) improve DNN verification efficiency.
Gating is a key technique used for integrating information from multiple sources by long short-term memory (LSTM) models and has recently also been applied to other models such as the highway network. Although gating is powerful, it is rather expensive in terms of both computation and storage as each gating unit uses a…
Study extends GNN VC dimension bounds to Pfaffian activation functions.
Projective manifolds with specific bundles are isomorphic to simpler spaces.
Combining M-algebra and hyperbolic involutory algebra extends exceptional tangent spaces to 11 dimensions.
Sigmoid-type networks avoid vanishing gradients with regularization and rescaling.
The paper finds lower bounds for volumes of complex geometric structures.
This paper is a continuation of the previous paper of the author[M]. We show that an affine deformation space of a hyperbolic surface of type (g,b) can be parametrized by Margulis invariants and affine twist parameters with a certain decomposition of the surface, which are associated with the Fenchel-Nielsen coordinate…
For an -dimensional real hyperbolic manifold , we calculate the Zariski tangent space of a character variety at Fuchisan loci to show that the tangent space consists of cubic forms. Furthermore we prove the Weil's local rigidity theorem for uniforml hyperbolic lattices using rea…
Partial coverings of hyperbolic surfaces equidistribute with geodesics.
The paper extends Descartes' circle theorem to n-flower configurations using hyperbolic geometry.
The paper examines when NTK theory applies to real finite-width neural networks.
Sigmoid gating is more sample efficient than softmax in mixture of experts.
New normalizing flows in hyperbolic space improve posterior modeling for hierarchical data.
Study shows MSE with sigmoid can match SCE in classification tasks, especially with noisy data.
New method initializes sigmoidal MLPs for interpretable shapes.
New surfaces show horocyclic flow isn't always minimal.
S-GAI initializes MLPs using spectral geometry from data, improving performance.
Deep convolutional neural networks have led to breakthrough results in numerous practical machine learning tasks such as classification of images in the ImageNet data set, control-policy-learning to play Atari games or the board game Go, and image captioning. Many of these applications first perform feature extraction …
For a compact riemannian manifold of negative curvature, the geodesic foliation of its unit tangent bundle is independent of the negatively curved metric, up to Holder bicontinuous homeomorphism. However, the riemannian metric defines a natural transverse measure to this foliation, the Liouville transverse measure, whi…
Study of harmonic maps with extreme Kerr-like singularities.
DeepSeekMoE improves language model efficiency with shared experts and normalized gating.
We study layered neural networks of rectified linear units (ReLU) in a modelling framework for stochastic training processes. The comparison with sigmoidal activation functions is in the center of interest. We compute typical learning curves for shallow networks with K hidden units in matching student teacher scenarios…
We develop a new approach to learn the parameters of regression models with hidden variables. In a nutshell, we estimate the gradient of the regression function at a set of random points, and cluster the estimated gradients. The centers of the clusters are used as estimates for the parameters of hidden units. We justif…
We consider the normalized Ricci flow evolving from an initial metric which is conformally compactifiable and asymptotically hyperbolic. We show that there is a unique evolving metric which remains in this class, and that the flow exists up to the time where the norm of the Riemann tensor diverges. Restricting to initi…
Defines horocyclic evolutes, parallels, and involutes of spacelike frontals in hyperbolic 2-space.
Long Short-Term Memory (LSTM) infers the long term dependency through a cell state maintained by the input and the forget gate structures, which models a gate output as a value in [0,1] through a sigmoid function. However, due to the graduality of the sigmoid function, the sigmoid gate is not flexible in representing m…
We show that the empirical risk minimization (ERM) problem for neural networks has no solution in general. Given a training set with corresponding responses , fitting a -layer neural network involves estimation of…
The study limits the cohomological dimension of certain affine manifolds with partially hyperbolic holonomy groups.
New examples of real hypersurfaces found in complex hyperbolic quadrics.
A new framework for hyperbolic neural networks using the Klein model is introduced.
Deep conditional generative models are developed to simultaneously learn the temporal dependencies of multiple sequences. The model is designed by introducing a three-way weight tensor to capture the multiplicative interactions between side information and sequences. The proposed model builds on the Temporal Sigmoid Be…
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
Explains exceptional surgeries connecting maps and knot orbifolds.
For a non-vanishing gradient-like vector field on a compact manifold with boundary, a discrete set of trajectories may be tangent to the boundary with reduced multiplicity , which is the maximum possible. (Among them are trajectories that are tangent to exactly times.) We prove a lower bou…