Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

239477716954 · Jun 202019922001200920182026
48 results for Sigmoid belief networks

Highly expressive directed latent variable models, such as sigmoid belief networks, are difficult to train on large datasets because exact inference in them is intractable and none of the approximate inference methods that have been applied to them scale well. We propose a fast non-iterative approximate inference metho…

2014-01-31abs ↗pdf ↗

New method reparameterizes discrete variables to reduce gradient variance.

problem Low variance gradient estimation for discrete variables in neural networks.
method Marginalizing out the variable of interest to bypass discontinuity, resulting in a new reparameterization trick.
result The new reparameterization reduces gradient variance significantly, theoretically not larger than likelihood-ratio method.

A neural network derived from first principles using MaxEnt.

problem Developing a neural network from first principles.
method Derived a neural network using the principle of Maximum Entropy, with linear dimension-reducing transformations and conditional mean estimators.
result Unified theoretical justification for activation functions like sigmoid, softplus, and relu.

Python scripts analyze MRAM-based neuromorphic devices' process variation impacts on machine learning accuracy.

problem Impact of process variation on MRAM-based neuromorphic devices' performance in machine learning applications.
method Developed transportable Python scripts to analyze output variation under changes in device dimensions.
result Revealed impacts and limits for processing variation of device fabrication on energy vs. accuracy tradeoffs.

Iterative Proportional Fitting (IPF), combined with EM, is commonly used as an algorithm for likelihood maximization in undirected graphical models. In this paper, we present two iterative algorithms that generalize upon IPF. The first one is for likelihood maximization in discrete chain factor graphs, which we define …

2012-12-12abs ↗pdf ↗

We show that deep narrow Boltzmann machines are universal approximators of probability distributions on the activities of their visible units, provided they have sufficiently many hidden layers, each containing the same number of units as the visible layer. We show that, within certain parameter domains, deep Boltzmann…

2014-11-14abs ↗pdf ↗

We study two-layer belief networks of binary random variables in which the conditional probabilities Pr[childlparents] depend monotonically on weighted sums of the parents. In large networks where exact probabilistic inference is intractable, we show how to compute upper and lower bounds on many probabilities of intere…

2013-01-30abs ↗pdf ↗

Stochastic neural networks can approximate any function, even with correlated outputs.

problem Approximating functions with stochastic outputs and correlations.
method Investigating deep sigmoid belief networks to approximate any stochastic mapping.
result Minimal number of layers and units needed for approximation.

ReLU activations lead to smoother learning curves compared to sigmoidal activations in neural networks.

problem Comparing the performance of ReLU and sigmoidal activations in neural networks.
method Analytical computation of learning curves in shallow networks with different activation functions.
result ReLU networks exhibit continuous transitions in performance, while sigmoidal networks show discontinuous transitions.

Develops ADMM for deep neural networks with sigmoid activations to avoid saturation and improve approximation.

problem Gradient saturation in deep neural networks with sigmoid activations.
method Introduces sigmoid-ADMM pair for training deep sigmoid nets and proves its convergence.
result ADMM avoids saturation and improves approximation of deep sigmoid nets compared to ReLU nets.

Study shows MSE with sigmoid can match SCE in classification tasks, especially with noisy data.

problem Inconsistent errors in neural network classification tasks.
method Introduced Output Reset algorithm to use MSE with sigmoid activation.
result MSE with sigmoid activation achieves comparable accuracy and convergence rates to Softmax Cross-Entropy, especially in noisy data scenarios.

Sigmoid autoencoders can implement associative memory with certain conditions.

problem Implementing associative memory in neural networks.
method Theoretical analysis of overparameterized sigmoid autoencoders using the NTK and iterative maps.
result Overparameterized sigmoid autoencoders can have attractors in the NTK limit, leading to associative memory.

A novel rejection sampling step improves variational inference for latent variable models.

problem High variance in gradient estimates for approximate posterior in stochastic variational inference.
method Rejection sampling to discard low-likelihood samples and a new gradient estimator.
result Improves marginal log-likelihood estimation by 3.71 nats and 0.21 nats.

Improved logistic MoE with sigmoid gate shows better sample efficiency.

problem Improving sample efficiency in logistic MoE models.
method Comprehensive analysis of multinomial logistic MoE with modified sigmoid gate, incorporating temperature parameter and using Euclidean score.
result The sigmoid gate leads to lower sample complexity than softmax gate for both parameter and expert estimation.

New method initializes sigmoidal MLPs for interpretable shapes.

problem Creating interpretable decision boundaries in neural networks.
method Introducing a geometry-aware initialization for sigmoidal multi-layer perceptrons (MLPs) using tropical geometry.
result Sigmoidal MLPs can have decision boundaries aligned with prescribed shapes at initialization.

Shallow neural nets classify objects perfectly if their distribution is linearly separable.

problem Designing efficient neural networks for classification.
method Constructed shallow sigmoid-type neural networks.
result Achieves 100% accuracy for datasets following a linear separability condition.

Deep sigmoidal networks can achieve dynamical isometry with orthogonal weight initialization, significantly speeding up learning.

problem Ensuring efficient learning in deep neural networks, especially with nonlinear activation functions.
method Employing free probability theory to compute the singular value distribution of a deep network's input-output Jacobian.
result Deep sigmoidal networks can achieve dynamical isometry with orthogonal weight initialization, leading to faster learning.

New method optimizes neural network learning by adjusting random parameters to target function features.

problem Difficulty in setting optimal random parameters for neural network learning.
method Adjusts sigmoid slopes and positions to target function features in a randomized learning method.
result Significantly better approximation of complex target functions compared to standard methods.

Unified Bayesian model explains in-context learning and activation steering in LLMs.

problem Understanding and controlling the behavior of large language models (LLMs) through prompts and activations.
method Developed a Bayesian model to explain and predict the effects of in-context learning and activation steering.
result Unified model predicts distinct phases and sudden shifts in LLM behavior, explaining prior empirical phenomena.

We propose a new class of learning algorithms that combines variational approximation and Markov chain Monte Carlo (MCMC) simulation. Naive algorithms that use the variational approximation as proposal distribution can perform poorly because this approximation tends to underestimate the true variance and other features…

2013-01-10abs ↗pdf ↗

Researchers use statistical physics to model neural network learning dynamics.

problem Understanding the learning dynamics of ReLU neural networks.
method Developed a system of differential equations using statistical physics techniques.
result ReLU networks exhibit distinct learning behavior compared to sigmoidal networks.

We developed safe and ranged approximations for ReLU, tanh, and sigmoid functions to reduce neural network training time.

problem Expensive computation of hyperbolic tangent and sigmoid functions in neural networks.
method Function approximation techniques to create safe and ranged approximations.
result 10% to 37% improvement in training times on CPU and 20% to 53% improvement in ranged cases.

Sigmoid gating is more sample efficient than softmax in mixture of experts.

problem Softmax gating leads to unnecessary competition among experts, causing representation collapse.
method Theoretical analysis of a regression framework with mixture of experts, identifying identifiability conditions and convergence rates.
result Sigmoid gating requires fewer samples to achieve the same expert estimation error as softmax gating.

Belief Propagation solves a relaxed network flow problem.

problem Generalized Min-Cost Network Flow with relaxed flow conservation constraints.
method Extends Belief Propagation to solve a new class of network flow problems.
result Belief Propagation converges to the exact solution of the relaxed network flow problem.

Deep belief networks are a powerful way to model complex probability distributions. However, learning the structure of a belief network, particularly one with hidden units, is difficult. The Indian buffet process has been used as a nonparametric Bayesian prior on the directed structure of a belief network with a single…

2009-12-31abs ↗pdf ↗

New method learns belief representations for GAIL in POMDPs.

problem Imitation learning in partially observable Markov decision processes (POMDPs).
method Joint learning of belief module and policy with task-aware imitation loss and belief regularization.
result Our BMIL approach outperforms GAIL and task-agnostic belief learning.

Unified probabilistic framework for nonlinearities in neural networks.

problem Lack of a unified approach to incorporating nonlinearities in neural networks.
method Doubly truncated Gaussian distributions for generating various nonlinearities.
result Performance improvements in RBM, temporal RBM, and TGGM when nonlinearities are learned alongside weights.

S-GAI initializes MLPs using spectral geometry from data, improving performance.

problem Lack of guidance on initial weights encoding data geometry.
method S-GAI uses SVD to estimate spectral class geometry, initializing MLPs from training data.
result S-GAI-initialized MLPs start from a more informative hidden state and achieve comparable accuracy.

The study explores how Matrix Product States can represent boolean and continuous functions.

problem Representing arbitrary boolean and continuous functions using Matrix Product States.
method Developed a construction method for MPS to represent boolean gates and proved density in continuous function space.
result MPS can accurately represent arbitrary boolean functions and continuous functions densely.

We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.

problem Improving word embeddings quality and understanding their relationship with hyperbolic spaces.
method Analyzing squashed shifted PMI matrix and its relation to graph properties and hyperbolic geometry.
result Word embeddings can be connected to hyperbolic spaces through squashed shifted PMI matrix.

SOLBP extends efficient inference to uncertain Bayesian networks.

problem Inference in uncertain Bayesian networks with second-order probabilities.
method Extends Loopy Belief Propagation to second-order Bayesian networks.
result Generates inferences consistent with sum-product networks, more efficient and scalable.

New method improves random parameter generation in neural networks.

problem Standard method of generating random weights and biases in neural networks has drawbacks.
method Proposes a new method to generate random parameters ensuring nonlinear sigmoids remain in the input hypercube and uniformly distributed slope angles for activation functions.
result Ensures the most useful nonlinear fragments of sigmoids remain in the input hypercube.

The paper uses attention networks for character-based handwritten text transcription.

problem Handwritten text recognition with improved character-level alignment.
method Attentional encoder-decoder networks trained on character sequences, comparing different activation functions.
result Softmax attention provides more precise character alignment than sigmoid attention.