A new softmax operator improves reinforcement learning algorithms.
problem Softmax operators can misbehave in reinforcement learning, leading to suboptimal policies.
method Developed a differentiable softmax operator and introduced a SARSA variant using it.
result The new algorithm converges and performs well in practice.
Softmax Bellman operator improves Q-function performance in RL despite sub-optimality.
problem Softmax Bellman operator's impact on value functions in RL is problematic.
method Revisited theoretical properties of softmax Bellman operator, proving convergence and overestimation reduction.
result Softmax Bellman operator leads to superior policies in practice, even outperforming double Q-learning.
DBS improves convergence in reinforcement learning.
problem Softmax operator convergence issues in reinforcement learning.
method Dynamic Boltzmann Softmax (DBS) updates value function.
result DBS enables better value function estimation and convergence.
Softmax is found ineffective for NL block, leading to improved performance.
problem Inefficiency of softmax in NL block for global context modeling.
method Empirical analysis and replacement of softmax with scaling factor.
result Improved performance on various datasets with reduced computational cost.
Unified framework for studying softmax attention under large prompts.
problem Challenges in theoretical analysis of softmax attention.
method Measure-based framework for finite and infinite prompts.
result Softmax attention converges to linear attention in the large-prompt regime.
A new operator based on t-distributions improves NN classifiers' robustness to out-of-distribution samples.
problem NN classifiers assign extreme probabilities to out-of-distribution samples, leading to unreliable predictions.
method Derive a novel operator using t-distributions to model uncertainty more accurately.
result Classifiers using the new operator are more robust to out-of-distribution samples.
ZeroS improves Transformers by adding negative weights, matching or beating softmax attention.
problem Limited performance of linear attention methods, especially in long context sequences.
method Proposes Zero-Sum Linear Attention (ZeroS) that removes the zero-order term and reweights zero-sum softmax residuals.
result ZeroS matches or exceeds standard softmax attention across various benchmarks, theoretically expanding representable functions.
Paper proposes Sinkformers for Transformers with doubly stochastic attention.
problem Improving Transformer models' accuracy in vision and natural language processing.
method Using Sinkhorn's algorithm to make attention matrices doubly stochastic instead of SoftMax normalization.
result Sinkformers enhance model accuracy in vision and natural language processing tasks.
A method to reduce computation by dynamically sacrificing accuracy in deep neural networks.
problem Balancing computational effort and classification accuracy in deep neural networks.
method A cascade of deep neural networks with dynamically set confidence thresholds based on softmax outputs.
result Reduces 15%-50% in MAC operations with a 1% accuracy degradation.
Efficient algorithm for evaluating hierarchical classification methods at multiple operating points.
problem Evaluating hierarchical classification methods at multiple operating points.
method Efficient algorithm to produce operating characteristic curves for any method that assigns scores to every class in the hierarchy.
result Top-down classifiers are dominated by a naive flat softmax classifier across the entire operating range.
This paper introduces Gumbel-Sinkhorn networks for learning latent matchings.
problem Learning in latent variable models with permutations is difficult due to combinatorial intractability.
method Approximates maximum-weight matching using the Sinkhorn operator, extending Gumbel-Softmax.
result Demonstrates effectiveness on sorting, jigsaw puzzles, and neural signal identification tasks.
A new framework for sparse and structured neural attention.
problem Improving interpretability and performance of neural networks.
method Proposes a smoothed max operator framework for attention mechanisms.
result Improved interpretability without sacrificing performance.
Proposes sigsoftmax to overcome the softmax bottleneck in language models.
problem Softmax function acts as a bottleneck in neural network representational capacity.
method Identifies the cause of softmax bottleneck and proposes sigsoftmax as a new activation function.
result Sigsoftmax outperforms softmax in language modeling tasks.
Develops RF-softmax for faster training with softmax cross entropy.
problem High computational cost of training with softmax cross entropy.
method Random Fourier Features for efficient sampling from approximate softmax distribution.
result RF-softmax provides low bias in estimating both softmax distribution and its gradient.
New RFs reduce kernel approximation variance and improve Transformer performance.
problem Efficient approximation of Gaussian and softmax kernels for kernel methods and Transformers.
method Parameterized, positive, non-trigonometric RFs optimized for variance reduction.
result Significant variance reduction in practice, outperforming previous methods.
LARF improves random forests with attention mechanisms and contamination models.
problem Improving accuracy in classification tasks with random forests.
method Introduces a two-level attention mechanism and uses a mixture of contamination models.
result Significantly improved classification performance on various datasets.
Hierarchical Softmax approximates class probabilities for large datasets efficiently.
problem Computational inefficiency of Softmax for large-scale classification tasks.
method Used Hierarchical Softmax to approximate class probabilities efficiently.
result Hierarchical Softmax performance degrades as the number of classes increases.
Revises logistic-softmax likelihood for Bayesian meta-learning in few-shot classification.
problem Inherent uncertainty in logistic-softmax leads to suboptimal performance in meta-learning.
method Redesigns logistic-softmax likelihood with a temperature parameter for better control of prior confidence.
result Achieves well-calibrated uncertainty estimates and comparable/superior performance on benchmark datasets.
New insights into CE dynamics reveal how Hadamard initialization simplifies softmax.
problem Understanding the dynamics of cross-entropy training loss in deep learning.
method Analyzing a two-layer linear neural network with standard-basis vectors as inputs.
result Gradient flow on cross-entropy converges to neural collapse geometry, proving global convergence.
Paper introduces Balanced Meta-Softmax for better long-tailed visual recognition.
problem Long-tailed distribution mismatch between training and testing data.
method Balanced Meta-Softmax, an unbiased extension of Softmax, using a Meta Sampler.
result Balanced Meta-Softmax outperforms state-of-the-art solutions on visual recognition and instance segmentation.
Paper introduces hierarchical softmax for global hierarchical classification tasks.
problem Improving classification accuracy in tasks with class hierarchies.
method Global hierarchical neural networks using hierarchical softmax.
result Hierarchical softmax outperforms regular softmax in multiple datasets.
DS-Softmax speeds up softmax inference by learning sparse experts.
problem Expensive softmax computations for large output classes.
method Sparse mixture of sparse experts for efficient top-k class retrieval.
result Significant computation reductions achieved at no performance loss.
Softmax emerges naturally in neural networks as a measure of conditional mutual information.
problem The artificial nature of softmax in neural networks.
method Information-theoretic perspective to derive log-softmax and evaluate conditional mutual information.
result Training deterministic neural networks through log-softmax maximises conditional mutual information.
Proposes L-Softmax loss for CNNs to improve feature discriminativeness.
problem Lack of explicit feature discriminativeness in cross-entropy loss.
method Introduces L-Softmax loss that encourages intra-class compactness and inter-class separability.
result Deeply learned features with L-Softmax loss are more discriminative, boosting performance.
Binary testing for softmax models requires many samples, similar to leverage score models.
problem Binary hypothesis testing for softmax models and leverage score models.
method Analyzing sample complexity and drawing analogies between models.
result Sample complexity is asymptotically \(O(ε^{-2})\), where \(ε\) is the distance between model parameters.
New algorithms make softmax optimization unbiased and scalable.
problem Efficiently computing softmax distributions with large categories.
method Proposed unbiased algorithms for maximizing softmax likelihood.
result Comprehensive outperformance on seven real-world datasets.
The paper investigates polynomial alternatives to softmax in transformer models.
problem The effectiveness of softmax attention in transformers is questioned.
method The authors explore polynomial activations as alternatives to softmax, focusing on their ability to regularize the attention matrix.
result Certain polynomials can serve as effective substitutes for softmax in transformer applications, achieving strong performance.
Softmax temperature influences model representation rank and performance.
problem Understanding and optimizing softmax function's impact on model representations.
method Investigated softmax function's role in deep neural networks, introduced rank deficit bias.
result Softmax temperature affects model representation rank and can improve performance.
Improved classifier accuracy by using more of the class-specific structure in trained models.
problem Softmax ignores valuable information encoded in the full array of class response distributions.
method Developed a hybrid classifier (Softmax-Pooling Hybrid, SPH) that uses Softmax on high-scoring samples and a log-likelihood method on low-scoring samples. result Reduces test set error by 6% to 23% using the exact same trained model.
Paper proposes a new softmax loss for better performance in Positive and Unlabeled data tasks.
problem Current softmax losses and sampling schemes have drawbacks in Positive and Unlabeled learning.
method Proposes Relaxed Softmax (RS) loss and a new negative sampling scheme.
result New training objective drives uplifts in performance on textual and recommendation datasets.
Softmax confidence misrepresents uncertainty in neural networks.
problem Neural networks fail to increase uncertainty on out-of-distribution data.
method Investigates two implicit biases in softmax confidence.
result Softmax confidence correlates with epistemic uncertainty due to decision boundary structure and deep network filtering.
In a multi-class classification problem, it is standard to model the output of a neural network as a categorical distribution conditioned on the inputs. The output must therefore be positive and sum to one, which is traditionally enforced by a softmax. This probabilistic mapping allows to use the maximum likelihood pri…
GANs use Gumbel-softmax for generating sequences of discrete elements.
problem GANs struggle with discrete sequences due to non-differentiability of multinomial distributions.
method Used Gumbel-softmax distribution as a continuous approximation to a multinomial distribution for discrete elements.
result GANS with Gumbel-softmax outperform traditional GANs in generating sequences of discrete elements.
Estimates softmax parameters without data, using class geometry.
problem Softmax parameter estimation with limited labeled data.
method Solves linear equations based on class geometry specifications.
result Closed-form solutions possible without data sampling.
Transformers use ReLUs to approximate softmax efficiently.
problem Analyzing resource usage in softmax transformer models.
method Translating ReLU approximation results to softmax attention mechanisms.
result Economic resource bounds for softmax attention mechanisms.
Paper shows softmax output misleads in evaluating adversarial example strength.
problem Softmax output misleads in evaluating adversarial example strength.
method Demonstrates how adversarial examples can exploit softmax properties.
result Softmax output is a poor indicator of adversarial example strength.
Improved SincNet for better speaker recognition.
problem Speaker recognition challenges and the need for better deep learning models.
method Proposes AM-SincNet, a SincNet-based model with an improved AM-Softmax layer.
result Improved speaker recognition performance, achieving a 40% Frame Error Rate reduction.
New method adds uncertainty estimation to softmax outputs.
problem Uncertainty estimation in neural networks.
method Extend softmax layer with an additional constant input.
result Performs comparably to more computationally expensive methods.
Solves challenges in estimating parameters of softmax gating Gaussian mixture models.
problem Identifiability issues and complex interactions in Gaussian mixture of experts.
method Proposes novel Voronoi loss functions and establishes convergence rates of MLE.
result Connects convergence rate of MLE to a solvability problem of polynomial equations.
Quaternion self-attention reduces computational cost and improves performance.
problem Existing quaternion self-attention increases computational cost and diverges attention distributions.
method Proposes a shared-score quaternion self-attention mechanism.
result Reduces score-computation multiplications by 75% and softmax operations from four to one.
The paper proves neural networks with ReLU and softmax can approximate any function.
problem Approximating functions and class labels in neural networks.
method Extended universal approximator theory to neural networks with ReLU and softmax.
result Neural networks with ReLU and softmax can approximate any function and class labels.
Linear Q-learning converges to a bounded set without divergence.
problem Proving linear Q-learning does not diverge and converges to a bounded set.
method No modifications to the original linear Q-learning algorithm, no Bellman completeness or near-optimality assumptions, only an ε-softmax behavior policy with adaptive temperature.
result First L2 convergence rate of linear Q-learning iterates to a bounded set. Proposes learnable monotonic functions to improve softmax's limitations.
problem Softmax's limited representational capacity in large output vocabularies.
method Learn parametric monotonic functions on logits.
result Improves quality metrics over traditional Linear-Softmax in language models.
Efficiently approximates softmax probabilities for large-scale inference.
problem High cost of computing softmax probabilities for large-scale inference.
method Introduces a lower bound on softmax probabilities as a product of pairwise probabilities, scalable through stochastic optimization and subsampling.
result Demonstrates that the new bound has interesting theoretical properties and can be used in classification problems.
Introduces gradient decay in Softmax for better generalization.
problem Improving generalization performance in neural networks.
method Gradient decay hyperparameter in Softmax for varying gradient rates based on probability.
result Gradient decay rate affects generalization performance and can be tuned for better optimization.
Unified interpretation of softmax cross-entropy and negative sampling for knowledge graph embedding.
problem Lack of theoretical relationship between softmax cross-entropy and negative sampling loss functions in knowledge graph embedding.
method Used Bregman divergence to provide a unified interpretation of the two loss functions.
result Theoretical findings for fair comparison of softmax cross-entropy and negative sampling are derived.
New L2 regularization improves softmax MAB performance.
problem Improving softmax MAB performance with vanishing regularization.
method L2 regularization with vanishing parameter analyzed and proven convergent.
result Vanishing L2 regularization makes softmax MAB more numerically advantageous.
Gradient flow in softmax models tends to produce low-entropy outputs.
problem Understanding the training dynamics of softmax-based models.
method Analysis of gradient flow dynamics in the value-softmax model.
result Gradient flow drives optimization towards low-entropy solutions.