Paper introduces hierarchical softmax for global hierarchical classification tasks.
problem Improving classification accuracy in tasks with class hierarchies.
method Global hierarchical neural networks using hierarchical softmax.
result Hierarchical softmax outperforms regular softmax in multiple datasets.
Improved classifier accuracy by using more of the class-specific structure in trained models.
problem Softmax ignores valuable information encoded in the full array of class response distributions.
method Developed a hybrid classifier (Softmax-Pooling Hybrid, SPH) that uses Softmax on high-scoring samples and a log-likelihood method on low-scoring samples. result Reduces test set error by 6% to 23% using the exact same trained model.
Efficiently solves inverse classification problems for logistic and softmax models.
problem Finding instances that change classifier predictions.
method Closed-form solution for logistic regression, iterative optimization for softmax.
result Fast, exact solutions for high-dimensional instances and many classes.
New neural network approach using mutual information.
problem Training neural networks for imbalanced datasets.
method Converts neural network classifiers to mutual information evaluators.
result New form of softmax leads to better classification accuracy, especially for imbalanced datasets.
Heating up softmax improves feature compactness for better metric learning.
problem Learning embeddings for samples of the same category to be compact while different categories spread out.
method Training classifiers with different temperature values of softmax function, then heating up the classifier.
result Classifier with increasing temperatures achieves state-of-the-art performance on metric learning benchmarks.
Paper introduces Balanced Meta-Softmax for better long-tailed visual recognition.
problem Long-tailed distribution mismatch between training and testing data.
method Balanced Meta-Softmax, an unbiased extension of Softmax, using a Meta Sampler.
result Balanced Meta-Softmax outperforms state-of-the-art solutions on visual recognition and instance segmentation.
A new operator based on t-distributions improves NN classifiers' robustness to out-of-distribution samples.
problem NN classifiers assign extreme probabilities to out-of-distribution samples, leading to unreliable predictions.
method Derive a novel operator using t-distributions to model uncertainty more accurately.
result Classifiers using the new operator are more robust to out-of-distribution samples.
Simple method improves deep classifier accuracy under noisy labels.
problem Training deep classifiers with noisy labels.
method Probabilistic approach using temperature parameterized softmax.
result Improves accuracy, log-likelihood and calibration on noisy datasets.
A method to reduce computation by dynamically sacrificing accuracy in deep neural networks.
problem Balancing computational effort and classification accuracy in deep neural networks.
method A cascade of deep neural networks with dynamically set confidence thresholds based on softmax outputs.
result Reduces 15%-50% in MAC operations with a 1% accuracy degradation.
Proposes a new method to improve multiclass probability calibration.
problem Uncalibrated class probabilities in multiclass classifiers leading to over-confidence.
method Dirichlet calibration method applicable to any model class, derived from Dirichlet distributions.
result Improved probabilistic predictions across various datasets and classifiers.
Simplified plug-in loss approximates EDL for reliable uncertainty estimation.
problem Efficient and reliable uncertainty estimation in real-world sensor-based learning systems.
method Approximate Dirichlet expected objectives with plug-in losses evaluated at the Dirichlet mean.
result Plug-in losses provide comparable predictive accuracy and selective prediction performance to classical EDL, while being simpler to implement.
Proposes a sparse classifier for discriminative Gaussian Mixture Models.
problem Softmax-based discriminative models assume unimodality, leading to parameter redundancy.
method Sparse Bayesian learning for GMM-based discriminative model, reducing parameters and complexity.
result The SDGM outperforms existing softmax-based discriminative models.
Softmax cross-entropy optimizes mutual information in neural networks.
problem Understanding the relationship between mutual information and classification neural networks.
method Demonstrated that optimizing softmax cross-entropy maximizes mutual information between inputs and labels.
result Softmax cross-entropy can approximate mutual information and highlight relevant image regions.
Study shows MSE with sigmoid can match SCE in classification tasks, especially with noisy data.
problem Inconsistent errors in neural network classification tasks.
method Introduced Output Reset algorithm to use MSE with sigmoid activation.
result MSE with sigmoid activation achieves comparable accuracy and convergence rates to Softmax Cross-Entropy, especially in noisy data scenarios.
New method improves object detection models for long-tailed datasets.
problem Classifier imbalance in long-tail object detection datasets.
method Balanced Group Softmax (BAGS) module for balanced training of classifiers.
result Significantly improves performance of object detection models.
Bayesian approach updates pretrained convnet for new image categories.
problem Learning new categories with limited data.
method Bayesian procedure using pretrained convnet weights as prior.
result Competitive performance with state-of-the-art methods.
New neural network approach for projection-free optimization.
problem Feasibility constraints in optimization problems.
method Designing projection-free convex optimization algorithms as Frank-Wolfe Networks.
result LSTM-learned optimizers outperform hand-designed and unconstrained optimizers.
Balanced Activation improves object detection performance on long-tailed datasets.
problem Mismatch between training and testing label distributions in object detection.
method Introduces Balanced Activation (Balanced Softmax and Balanced Sigmoid) to address label distribution shift.
result Balanced Activation provides ~3% gain in mAP on LVIS-1.0 compared to state-of-the-art methods.
Improves uncertainty estimation and OOD detection in neural networks.
problem Accurate uncertainty estimation and OOD detection in neural networks.
method Investigates one-vs-all and distance-based logit representations for probabilities.
result One-vs-all formulations improve calibration without additional complexity.
Efficient algorithm for evaluating hierarchical classification methods at multiple operating points.
problem Evaluating hierarchical classification methods at multiple operating points.
method Efficient algorithm to produce operating characteristic curves for any method that assigns scores to every class in the hierarchy.
result Top-down classifiers are dominated by a naive flat softmax classifier across the entire operating range.
Paper proposes an adversarial sampling method for efficient extreme classification.
problem Training classifiers over many classes is computationally expensive.
method Adversarial sampling to draw negative samples from an adversarial model.
result Significantly reduces training time by an order of magnitude.
BERT improves Chinese word segmentation performance.
problem Chinese word segmentation task.
method Applying BERT to CWS task using benchmark datasets.
result BERT can improve performance even with inconsistent labels.
Probabilistic label trees improve XMLC by organizing labels hierarchically.
problem Efficiently tagging instances with a small subset of relevant labels from a large pool.
method Introduce and analyze probabilistic label trees (PLTs) as a generalization of hierarchical softmax for multi-label problems.
result PLTs are consistent for various performance metrics and can be trained online without prior knowledge.
Proposes sigsoftmax to overcome the softmax bottleneck in language models.
problem Softmax function acts as a bottleneck in neural network representational capacity.
method Identifies the cause of softmax bottleneck and proposes sigsoftmax as a new activation function.
result Sigsoftmax outperforms softmax in language modeling tasks.
Develops RF-softmax for faster training with softmax cross entropy.
problem High computational cost of training with softmax cross entropy.
method Random Fourier Features for efficient sampling from approximate softmax distribution.
result RF-softmax provides low bias in estimating both softmax distribution and its gradient.
Bayesian method improves few-shot classification accuracy.
problem Few-shot classification with small labeled datasets.
method Gaussian process classifier with Pólya-Gamma augmentation and one-vs-each softmax.
result Improved accuracy and uncertainty quantification.
Hierarchical Softmax approximates class probabilities for large datasets efficiently.
problem Computational inefficiency of Softmax for large-scale classification tasks.
method Used Hierarchical Softmax to approximate class probabilities efficiently.
result Hierarchical Softmax performance degrades as the number of classes increases.
Improves VAE by adding a softmax classifier to enhance mutual information.
problem Posterior collapse and blurred reconstructions in VAE.
method Integrates an auxiliary softmax multi-classifier to improve VAE training.
result VAE-AS significantly improves mutual information and solves posterior collapse.
Revises logistic-softmax likelihood for Bayesian meta-learning in few-shot classification.
problem Inherent uncertainty in logistic-softmax leads to suboptimal performance in meta-learning.
method Redesigns logistic-softmax likelihood with a temperature parameter for better control of prior confidence.
result Achieves well-calibrated uncertainty estimates and comparable/superior performance on benchmark datasets.
EBNCs are a new model for Bayesian networks, derived from different networks.
problem Creating efficient models for Bayesian networks with discrete variables.
method Developed an EBNC model derived from different Bayesian networks, showing it's a special case of softmax polynomial regression.
result EBNCs can be used as a special case of softmax polynomial regression models.
DS-Softmax speeds up softmax inference by learning sparse experts.
problem Expensive softmax computations for large output classes.
method Sparse mixture of sparse experts for efficient top-k class retrieval.
result Significant computation reductions achieved at no performance loss.
Softmax emerges naturally in neural networks as a measure of conditional mutual information.
problem The artificial nature of softmax in neural networks.
method Information-theoretic perspective to derive log-softmax and evaluate conditional mutual information.
result Training deterministic neural networks through log-softmax maximises conditional mutual information.
Binary testing for softmax models requires many samples, similar to leverage score models.
problem Binary hypothesis testing for softmax models and leverage score models.
method Analyzing sample complexity and drawing analogies between models.
result Sample complexity is asymptotically \(O(ε^{-2})\), where \(ε\) is the distance between model parameters.
New algorithms make softmax optimization unbiased and scalable.
problem Efficiently computing softmax distributions with large categories.
method Proposed unbiased algorithms for maximizing softmax likelihood.
result Comprehensive outperformance on seven real-world datasets.
Topology aids in solving machine learning classification problems.
problem Machine learning classification problems.
method Classical topology applied to neural networks.
result Topology guides neural network architecture and training.
The paper investigates polynomial alternatives to softmax in transformer models.
problem The effectiveness of softmax attention in transformers is questioned.
method The authors explore polynomial activations as alternatives to softmax, focusing on their ability to regularize the attention matrix.
result Certain polynomials can serve as effective substitutes for softmax in transformer applications, achieving strong performance.
Softmax temperature influences model representation rank and performance.
problem Understanding and optimizing softmax function's impact on model representations.
method Investigated softmax function's role in deep neural networks, introduced rank deficit bias.
result Softmax temperature affects model representation rank and can improve performance.
Cross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a gene…
Paper proposes a new softmax loss for better performance in Positive and Unlabeled data tasks.
problem Current softmax losses and sampling schemes have drawbacks in Positive and Unlabeled learning.
method Proposes Relaxed Softmax (RS) loss and a new negative sampling scheme.
result New training objective drives uplifts in performance on textual and recommendation datasets.
New ACE cost function encourages diversity in neural networks.
problem Training multiple classifiers with controlled diversity.
method Mathematical derivation and gradient control.
result ACE yields better ensemble results than vanilla.
This research improves neural network uncertainty estimates and reliability.
problem Lack of inherent uncertainty estimates and variability in softmax scores.
method Ensemble-based Dirichlet modeling with method of moments estimator.
result Improved stability and predictive uncertainty estimates.
Softmax confidence misrepresents uncertainty in neural networks.
problem Neural networks fail to increase uncertainty on out-of-distribution data.
method Investigates two implicit biases in softmax confidence.
result Softmax confidence correlates with epistemic uncertainty due to decision boundary structure and deep network filtering.
Unified framework for studying softmax attention under large prompts.
problem Challenges in theoretical analysis of softmax attention.
method Measure-based framework for finite and infinite prompts.
result Softmax attention converges to linear attention in the large-prompt regime.
A method detects abnormal samples and adversarial attacks.
problem Detecting abnormal samples and adversarial attacks in deep neural networks.
method Obtain class conditional Gaussian distributions and use Mahalanobis distance for confidence score.
result Achieves state-of-the-art performance for both out-of-distribution and adversarial samples.
In a multi-class classification problem, it is standard to model the output of a neural network as a categorical distribution conditioned on the inputs. The output must therefore be positive and sum to one, which is traditionally enforced by a softmax. This probabilistic mapping allows to use the maximum likelihood pri…
Estimates softmax parameters without data, using class geometry.
problem Softmax parameter estimation with limited labeled data.
method Solves linear equations based on class geometry specifications.
result Closed-form solutions possible without data sampling.
Transformers use ReLUs to approximate softmax efficiently.
problem Analyzing resource usage in softmax transformer models.
method Translating ReLU approximation results to softmax attention mechanisms.
result Economic resource bounds for softmax attention mechanisms.
Paper shows softmax output misleads in evaluating adversarial example strength.
problem Softmax output misleads in evaluating adversarial example strength.
method Demonstrates how adversarial examples can exploit softmax properties.
result Softmax output is a poor indicator of adversarial example strength.