MELC uses entropy for multithreshold classification, showing consistency similar to SVM.
problem Consistency of multithreshold linear classifiers.
method Employed multithreshold maximum margin model based on information theory.
result Objective function upper bounds misclassified points, similar to hinge loss.
Linear classifiers separate the data with a hyperplane. In this paper we focus on the novel method of construction of multithreshold linear classifier, which separates the data with multiple parallel hyperplanes. Proposed model is based on the information theory concepts -- namely Renyi's quadratic entropy and Cauchy-S…
Paper speeds up optimization of MELC classifier.
problem Optimization speed of MELC classifier.
method Approximate solutions of Kernel Density Estimation, conjugate gradients, L-BFGS optimizers.
result Improved optimization speed of MELC classifier.
Optimal policy reduces entropy in preference learning.
problem Learning user preferences via choice-based queries.
method Parametric model with Bayesian prior, greedy policy for entropy reduction.
result Greedy policy achieves linear lower bound on entropy reduction.
The paper introduces a minimax approach to supervised learning problems.
problem Optimal decision rule minimizing worst-case expected loss over probability distributions.
method Generalization of maximum entropy principle applied to constrained distributions.
result Developed a new linear classifier called the maximum entropy machine.
Deep neural networks converge quickly for classification tasks.
problem Classifying data with smooth decision boundaries, probabilities, or margins.
method Hinge loss and cross-entropy for training; analysis of convergence rates.
result DNNs achieve fast convergence rates under various conditions.
Proposes a method to robustly classify and detect anomalies.
problem Learning robust binary classifiers in the presence of corrupted measurements.
method Geometric-Entropy-Minimization regularized Maximum Entropy Discrimination (GEM-MED) method.
result Improved performance in classification accuracy and anomaly detection rate.
Proposes a method to classify with sensor failures, improving accuracy and anomaly detection.
problem Learning robust binary classifier with possible sensor failures.
method Geometric-Entropy-Minimization regularized Maximum Entropy Discrimination (GEM-MED) method.
result Improved performance in classification accuracy and anomaly detection rate.
Temperature scaling improves model uncertainty but not diversity in LLMs.
problem Improving the calibration and stochasticity of probabilistic models.
method Investigates theoretical properties of temperature scaling in classification and LLMs.
result Temperature scaling increases model uncertainty but not diversity in LLMs.
New ACE cost function encourages diversity in neural networks.
problem Training multiple classifiers with controlled diversity.
method Mathematical derivation and gradient control.
result ACE yields better ensemble results than vanilla.
EAST aligns neural network classifiers with user-defined evaluation metrics.
problem Mismatch between neural network training and evaluation metrics leads to suboptimal performance.
method EAST uses dynamic thresholding, soft-set confusion matrix, and annealing to align neural network predictions with target evaluation metrics.
result EAST improves alignment between training objectives and evaluation metrics, outperforming existing methods.
Paper studies Fenchel-Young losses for classifier construction.
problem Creating effective loss functions for classifiers.
method Analyzes Fenchel-Young losses from generalized entropies, formulates conditions for separation margins and sparse support.
result Fenchel-Young losses can induce predictive distributions with separation margins and sparse support.
New neural network approach using mutual information.
problem Training neural networks for imbalanced datasets.
method Converts neural network classifiers to mutual information evaluators.
result New form of softmax leads to better classification accuracy, especially for imbalanced datasets.
A new classifier method detects out-of-distribution samples by minimizing KL divergence.
problem Detecting out-of-distribution samples in neural networks.
method Training a confident-classifier by minimizing KL divergence and maximizing entropy, or adding a reject class.
result The confident-classifier still yields high confidence for OOD samples far from the in-distribution.
Paper develops MRCs for supervised classification using generalized maximum entropy.
problem Developing robust classifiers for decision problems.
method Generalized maximum entropy principle applied to minimax risk classifiers.
result Learning techniques for determining MRCs with performance guarantees.
New method controls classifier guidance in diffusion models.
problem Improving classifier guidance in diffusion models.
method Cross-entropy control of classifier gradients.
result Effective guidance vectors with mean squared error O ( d ε ) O(d \varepsilon ) O ( d ε ) . Optimized fuzzy entropy framework improves feature selection and classification performance.
problem Improving feature selection and classification in fuzzy entropy frameworks.
method Implemented and compared combinations of ideal vectors, maximal similarity classifiers, and fuzzy entropy functions.
result Optimized combination of ideal vector, similarity classifier, and fuzzy entropy function achieved the most stable performance for all three datasets.
New entropy formulae for heat equation on manifolds.
problem Entropy formulae for linear heat equation on Riemannian manifolds.
method Proved new entropy formulae for linear heat equation on static Riemannian manifolds with nonnegative Ricci curvature.
result Results are analogies of Cao and Hamilton's entropies for Ricci flow.
Study uses EEG features HFD and SampEn to detect depression with high accuracy.
problem Diagnosing depression reliably and accurately.
method Applied Higuchi Fractal Dimension and Sample Entropy on EEG signals using seven machine learning algorithms.
result Good classification possible even with small EEG data, achieving high accuracy.
Generates confident out-of-distribution samples to improve classifier robustness.
problem Overconfidence in deep learning models on out-of-distribution inputs.
method Uses a GAN to generate out-of-distribution samples that the classifier is confident on, maximizing entropy.
result Shows effectiveness on handwritten characters and natural images datasets.
Revises Bayesian model averaging for foundation models.
problem Ensemble pre-trained and lightly-finetuned foundation models for improved classification performance.
method Introduces trainable linear classifiers and computationally cheaper model averaging scheme (OMA).
result Ensembled models can better predict on various datasets.
Entropy-regularized NPG converges linearly with linear function approximation.
problem Analyzing convergence of entropy-regularized NPG with function approximation.
method Established finite-time convergence analyses with entropy regularization and linear function approximation.
result Entropy-regularized NPG achieves linear convergence up to a function approximation error.
A new approach for test-time adaptation detects and reacts to distribution shifts.
problem Improving test-time accuracy under distribution shifts.
method Online self-training with a detection tool based on entropy values and betting martingales.
result The classifier's entropy values match those of the source domain, building invariance to distribution shifts.
Paper proposes mutual information learning for deep learning classifiers.
problem Overfitting in deep learning models.
method Mutual information learning framework to train classifiers.
result MILCs achieve better generalization than conditional entropy classifiers.
Entropy-SGD optimizes a PAC-Bayes bound, leading to improved generalization.
problem Improving generalization in machine learning models.
method Entropy-SGD optimizes a PAC-Bayes bound by adjusting the prior, which is typically chosen independently of the data.
result Entropy-SGD can yield relatively tight generalization bounds and still fit real labels.
Unified entropy formula for real, complex, and quaternionic DLNs.
problem Deriving a formula for DLNs over different fields.
method Extending Menon and Yu's formula to complex and quaternionic DLNs.
result Unified entropy formula for DLNs over R \mathbb{R} R , C \mathbb{C} C , and H \mathbb{H} H . Self-training avoids spurious features in domain adaptation.
problem Domain shift with large differences between source and target domains.
method Entropy minimization on unlabeled target data, initialized with a source classifier.
result Entropy minimization avoids using spurious features in large domain shifts.
The paper classifies contact 3-manifolds with critical metrics and connects entropy to optimization.
problem Classifying contact 3-manifolds with critical metrics and understanding their entropy.
method Critical metrics optimization and entropy analysis.
result Anosov contact metrics' optimization is linked to Reeb dynamics and entropy.
Paper optimizes score transformation for fair binary classification.
problem Ensuring fairness in binary classification with predicted scores.
method Formulates and solves a convex optimization problem for transforming scores to meet fairness constraints.
result Derives a closed-form expression for optimal transformed scores and provides guarantees for finite sample settings.
Ancient flows by curvature powers in 2D have finite entropy.
problem Existence of non-homothetic ancient flows by powers of curvature in R 2 \mathbb{R}^2 R 2 . method Determined Morse indices and kernels of the linearized operator of shrinkers. Constructed flows using unstable eigenfunctions.
result Existence of ancient flows with finite entropy.
eSPA breaches overfitting barriers in ML with 10^-12 cost.
problem Overfitting in machine learning with small data.
method eSPA, entropy-optimal Scalable Probabilistic Approximations algorithm.
result 30-fold boost in classification performance with eSPA.
The concept of refinement from probability elicitation is considered for proper scoring rules. Taking directions from the axioms of probability, refinement is further clarified using a Hilbert space interpretation and reformulated into the underlying data distribution setting where connections to maximal marginal diver…
A new method ranks autoencoder hidden nodes based on their saliency.
problem Autoencoders lack saliency measures similar to PCA.
method Supervised node saliency (SNS) method using normalized entropy difference (NED).
result SNS identifies salient hidden nodes in autoencoders.
Uniform bounds derived for fully non-linear equations.
problem Bounding fully non-linear equations uniformly in background metrics.
method Auxiliary Monge-Ampère equations and entropy-like quantities.
result Uniform L ∞ L^\infty L ∞ bounds for systems coupling fully non-linear equations to their linearizations. A new classifier updates sequentially using maximum margin principles.
problem Sequential data collection and partial labeling.
method Maximum margin classifier with Maximum Entropy Discrimination principle, kernel representation, and regularization.
result Improved performance compared to non-sequential classifiers.
PEOC uses policy entropy to detect untrained states in RL.
problem Detecting untrained states in reinforcement learning for safety.
method Policy entropy based one-class classifier.
result PEOC is highly competitive and reliable.
Graph convolution improves linear separability and generalizes to out-of-distribution data.
problem Improving linear separability in semi-supervised classification.
method Applying graph convolution to mixtures of Gaussians in a stochastic block model.
result Graph convolution extends the linear separability regime by a factor of 1 / D 1/\sqrt{D} 1/ D . New method discovers nonlinear associations using copula entropy.
problem Linear association measures have theoretical limitations.
method Proposes a new method using copula entropy.
result Demonstrates more meaningful nonlinear associations.
We derive the entropy formula for the linear heat equaiton on complete Riemannian manifolds with nonnegative Ricci curvature. As applications, we study the relation between the value of entropy and the volume of balls of various scales. The results are simpler version, without Ricci flow, of Perelman's recent results o…
Machine learning for entropy calculation from binary signals.
problem Calculating entropy from binary configurations/signals.
method Transformed entropy calculation into supervised classification tasks using machine learning.
result Reproduced entropy and free energy of the 2D Ising model.
Deep linear networks exhibit collapsing features and classifiers across datasets.
problem Understanding the collapse of features and classifiers in deep linear networks.
method Theoretical and empirical analysis of deep linear networks with MSE and CE losses.
result Deep linear networks exhibit NC properties, collapsing features and classifiers to orthogonal vectors.
DNLL loss improves deep LDA accuracy and consistency.
problem Pathological solutions in unconstrained Deep LDA.
method Introducing Discriminative Negative Log-Likelihood (DNLL) loss.
result Deep LDA trained with DNLL produces clean latent spaces and better calibrated probabilities.
New risk bound derived for multi-category margin classifiers.
problem Guaranteed risk dependency on categories, sample size, and margin parameter.
method Derived a new risk bound using Rademacher complexity and chaining method.
result Improved dependency on categories over state of the art.
This paper introduces a new potential function using Tsallis entropy for neural network optimization.
problem The challenge of obtaining exponential convergence in neural network optimization.
method Utilizes a linearized potential function based on Csiszár type of Tsallis entropy.
result Derives an exponential convergence result in neural network optimization.
CPR adds entropy maximization to improve continual learning methods.
problem Catastrophic forgetting in continual learning.
method Classifier-Projection Regularization (CPR) adds an entropy maximization term to existing regularization methods.
result CPR improves accuracy and plasticity in continual learning methods.
Entropy formula derived for deep linear networks using geometric analysis.
problem Thermodynamic description of learning in deep linear networks.
method Group actions, Riemannian submersion, foliation, Jacobi matrices.
result Entropy formula defined on the balanced manifold of DLN.
Study uncertainty metrics from CNN gradients for deep learning tasks.
problem Uncertainty quantification in deep neural networks.
method Gradient metrics for uncertainty quantification, meta classification of predictions.
result Meta classification accuracy similar to entropy thresholding for CNN predictions.
Following Cao-Hamilton-Ilmanen, in this paper we study the linear stability of Perelman's ν ν ν -entropy on Einstein manifolds with positive Ricci curvature. We observe the equivalence between the linear stability restricted to the transversal traceless symmetric 2-tensors and the stability of Einstein manifolds with resp…