Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4181122162 · May 202619922001200920182026
48 results for Cross entropy

Quantum machine learning uses quantum cross entropy to minimize loss, but measurement loss affects this process.

problem Quantum machine learning's loss minimization through cross entropy is affected by measurement outcomes.
method Defined quantum cross entropy, proved its lower bounds, and investigated its relation to quantum fidelity and likelihood.
result Quantum cross entropy is lower-bounded by negative log-likelihood when derived from quantum data, but measurement outcomes can cause loss.

Cross-entropy loss linked to metric learning, outperforming complex pairwise losses.

problem Improving metric learning performance without complex optimization schemes.
method Theoretical analysis linking cross-entropy to pairwise losses, showing cross-entropy as an upper bound and equivalent to mutual information maximization.
result Minimizing cross-entropy is equivalent to maximizing mutual information, leading to state-of-the-art performance.

Unified interpretation of softmax cross-entropy and negative sampling for knowledge graph embedding.

problem Lack of theoretical relationship between softmax cross-entropy and negative sampling loss functions in knowledge graph embedding.
method Used Bregman divergence to provide a unified interpretation of the two loss functions.
result Theoretical findings for fair comparison of softmax cross-entropy and negative sampling are derived.

Max-pooling loss improves LSTM KWS models with lower resource usage.

problem Training efficient LSTM networks for small-footprint keyword spotting.
method Max-pooling loss training guided by cross-entropy initialization, posterior smoothing evaluation.
result Max-pooling loss trained LSTM models outperform baseline DNNs with significant resource savings.

Study Transformer layers under cross-entropy training using mean field control.

problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.

Square loss performs comparably or better than cross-entropy in neural architectures for various tasks.

problem The superiority of cross-entropy loss over square loss in classification tasks is debated.
method Comparison of several neural architectures on NLP, ASR, and computer vision datasets using both loss functions.
result Square loss often produces better results in the majority of tasks, especially in NLP and ASR.

Theoretical analysis of cross-entropy loss functions and their robustness.

problem Guarantees for using cross-entropy as a surrogate loss function.
method Theoretical analysis of a broad family of loss functions, including cross-entropy.
result First HH-consistency bounds for comp-sum losses and smooth adversarial comp-sum losses.

Improved CNN for HCCR with new loss function and ranking method.

problem Loss of inter-class information in traditional CNN models for HCCR.
method Combining cross entropy with a new similarity ranking function (Average variance similarity) as loss function.
result New loss function (SoftMax cross entropy with Average variance similarity) achieves highest accuracy in HCCR.

Proposes a method to solve deep neural networks' local minimum problem.

problem Local minimum problem in deep neural networks training.
method Transforms cross-entropy loss into risk-averse error criterion, adjusts RSI, and uses convexity region.
result Trained deep learning machine is expected to be inside a global minimum's attraction basin.

Paper proposes a new loss function for conditional models using soft targets.

problem Improving generalization performance of deep neural networks on supervised classification tasks.
method Introduces a new loss function compatible with soft targets, based on noise contrastive estimation.
result Soft target InfoNCE loss performs on par with cross-entropy baselines and outperforms other losses.

Softmax cross-entropy optimizes mutual information in neural networks.

problem Understanding the relationship between mutual information and classification neural networks.
method Demonstrated that optimizing softmax cross-entropy maximizes mutual information between inputs and labels.
result Softmax cross-entropy can approximate mutual information and highlight relevant image regions.

Structured entropy improves classification performance on structured targets.

problem Cross-entropy loss fails to account for target variable structure.
method Proposes structured entropy, a generalization of entropy using random partitions.
result Structured cross-entropy loss yields better results on classification problems with known structure.

Paper proposes AXE loss for non-autoregressive machine translation, improving performance.

problem Challenges in training non-autoregressive models due to lack of autoregressive factors and cross entropy loss penalties.
method Proposes aligned cross entropy (AXE) loss function using a differentiable dynamic program for better word order alignment.
result AXE-based training improves performance on major WMT benchmarks and sets a new state of the art for non-autoregressive models.

Proposes squentropy loss for improved classification accuracy and model calibration.

problem Theoretical and empirical evidence for cross-entropy loss is lacking.
method Introduces squentropy loss as the sum of cross-entropy and average square loss over incorrect classes.
result Squentropy loss outperforms cross-entropy and rescaled square losses in classification accuracy and model calibration.

Proposes a new loss function for deep neural networks.

problem Deep neural networks lack a direct method to discriminate between correct and competing classes.
method Introduces a discriminative loss function based on negative log likelihood ratio.
result Significantly outperforms cross-entropy loss on image classification tasks.

Study on bit threads and their locking properties in holographic spacetimes.

problem Understanding the conditions under which regions can be locked in holographic spacetimes.
method Investigation of different density bounds and their implications on the locking of regions.
result Non-crossing regions can be locked under the most stringent bound, but crossing regions cannot.

Generative Cross-Entropy improves classification with fewer labels.

problem Limited sample efficiency of cross-entropy loss in data-scarce scenarios.
method Proposes Generative Cross-Entropy (GenCE), a new loss function that incorporates generative principles into a standard discriminative network.
result Generative Cross-Entropy outperforms traditional cross-entropy loss across various datasets and conditions.

Proposes a cross entropy loss for better ranking algorithms.

problem Improving the theoretical understanding and performance of ranking algorithms.
method Introduces a cross entropy-based loss function that is a convex bound on NDCG and consistent with NDCG.
result Empirically, the proposed method outperforms existing algorithms in quality and robustness.

The paper analyzes how well classes are separated in neural network feature space.

problem Understanding class separability in neural network feature space.
method Theoretical analysis of intra-class and inter-class distances in feature space.
result A lower bound for the probability of inter-class distance being greater than intra-class distance as a function of loss value.

This paper shows using classification instead of regression improves deep RL scalability.

problem Challenges in training value functions for large networks in deep RL.
method Used categorical cross-entropy loss instead of mean squared error regression.
result Significant improvements in performance and scalability across various domains.

Develops a mean-field theory for multi-head self-attention under cross-entropy training.

problem Mean-field analysis of multi-head self-attention under cross-entropy training.
method Mean-field theory for a simplified single-layer causal multi-head self-attention model.
result Proves a static finite-head approximation bound for the optimal risk.

A new policy improvement method using CEM for Actor-Critic.

problem Improving policy efficiency and robustness in reinforcement learning.
method Greedy Actor-Critic (Greedy AC) using Conditional Cross-Entropy Method (CCEM).
result Greedy AC outperforms Soft Actor-Critic and is less sensitive to entropy regularization.

This work shows that supervised contrastive learning achieves similar results to cross-entropy but requires more iterations.

problem The question of whether there are fundamental differences in representation geometry between supervised contrastive learning and cross-entropy.
method The authors prove that both losses attain their minimum when representations of each class collapse to the vertices of a regular simplex, and they empirically validate this finding.
result Supervised contrastive learning requires more iterations to reach a close-to-optimal state compared to cross-entropy, indicating different optimization behavior.

New algorithms and guarantees for multiple-source adaptation.

problem Improving model performance on target mixtures from multiple sources.
method Normalized solutions with theoretical guarantees, algorithms for distribution-weighted combination.
result Our algorithm outperforms competing approaches by producing a robust model.

REMEDI improves neural entropy estimation across various tasks.

problem Challenges in estimating information theoretic quantities in high-dimensional data.
method Combines minimization of cross-entropy with estimation of deviation from data density.
result Improves accuracy in entropy estimation on synthetic and natural data.

The R Package CEC performs clustering based on the cross-entropy clustering (CEC) method, which was recently developed with the use of information theory. The main advantage of CEC is that it combines the speed and simplicity of kk-means with the ability to use various Gaussian mixture models and reduce unnecessary cl…

2015-08-19abs ↗pdf ↗

Simple linear relationship explains test performance differences in deep networks.

problem Understanding why two deep networks with identical training and architecture have different test performance.
method Showed that cross-entropy loss can lead to drastically different generalization performances for networks with different initialization or corrupted training.
result A linear relationship emerges between training and test losses, revealing the intrinsic problem of measuring test performance with cross-entropy loss.

New analysis shows how cross-entropy training shapes attention in transformers.

problem Understanding how gradient-based learning creates the required internal geometry in transformers.
method Developed a first-order analysis of cross-entropy training effects on attention scores and values in a transformer attention head.
result Introduced an advantage-based routing law and responsibility-weighted update for attention scores and values, respectively.

New algorithms predict reinforcement learning values efficiently.

problem Predicting reinforcement learning values with linear function approximation.
method Multi-timescale stochastic approximation of cross entropy method.
result Proved convergence and achieved good performance in experiments.