Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

149298447596 · Jun 202019922001200920172026
48 results for CE loss limitations

New method improves model calibration by adjusting confidence based on prediction correctness.

problem Improving model confidence alignment with true class probabilities.
method Post-hoc calibration objective using transformed samples for training.
result Competitive calibration performance on in-distribution and out-of-distribution test sets.

We present the Tamed Cross Entropy (TCE) loss function, a robust derivative of the standard Cross Entropy (CE) loss used in deep learning for classification tasks. However, unlike other robust losses, the TCE loss is designed to exhibit the same training properties than the CE loss in noiseless scenarios. Therefore, th…

2018-10-11abs ↗pdf ↗

New insights into CE dynamics reveal how Hadamard initialization simplifies softmax.

problem Understanding the dynamics of cross-entropy training loss in deep learning.
method Analyzing a two-layer linear neural network with standard-basis vectors as inputs.
result Gradient flow on cross-entropy converges to neural collapse geometry, proving global convergence.

Generative Cross-Entropy improves classification with fewer labels.

problem Limited sample efficiency of cross-entropy loss in data-scarce scenarios.
method Proposes Generative Cross-Entropy (GenCE), a new loss function that incorporates generative principles into a standard discriminative network.
result Generative Cross-Entropy outperforms traditional cross-entropy loss across various datasets and conditions.

The paper provides a uniform convergence bound for smooth calibration error and its relationship with functional gradient.

problem Limited theoretical understanding of learning algorithms achieving high accuracy and good calibration.
method Focuses on smooth calibration error, providing a uniform convergence bound and proving the relationship with functional gradient.
result Derives conditions for simultaneous classification and calibration guarantees in gradient boosting trees, kernel boosting, and neural networks.

Deep nets trained with MSE loss exhibit Neural Collapse, collapsing features and classifiers to class means.

problem Understanding Neural Collapse in MSE-trained deep nets.
method Developed a new MSE loss decomposition and introduced the central path concept.
result Exact dynamics of Neural Collapse along the central path can be predicted.

Large learning rates work surprisingly well in standard parameterization, contrary to theory.

problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.

This paper examines how different loss functions affect neural network features and performance.

problem Investigating which loss function is best for deep neural networks.
method Examining last-layer features of deep networks and drawing inspiration from the Neural Collapse phenomenon.
result All relevant loss functions (CE, LS, FL, MSE) produce equivalent features and similar performance.

CE improves climate uncertainty quantification using GCM ensembles and observational data.

problem Uncertainty in climate projections due to model inadequacies and variability.
method Conformal ensembles integrating GCM ensembles and observational data.
result CE generates statistically rigorous, easy-to-interpret uncertainty estimates.

This work improves adversarial robustness by boosting model ensembles with margin maximization.

problem Single models are insufficient for defending against adversarial attacks.
method Margin-boosting approach to learn ensembles with maximum margin.
result Our algorithm outperforms existing ensembling techniques and large models trained end-to-end.

The paper develops a method to optimize individualized treatment rules for cost-effectiveness.

problem Developing cost-effective individualized treatment rules for healthcare policy.
method Using conditional random forest and net-monetary-benefit (NMB) to estimate optimal CE-ITR.
result The approach optimizes healthcare resource allocation by maximizing health gains and minimizing costs.

Ensembles of neural networks improve training dynamics and performance.

problem Improving neural network performance through model size increase.
method Defining collegial ensembles (CE) as multiple independent models trained as a single model, and using theoretical results on NTK to optimize architecture search.
result CE dynamics simplify and scale favorably, resembling wide models, and can be efficiently implemented using group convolutions and block diagonal layers.

CONFETTI improves interpretability of deep learning models for MTS by providing counterfactual explanations.

problem Lack of transparency in deep learning models for multivariate time series classification.
method CONFETTI is a novel multi-objective counterfactual explanation method that balances prediction confidence, proximity, and sparsity.
result CONFETTI outperforms state-of-the-art methods in various metrics, improving interpretability and decision support.

Stochastic momentum methods trade compute efficiency for serial runtime.

problem Stochastic momentum methods trade compute efficiency for serial runtime.
method Stochastic HB and ASGD for consistent linear regression with Gaussian covariates.
result HB preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale.

Variable selection is of significant importance for classification and regression tasks in machine learning and statistical applications where both predictability and explainability are needed. In this paper, a Copula Entropy (CE) based method for variable selection which use CE based ranks to select variables is propo…

2019-10-28abs ↗pdf ↗

Many applications of Bayesian data analysis involve sensitive information, motivating methods which ensure that privacy is protected. We introduce a general privacy-preserving framework for Variational Bayes (VB), a widely used optimization-based Bayesian inference method. Our framework respects differential privacy, t…

2016-11-01abs ↗pdf ↗

In this article we study the limiting behavior of the Kähler Ricci flow on complete non-compact Kähler manifolds. We provide sufficient conditions under which a complete non-compact gradient Kähler-Ricci soliton is biholomorphic to $\ce^n$. We also discuss the uniformization conjecture by Yau \cite{Y} for complete non-…

2003-10-14abs ↗pdf ↗

Our previous exploration of the $\cE_g^{PD}$-geometry has shown that the field is promising. Namely, the $\cE_g^{PD}$-approach is amenable to development of novel trends in relativistic and metric differential geometry and can particularly be effective in context of the Finslerian or Minkowskian Geometries. The main po…

2003-10-13abs ↗pdf ↗

Sector specific multifactor CES elasticity of substitution and the corresponding productivity growths are jointly measured by regressing the growths of factor-wise cost shares against the growths of factor prices. We use linked input-output tables for Japan and the Republic of Korea as the data source for factor price …

2016-08-03abs ↗pdf ↗

Causal discovery is a fundamental problem in statistics and has wide applications in different fields. Transfer Entropy (TE) is a important notion defined for measuring causality, which is essentially conditional Mutual Information (MI). Copula Entropy (CE) is a theory on measurement of statistical independence and is …

2019-10-10abs ↗pdf ↗

This paper improves loss functions for deep learning with noisy labels.

problem Training deep neural networks with noisy labels.
method The paper introduces a normalization technique to make any loss function robust to noisy labels and proposes a framework called Active Passive Loss (APL) to combine robust loss functions.
result The proposed APL framework consistently outperforms state-of-the-art methods, especially under high noise rates.

ReQuestNet simplifies 5G channel estimation with a unified model.

problem Complex channel estimation in 5G systems with varying conditions.
method Unified neural architecture that handles dynamic resource blocks and transmit layers.
result Significantly outperforms legacy methods, achieving up to 10dB gain at high SNRs.

In this paper, we propose a novel implicit semantic data augmentation (ISDA) approach to complement traditional augmentation techniques like flipping, translation or rotation. Our work is motivated by the intriguing property that deep networks are surprisingly good at linearizing features, such that certain directions …

2019-09-26abs ↗pdf ↗

Estimates calibration error under label shift without labels.

problem Ensuring model reliability in the face of dataset shift without access to labels.
method Importance re-weighting of the labeled source distribution to estimate calibration error under label shift.
result Effective and reliable CE estimation with respect to the shifted target distribution.

Proposes CE-BASS for robust Kalman filtering with innovative and additive outliers.

problem Robustness to both innovative and additive outliers in Kalman filtering.
method Particle mixture Kalman filter with re-sampling of past states.
result CE-BASS efficiently handles multi-modality and trend changes in hidden state distributions.

Training accurate deep neural networks (DNNs) in the presence of noisy labels is an important and challenging task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy (CE) exhibits overfitting to noisy lab…

2019-08-16abs ↗pdf ↗

Novel CE-method variants reduce local minima convergence with fewer function evaluations.

problem Local minima and expensive function evaluations in optimization.
method Surrogate model-based CE-method variants to reduce local minima convergence.
result Surrogate model-based approach reduces local minima convergence using fewer function evaluations.

New bounds on geodesic dimension and curvature exponent in Carnot groups.

problem Characterizing geodesic dimension and curvature exponent in Carnot groups.
method Characterization and lower bound calculation for geodesic dimension and curvature exponent.
result Found an example where curvature exponent is greater than geodesic dimension.

To any g\mathfrak{g}-manifold MM are associated two dglas tot(ΛgkTpoly)\operatorname{tot}\big(Λ^{\bullet} \mathfrak{g}^\vee \otimes_{\Bbbk} T_{\operatorname{poly}}^{\bullet} \big) and tot(ΛgkDpoly)\operatorname{tot} \big(Λ^{\bullet} \mathfrak{g}^\vee\otimes_{\Bbbk} D_{\operatorname{poly}}^{\bullet} \big), whose cohomologies $H_{\operatorn…

2017-01-17abs ↗pdf ↗

Proposes sparse local and regional counterfactual rules for robust recourses.

problem Challenges in counterfactual explanations, especially stability, synthesis, and implementation.
method Probabilistic framework using Random Forest to derive sparse local and regional counterfactual rules.
result Effective recourses derived from high-density regions, providing sparse and robust counterfactual rules.

We review the properties of transversality of distributions with respect to submersions. This allows us to construct a convolution product for a large class of distributions on Lie groupoids. We get a unital involutive algebra $\cE\_{r,s}'(G,Ω^{1/2})$ enlarging the convolution algebra C_c(G,Ω1/2)C^\infty\_c(G,Ω^{1/2}) associate…

2015-02-06abs ↗pdf ↗

This paper introduces DCE for better counterfactual explanations using optimal transport.

problem Lack of nuanced distributional characteristics in existing counterfactual explanations.
method Formulates a chance-constrained optimization problem using optimal transport to derive counterfactual distributions.
result DCE provides deeper insights into decision-making models by aligning counterfactual distributions with factual ones.

We construct a non-abelian extension ΓΓ of S1S^1 by $\cy 3 \times \cy 3$, and prove that ΓΓ acts freely and smoothly on S5×S5S^{5} \times S^{5}. This gives new actions on S5×S5S^{5} \times S^{5} for an infinite family $\cP$ of finite 3-groups. We also show that any finite odd order subgroup of the exceptional Lie group $G_…

2007-05-28abs ↗pdf ↗