Proposes a generalized XGBoost method for nonconvex loss functions.
problem Limited to convex loss functions in XGBoost.
method Extends XGBoost to use nonconvex loss functions and multivariate loss functions.
result Generalized XGBoost method can model multiple parameters in various distributions.
New loss functions based on f-divergences improve language model performance.
problem Improving multiclass classification and language modeling performance.
method Constructing new convex loss functions using f-divergences and deriving an operator for computation.
result The α-divergence loss function with α=1.5 performs well across various tasks. Paper connects loss functions and t-norms for faster deep learning convergence.
problem Improving convergence rates in deep learning models.
method Interprets loss functions through t-norms and generator functions.
result Derives a general relation between loss functions and t-norms leading to faster convergence.
Paper proposes a new e-exponentiated transformation to make convex loss functions more robust to outliers.
problem Making convex loss functions robust to outliers in the presence of label noise.
method Introduces a novel e-exponentiated transformation for loss functions and proves its effectiveness through theoretical and empirical analysis. result The transformed loss function achieves tighter generalization error bounds and higher accuracy in noisy datasets.
Proposes a new loss function for robust learning.
problem Creating a robust loss function for machine learning.
method Extended pseudo Huber loss with log-exp transform and logistic function.
result Linear convergence algorithm for minimizer finding.
This paper explores neural network loss landscapes and their effects on generalization.
problem Understanding the structure of neural network loss functions and their impact on generalization.
method Simple filter normalization and various visualization methods to explore loss landscape structure and network architecture effects.
result Visualizations reveal how network architecture and training parameters affect loss landscape curvature and minimizers.
Paper studies Fenchel-Young losses for classifier construction.
problem Creating effective loss functions for classifiers.
method Analyzes Fenchel-Young losses from generalized entropies, formulates conditions for separation margins and sparse support.
result Fenchel-Young losses can induce predictive distributions with separation margins and sparse support.
Paper proposes new loss functions for training energy networks.
problem Challenges in computing gradients for training energy networks.
method Proposes generalized Fenchel-Young losses for efficient gradient computation.
result Demonstrates the calibration of excess risk for linear-concave energies.
Paper introduces Fenchel-Young losses for supervised learning tasks.
problem Choosing the right loss function for supervised learning tasks.
method Introduces Fenchel-Young losses as a generic way to construct convex loss functions.
result Fenchel-Young losses unify and create new loss functions.
Generalizes robust loss functions for improved performance in various tasks.
problem Improving performance on tasks like registration and clustering.
method Introduces a continuous robustness parameter into loss functions, allowing them to be generalized.
result Improves performance on learning-based tasks like generative image synthesis and unsupervised depth estimation.
Improves GAN training with a repulsive loss function.
problem Discourages learning of fine details in data.
method Proposes a repulsive loss function and a bounded Gaussian kernel.
result Significantly improves GAN performance without additional computational cost.
Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
problem Improving loss functions for machine learning.
method Introduces Fitzpatrick losses based on the Fitzpatrick function.
result Fitzpatrick losses are tighter than Fenchel-Young losses.
A new loss function for VAEs improves image quality and efficiency.
problem Training VAEs to generate realistic images requires a loss function that reflects human perception.
method Based on Watson's perceptual model, the loss function computes a weighted distance in frequency space, accounts for luminance and contrast masking, and is extended to color images.
result VAEs trained with the new loss function generated high-quality, less blurry images with fewer artifacts and less computational resources.
New bounds for general unbounded loss functions, optimizing for various estimators.
problem Excess risk bounds for general unbounded loss functions, including log loss and squared loss.
method Optimized bounds for η-generalized Bayesian, MDL, and empirical risk minimization estimators, using v-GRIP and witness conditions. result Achieves ildeO(1/n) rates for certain loss functions under favorable v and small model complexity. This paper introduces new loss functions for balanced multi-class classification.
problem Balancing class imbalance in multi-class classification.
method Introduces two new surrogate loss families: GLA and GCA.
result GCA losses offer stronger theoretical guarantees in imbalanced settings.
Paper studies non-convex truncated loss functions for robust learning.
problem Improving generalization with non-convex loss functions.
method Truncating traditional loss functions and using SGD.
result Excess risk bounds and stationary points found by SGD.
Study risk bounds for distributed ERM with general loss functions and hypothesis spaces.
problem Limited theoretical analysis for distributed ERM with general loss functions and hypothesis spaces.
method Derive tight risk bounds under assumptions on hypothesis space and loss function.
result Developed more general risk bound for distributed ERM without strong convexity restriction.
Paper introduces a new robust loss function for RL.
problem Heuristic selection of threshold parameters in quantile Huber loss.
method Derived from Wasserstein distance, captures noise in quantile values.
result Enhances robustness against outliers and enables parameter adjustment.
A meta-learning method learns adaptive robust loss functions for noisy labels.
problem Handling robust learning with noisy labels and optimizing hyperparameters.
method Adaptive learning of robust loss hyperparameters through mutual improvement with network parameters.
result Generalized and effective robust loss functions with good generalization capability.
The paper proves deep learning can be robust with certain loss functions.
problem The robustness of deep learning models under flawed data.
method Empirical-risk minimization with unbounded, Lipschitz-continuous loss functions.
result These loss functions provide efficient prediction under minimal data assumptions.
Optimizes ensemble predictions for various loss functions.
problem Minimizing prediction loss on unlabeled data.
method Minimax optimal aggregation of binary classifiers for general loss functions.
result Family of efficient semi-supervised ensemble algorithms.
Paper improves stability analysis of SGD for various loss functions and data distributions.
problem Improving stability analysis of SGD for non-convex loss functions and data distributions.
method Analyzes stability of SGD for convex and non-convex loss functions, and improves data-dependent bounds.
result Improved stability bounds for non-convex loss functions and convex regularized loss functions.
New GAN loss functions improve image generation quality and stability.
problem Improving the performance of GANs in generating high-quality images.
method Introducing least kth-order GAN (LkGAN) and Rényi-centric GAN loss functions. result The proposed loss functions lead to better image quality and stability.
Tilting loss functions improves machine learning performance.
problem Improving machine learning models, especially in under- and over-parameterized networks.
method Using evolving loss functions that emphasize different classes cyclically.
result Dynamical loss functions lead to better generalization and stability in training.
Proposes a new Huber loss combining absolute and quadratic properties.
problem Improving robustness in learning models.
method Introduces a generalized Huber loss with a log-exp transform and provides an efficient minimization algorithm.
result Shows that the new loss function can be minimized efficiently.
Novel convex surrogate for non-modular loss functions.
problem Computational tractability for non-modular loss functions.
method Submodular-supermodular decomposition, slack-rescaling, Lov{á}sz hinge.
result First tractable solution for non-modular loss functions.
Proposes a new loss function for learning with noisy labels.
problem Improving model learnability with noisy labels.
method Uses generalized Jensen-Shannon divergence as a noise-robust loss function.
result Shows state-of-the-art results on noisy data.
Logitron combines Perceptron and logistic loss for improved classification.
problem Non-convex and non-smooth zero-one loss function in classification models.
method Introduces a Perceptron-augmented convex classification framework with an extended logistic loss function.
result Hinge-Logitron outperforms logistic regression and SVM in classification accuracy.
Proposes new loss functions for GANs to improve estimation accuracy and robustness.
problem Improving the training of GANs to achieve more accurate and robust models.
method Introduces Hellinger-type loss functions and analyzes their statistical properties.
result Demonstrates improved estimation accuracy and robustness of the proposed loss functions.
GANs can generate realistic data without minimizing a divergence, contrary to current theory.
problem Current theory suggests GANs minimize a divergence to generate realistic data.
method Discussed various loss functions for G, showing they are not divergences and do not have the same equilibrium.
result GANs can use a wide range of loss functions, not just divergences, to generate realistic data.
SoftAdapt dynamically adjusts loss weights for multi-part functions.
problem Slow convergence and poor weight selection for multi-part loss functions.
method SoftAdapt dynamically changes weights based on live performance statistics.
result Improved convergence and better weight selection for multi-part loss functions.
A new loss function α-loss improves classification robustness and calibration.
problem Improving classification robustness and calibration in machine learning.
method Introduces a tunable loss function α-loss, parameterized by α, and analyzes its theoretical and practical properties. result The α-loss function can improve model robustness to label flips and sensitivity to imbalanced classes. Classification and regression tasks in overparameterized models show different generalization properties.
problem Comparing classification and regression in overparameterized models.
method Comparison of least-squares minimum-norm interpolation and hard-margin SVM using different loss functions.
result Interpolating solutions generalize well with 0-1 loss but not with square loss.
The paper sets lower bounds for adversarial robustness in multiclass classification.
problem Adversarial robustness in multiclass classification with arbitrary loss functions.
method Dual and barycentric reformulations for robust risk minimization.
result Sharp lower bounds for adversarial risks are computed efficiently.
Proposes a new loss function for distributional learning.
problem Learning sparse and singular distributions.
method Entropy-regularized optimal transport and Fenchel duality.
result Geometric loss results in unconstrained convex objective functions.
New GE2E loss improves speaker verification efficiency and accuracy.
problem Improving speaker verification accuracy and efficiency.
method Proposed a new loss function (GE2E) that updates the network to emphasize difficult examples.
result Decreased EER by more than 10% and reduced training time by 60%.
New loss functions make deep nets robust to noisy labels.
problem Label noise in training data affects deep neural networks.
method Developed conditions for loss functions to be robust to label noise.
result Mean absolute value loss is inherently robust to label noise.
VIABLE learns a loss function for better few-shot learning.
problem Few-shot learning underfits with standard loss functions.
method Meta-learning to learn a differentiable loss function.
result Learning a relational loss function improves performance and sample efficiency.
This paper analyzes output activation functions for adversarial losses.
problem Understanding which output activation functions form a well-behaved adversarial loss.
method Variational divergence minimization and a comparative framework for adversarial losses.
result There is no single winning combination of output activation functions and regularization approaches across all settings.
Proposes a new loss function for deep neural networks.
problem Deep neural networks lack a direct method to discriminate between correct and competing classes.
method Introduces a discriminative loss function based on negative log likelihood ratio.
result Significantly outperforms cross-entropy loss on image classification tasks.
A novel dictionary-based approach for predicting functions.
problem Functional-output regression with non-orthogonal dictionaries.
method Projection learning (PL) with reproducing kernel Hilbert spaces (KPL).
result KPL offers a flexible and computationally efficient solution.
Theoretical analysis of cross-entropy loss functions and their robustness.
problem Guarantees for using cross-entropy as a surrogate loss function.
method Theoretical analysis of a broad family of loss functions, including cross-entropy.
result First H-consistency bounds for comp-sum losses and smooth adversarial comp-sum losses. The MM algorithm improves robust penalized estimation for outlier-contaminated data.
problem Outliers in data affect the reliability of penalized estimation.
method Innovative MM algorithm for both convex and nonconvex loss functions.
result Established convergence theory for MM algorithm with various loss functions.
New algorithm FWBoost avoids overfitting in boosting for general loss functions.
problem Overfitting in boosting, especially in regression settings.
method Frank-Wolfe type boosting algorithm (FWBoost) applied to general loss functions.
result FWBoost maintains good test performance even with many boosting rounds.
New method finds better loss functions for neural nets.
problem Finding effective loss functions for deep neural networks.
method Optimizes multivariate Taylor polynomial parameterizations using CMA-ES.
result TaylorGLO finds loss functions that outperform existing methods.
Improved robustness in kernel-based regression via novel loss function and IRLS.
problem Noise sensitivity in kernel-based regression methods.
method Proposed ℓs-loss function and iteratively reweighted least squares (IRLS) optimization. result Improved noise robustness in kernel-based regression methods.
SGLB boosts machine learning with Langevin diffusion for multimodal loss functions.
problem Dealing with multimodal loss functions in machine learning.
method Stochastic Gradient Langevin Boosting (SGLB) based on Langevin diffusion equation.
result SGLB guarantees global convergence for multimodal loss functions.
Proposes a general model for plane-based clustering with a new loss function.
problem Clustering of data points in a plane-based approach.
method General model containing existing methods, optimization problem with total loss minimization.
result The proposed method effectively clusters data points with a new loss function.