Self-paced learning selects tasks in a human-like progression for better multitask machine learning.
problem Improving multitask machine learning performance through effective task selection.
method Iterative selection of most appropriate tasks, learning task parameters, and updating shared knowledge using a bi-convex loss function.
result Self-paced task selection outperforms baseline methods in various multitask learning scenarios.
Paper tackles non-convex inf-projection problems with stochastic optimization.
problem Non-convex and possibly non-smooth inf-projection minimization problems.
method Developed stochastic algorithms for finding (nearly) stationary solutions.
result Established first-order convergence for non-convex inf-projection problems.
New algorithm recovers matrices that are both low rank and sparse in rows and columns.
problem Recovering matrices that are simultaneously low rank and row/column sparse.
method Gradient Descent with hard Thresholding (GDT) algorithm to minimize a bi-convex function over a nonconvex set of constraints.
result GDT achieves linear convergence to near optimal solutions with statistical error.
New method for completing binary matrices using machine learning theory.
problem Completing binary matrices with low-rank structure.
method Variational approximation of pseudo-posterior with convex relaxation.
result PAC-Bayesian learning bounds on prediction error.
Solves imaging inverse problems using a VAE prior and joint MAP optimization.
problem Solving ill-posed inverse problems in imaging.
method Joint Posterior Maximization with a VAE prior, using alternate optimization algorithms and stochastic encoding.
result Converges to high-quality solutions close to bi-convex, outperforming non-convex MAP approaches.
Solves imaging inverse problems using autoencoding priors.
problem Ill-posed inverse problems in imaging.
method Joint Posterior Maximization with Autoencoding Prior (JPMAP).
result JPMAP converges to a stationary point and provides robust solutions.
The paper improves multi-task learning by selecting variables and grouping tasks.
problem Improving generalization performance in multi-task learning.
method Factorizes a coefficient matrix into two matrices with sparsity for variable selection and overlapping group structure among tasks. Minimized using alternating optimization methods.
result Validated the effectiveness of the method on both synthetic and real-world datasets.
New findings on convexity of special Lagrangian geodesics.
problem Convexity of special Lagrangian geodesics in space-time.
method Space-time coordinate transformation preserving Lagrangian angle, leading to C2 estimate. result Subsolutions in all branches of the degenerate special Lagrangian equation are bi-convex.
DNCF framework recovers real scenes from imperfect images robustly.
problem Recovering real scenes from imperfect images.
method Nonparametric deep network that learns physical image formation equations.
result DNCF framework robustly defends against adversarial attacks.
New method infers graph from dependent matrix data.
problem Inferring graph from dependent matrix data.
method Sparse-group lasso-based frequency-domain formulation with ADMM approach.
result Local convergence of inverse PSD estimators to true value.
This work closes the theory-practice gap for distributed optimization methods by introducing a new regularity condition.
problem Existing convergence conditions for distributed optimization methods are violated by nearly all kernels used in practice.
method Introduces Hessian relative uniform continuity (HRUC) to guarantee convergence under mild conditions.
result Derives convergence guarantees for mirror descent-based gradient tracking without restrictive assumptions.
Model selects and generates music rules for realization and understanding.
problem Selecting and generating music rules that follow given constraints.
method Formulated as a bi-convex problem, derived efficient algorithm.
result Demonstrated model's effectiveness in music composition and analysis.
We develop a new model and algorithms for machine learning-based learning analytics, which estimate a learner's knowledge of the concepts underlying a domain, and content analytics, which estimate the relationships among a collection of questions and those concepts. Our model represents the probability that a learner p…
BLOCCS improves sparse CCA for better interpretation of multi-omics data.
problem Improving interpretation of multi-omics data.
method Block Sparse Canonical Correlation Analysis (BLOCCS) using a bi-convex objective and gradient descent.
result BLOCCS provides more interpretable solutions with improved orthogonality of sparse directions.
This paper tackles permutation recovery in unlabeled sensing from multiple measurement vectors.
problem Permutation recovery in unlabeled sensing from multiple measurement vectors.
method The paper studies the case of multiple noisy measurement vectors (MMVs) resulting from a common permutation and proposes computational schemes for permutation recovery.
result A large stable rank of the signal significantly reduces the required signal-to-noise ratio (SNR) for permutation recovery, and the problem can be solved efficiently using ADMM.
A new loss function α-loss bridges log-loss and 0-1 loss for binary classification.
problem Improving binary classification performance using a tunable loss function.
method Introducing α-loss, proving its margin-based form and classification-calibration, and providing an upper bound on empirical risk. result Empirical and expected risk difference upper bound for logistic regression-based classification.
Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
problem Improving loss functions for machine learning.
method Introduces Fitzpatrick losses based on the Fitzpatrick function.
result Fitzpatrick losses are tighter than Fenchel-Young losses.
We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise when margin losses can be proper composite losses, explicitly show how to determ…
Tamed Cross Entropy (TCE) loss outperforms standard CE loss in noisy classification tasks.
problem Improving classification performance in noisy data scenarios.
method Introducing Tamed Cross Entropy (TCE) loss, a derivative of Cross Entropy (CE) loss.
result TCE loss outperforms CE loss in all tested noisy classification scenarios.
Unified surrogate loss framework for multi-label learning with strong consistency guarantees.
problem Improving consistency and accounting for label correlations in multi-label learning.
method Introducing multi-label logistic loss and extending it to comprehensive multi-label comp-sum losses, proving strong consistency guarantees for any multi-label loss.
result Unified surrogate loss framework benefiting from strong consistency guarantees for any multi-label loss.
This paper introduces new loss functions for balanced multi-class classification.
problem Balancing class imbalance in multi-class classification.
method Introduces two new surrogate loss families: GLA and GCA.
result GCA losses offer stronger theoretical guarantees in imbalanced settings.
Visualizes basins of attraction for neural network loss functions.
problem Understanding the nature of neural network loss surfaces and basins of attraction.
method Gradient-based random sampling to visualize basins of attraction and stationary points.
result Entropic loss has a more searchable landscape with fewer stationary points than quadratic loss.
This work broadens calibeating to various proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses.
method Regret minimization and Bregman divergence approach.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
This work generalizes calibeating for a broader range of proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses beyond Brier and log loss.
method Regret minimization based on Bregman divergence for a family of proper losses.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
Proposes squentropy loss for improved classification accuracy and model calibration.
problem Theoretical and empirical evidence for cross-entropy loss is lacking.
method Introduces squentropy loss as the sum of cross-entropy and average square loss over incorrect classes.
result Squentropy loss outperforms cross-entropy and rescaled square losses in classification accuracy and model calibration.
New loss function calibrates WW-hinge loss for multiclass SVM.
problem WW-hinge loss not calibrated with 0-1 loss.
method Introduced ordered partition loss and proved WW-hinge loss is calibrated.
result WW-hinge loss is calibrated with ordered partition loss.
The study analyzes a model for aggregate losses with dependent and overdispersed inter-losses times.
problem Analyzing aggregate loss models with dependent and overdispersed inter-losses times.
method The study uses a two-state Markovian arrival process (MAP2) and a Markov renewal process to model the inter-losses times. Severities are modeled using a heavy-tailed, double-Pareto Lognormal distribution. The model is estimated via direct maximization of the likelihood function.
result The model with dependence and overdispersion in inter-losses times leads to higher capital charges compared to a Poisson process.
Logitron combines Perceptron and logistic loss for improved classification.
problem Non-convex and non-smooth zero-one loss function in classification models.
method Introduces a Perceptron-augmented convex classification framework with an extended logistic loss function.
result Hinge-Logitron outperforms logistic regression and SVM in classification accuracy.
Symmetric losses improve classifier robustness from corrupted labels.
problem Improving classifier performance from corrupted labels.
method Symmetric losses that satisfy a certain condition.
result Symmetric losses enhance robust classification from corrupted labels.
Two new algorithms improve performance in adversarial bandits with unbounded losses.
problem Adversarial Multi-Armed Bandits with unbounded losses.
method Developed UMAB-NN and UMAB-G for non-negative and general unbounded losses respectively.
result UMAB-NN achieves the first adaptive and scale-free regret bound for non-negative unbounded losses.
Paper explores connections between loss functions and consistency in binary classification and regression.
problem Consistency in binary classification and regression applications.
method Characterization of conformable loss functions and derivation of a new Huber-type loss function.
result Margin-based loss functions are equivalent to loss functions of squared standardized logistic regression residuals.
A new loss function combines advantages of average and max losses for better adaptability.
problem Improving supervised learning performance across various data distributions.
method Introducing average top-k loss, a convex function that averages the top-k losses. result The \atk loss function can adapt better to different data distributions and is convex.
Paper introduces a new topological loss for better convergence.
problem Optimizing topological losses for model's desired topological behavior.
method Introduces a new regularized topology-aware loss function.
result Guarantees efficient optimization of the new loss function.
Novel loss functions improve decision tree learning from noisy data.
problem Training decision trees with noisy labels.
method Introducing distribution losses and a new negative exponential loss.
result The negative exponential loss leads to efficient and robust decision tree learning.
Theoretical analysis of cross-entropy loss functions and their robustness.
problem Guarantees for using cross-entropy as a surrogate loss function.
method Theoretical analysis of a broad family of loss functions, including cross-entropy.
result First H-consistency bounds for comp-sum losses and smooth adversarial comp-sum losses. A new loss function α-loss improves classification robustness and calibration.
problem Improving classification robustness and calibration in machine learning.
method Introduces a tunable loss function α-loss, parameterized by α, and analyzes its theoretical and practical properties. result The α-loss function can improve model robustness to label flips and sensitivity to imbalanced classes. This paper improves operational risk modeling by selecting better loss severity distributions.
problem Inconsistent regulatory capital calculations due to changing loss severity distribution families.
method Presented truncation probability estimates and a consistent quantile scoring function for selection criteria. Also, recommended collecting loss frequencies below the minimum reporting threshold.
result More stable regulatory capital calculations through better selection of loss severity distributions.
This paper improves loss functions for deep learning with noisy labels.
problem Training deep neural networks with noisy labels.
method The paper introduces a normalization technique to make any loss function robust to noisy labels and proposes a framework called Active Passive Loss (APL) to combine robust loss functions.
result The proposed APL framework consistently outperforms state-of-the-art methods, especially under high noise rates.
Introduces MWLD to measure loss inequality across groups.
problem Machine learning's focus on average loss can lead to large group loss discrepancies.
method Defines MWLD, relates it to fairness and robustness, and provides estimation methods.
result MWLD can be estimated efficiently under certain weighting functions and reduces loss variance without significant accuracy loss.
We study cross-country GDP losses due to financial crises in terms of frequency (number of loss events per period) and severity (loss per occurrence). We perform the Loss Distribution Approach (LDA) to estimate a multi-country aggregate GDP loss probability density function and the percentiles associated to extreme eve…
The paper explores transferability of adversarial examples between convex and 01 loss models, finding non-transferability due to different decision boundaries caused by outliers.
problem Transferability of adversarial examples between convex and 01 loss models.
method Empirical study of transferability between linear 01 loss and convex (hinge) loss models, and between neural networks with different activation functions.
result Adversarial examples are non-transferable between convex and 01 loss models due to different decision boundaries caused by outliers.
Paper introduces Fenchel-Young losses for supervised learning tasks.
problem Choosing the right loss function for supervised learning tasks.
method Introduces Fenchel-Young losses as a generic way to construct convex loss functions.
result Fenchel-Young losses unify and create new loss functions.
The paper proves deep learning can be robust with certain loss functions.
problem The robustness of deep learning models under flawed data.
method Empirical-risk minimization with unbounded, Lipschitz-continuous loss functions.
result These loss functions provide efficient prediction under minimal data assumptions.
Symmetrizes loss functions to improve neural network robustness against noisy labels.
problem Designing robust loss functions for noisy labels in neural networks.
method Symmetrization of multi-class loss functions, focusing on cross-entropy and unhinged loss.
result The multi-class unhinged loss is the unique convex symmetric loss under suitable assumptions.
The study explores loss functions for learning distributions, finding the log loss and others are sufficient under certain conditions.
problem Understanding loss functions for distribution learning and density estimation.
method An axiomatic approach to design loss functions, proposing criteria and showing that no single loss function satisfies all criteria.
result No loss function satisfies all criteria, but the log loss and others do under the condition of candidate distributions being calibrated.
A new loss function improves classification accuracy in imbalanced datasets.
problem Suboptimal decision boundaries in classification with average, maximal, and average top-k losses. method Proposes a new classification objective called the close-k aggregate loss, which minimizes the loss for points close to the decision boundary. result Close-k aggregate loss achieves significant gains in 0-1 test accuracy compared to average, maximal, and average top-k losses. EnsLoss combines multiple loss functions to prevent overfitting in classification.
problem Preventing overfitting in classification models.
method EnsLoss is an ensemble method that combines loss functions, ensuring calibration and consistency.
result EnsLoss improves classification accuracy compared to fixed loss methods.
Study compares metric learning loss functions for speaker verification.
problem Comparing metric learning loss functions for end-to-end speaker verification.
method Cross entropy loss, cosine loss, angular margin loss, center loss, contrastive loss, triplet loss.
result Additive angular margin loss outperforms other loss functions.