Uncertainty sampling, a popular active learning algorithm, is used to reduce the amount of data required to learn a classifier, but it has been observed in practice to converge to different parameters depending on the initialization and sometimes to even better parameters than standard training on all the data. In this…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Boosted CVaR Classification improves tail performance in classification tasks.
Introduces a differentiable approximation to the zero-one loss.
Adversarial training achieves optimal test error for shallow networks.
Gradient descent finds halfspaces with low error for agnostic learning.
Study of estimation errors in surrogate loss minimizers, providing stronger guarantees than existing methods.
In this work, we addressed the issue of applying a stochastic classifier and a local, fuzzy confusion matrix under the framework of multi-label classification. We proposed a novel solution to the problem of correcting label pairwise ensembles. The main step of the correction procedure is to compute classifier- specific…
This manuscript provides optimization guarantees, generalization bounds, and statistical consistency results for AdaBoost variants which replace the exponential loss with the logistic and similar losses (specifically, twice differentiable convex losses which are Lipschitz and tend to zero on one side). The heart of the…
We define two types of local indices of a vector field at an isolated zero on the boundary, and prove Poincare-Hopf-type index theorems for certain vector fields on a compact smooth manifold which have only isolated zeros.
There is growing evidence that converting targets to soft targets in supervised learning can provide considerable gains in performance. Much of this work has considered classification, converting hard zero-one values to soft labels---such as by adding label noise, incorporating label ambiguity or using distillation. In…
Gaptron algorithm reduces mistakes in online multiclass classification.
Symmetric losses improve classifier robustness from corrupted labels.
Enhanced -consistency bounds derived under relaxed conditions.
Analyzes SGD dynamics on multi-class problems with exact expressions.
Binary classification rules based on covariates typically depend on simple loss functions such as zero-one misclassification. Some cases may require more complex loss functions. For example, individual-level monitoring of HIV-infected individuals on antiretroviral therapy (ART) requires periodic assessment of treatment…
Proposes measures for uncertainty quantification using proper scoring rules.
Class ambiguity is typical in image classification problems with a large number of classes. When classes are difficult to discriminate, it makes sense to allow k guesses and evaluate classifiers based on the top-k error instead of the standard zero-one loss. We propose top-k multiclass SVM as a direct method to optimiz…
Classification is the most important process in data analysis. However, due to the inherent non-convex and non-smooth structure of the zero-one loss function of the classification model, various convex surrogate loss functions such as hinge loss, squared hinge loss, logistic loss, and exponential loss are introduced. T…
Bayes-consistent disagreement discrepancy loss improves model robustness.
Label aggregation makes learning robust to noisy labels.
Theoretical analysis of cross-entropy loss functions and their robustness.
The F-measure, which has originally been introduced in information retrieval, is nowadays routinely used as a performance metric for problems such as binary classification, multi-label classification, and structured output prediction. Optimizing this measure is a statistically and computationally challenging problem, s…
The projection of a compact oriented submanifold M^{n-1} in R^{n+1} on a hyperplane P^{n} can fail to bound any region in P. We call this ``projecting to zero.'' Example: The equatorial S^1 in S^2 projects to zero in any plane containing the x_3-axis. Using currents to make this precise, we show: A lipschitz (homology)…
We study the Stochastic Gradient Langevin Dynamics (SGLD) algorithm for non-convex optimization. The algorithm performs stochastic gradient descent, where in each step it injects appropriately scaled Gaussian noise to the update. We analyze the algorithm's hitting time to an arbitrary subset of the parameter space. Two…
The paper argues that uncertainty quantification in ML is application-specific and proposes a flexible family of measures.
A censored transformed model for proportional outcomes with boundary mass and an application to loss given default modeling.
New algorithms minimize PAC-Bayesian C-Bound for majority voting, leading to scalable and accurate predictors.
Principal Component Analysis (PCA) is a very successful dimensionality reduction technique, widely used in predictive modeling. A key factor in its widespread use in this domain is the fact that the projection of a dataset onto its first principal components minimizes the sum of squared errors between the original …
Paper develops tighter risk certificates for contrastive learning models.
The predictive quality of machine learning models is typically measured in terms of their (approximate) expected prediction accuracy or the so-called Area Under the Curve (AUC). Minimizing the reciprocals of these measures are the goals of supervised learning. However, when the models are constructed by the means of em…
How initialization and loss function affect the learning of a deep neural network (DNN), specifically its generalization error, is an important problem in practice. In this work, by exploiting the linearity of DNN training dynamics in the NTK regime \citep{jacot2018neural,lee2019wide}, we provide an explicit and quanti…
The paper revisits discriminative vs. generative classifiers, showing naive Bayes requires fewer samples.
Differentiable Architecture Search (DARTS) is now a widely disseminated weight-sharing neural architecture search method. However, it suffers from well-known performance collapse due to an inevitable aggregation of skip connections. In this paper, we first disclose that its root cause lies in an unfair advantage in exc…
We show that a smooth complex projective threefold admits a holomorphic one-form without zeros if and only if the underlying real 6-manifold fibres smoothly over the circle, and we give a complete classification of all threefolds with that property. Our results prove a conjecture of Kotschick in dimension three.
Study on Privileged ERM showing limitations and providing capacity analysis.
New research shows logistic regression can achieve optimal error rate for agnostic learning of halfspaces.
Data-dependent hashing has recently attracted attention due to being able to support efficient retrieval and storage of high-dimensional data such as documents, images, and videos. In this paper, we propose a novel learning-based hashing method called "Supervised Discrete Hashing with Relaxation" (SDHR) based on "Super…
Ordinal regression is aimed at predicting an ordinal class label. In this paper, we consider its semi-supervised formulation, in which we have unlabeled data along with ordinal-labeled data to train an ordinal regressor. There are several metrics to evaluate the performance of ordinal regression, such as the mean absol…
The study counts curves on a once-punctured torus with self-intersections.
We classify the rational differential 1-forms with simple poles and simple zeros on the Riemann sphere according to their isotropy group; when the 1-form has exactly two poles the isotropy group is isomorphic to , namely , and when the 1-form has …
We present in this paper a -metric on an open neighbourhood of the origin in $\RR^{5}$. The metric is of Lorentzian signature and admits a solution to the twistor equation for spinors with a unique isolated zero at the origin. The metric is not conformally flat in any neighbourhood of the origin. The const…
Lower bounds on Bayes risk for realizable models derived using information theory.
We show that if an asymptotically flat manifold with horizon boundary admits a global static potential, then the static potential must be zero on the boundary. We also show that if an asymptotically flat manifold with horizon boundary admits an unbounded static potential in the exterior region, then the manifold must c…
Legendrian Lavrentiev links are shown to be equivalent to smooth links.
We prove a new upper bound for the first eigenvalue of the Dirac operator of a compact hypersurface in any Riemannian spin manifold carrying a non-trivial twistor spinor without zeros on the hypersurface. The upper bound is expressed as the first eigenvalue of a drifting Schrödinger operator on the hypersurface. Moreov…
Alternative model predicts health insurance reimbursement based on contract limitations.
Study proves Kotschick's conjecture for certain compact Kähler manifolds.
Let be a smooth compact Riemannian surface with no boundary. Given a smooth vector field with finitely many zeroes on , we study the distribution of the number of tangencies to of the nodal components of random band-limited functions. It is determined that in the high-energy limit, these obey a unive…