We analyze bias-variance of margin losses.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Support vector regression (SVR) is one of the most popular machine learning algorithms aiming to generate the optimal regression curve through maximizing the minimal margin of selected training samples, i.e., support vectors. Recent researchers reveal that maximizing the margin distribution of whole training dataset ra…
We present an improved algorithm for {\em quasi-properly} learning convex polyhedra in the realizable PAC setting from data with a margin. Our learning algorithm constructs a consistent polyhedron as an intersection of about halfspaces with constant-size margins in time polynomial in (where is the nu…
Graphical models trained using maximum likelihood are a common tool for probabilistic inference of marginal distributions. However, this approach suffers difficulties when either the inference process or the model is approximate. In this paper, the inference process is first defined to be the minimization of a convex f…
Paper proposes a new classifier for hyperbolic spaces using horospherical boundaries.
Consider a classification problem where we have both labeled and unlabeled data available. We show that for linear classifiers defined by convex margin-based surrogate losses that are decreasing, it is impossible to construct any semi-supervised approach that is able to guarantee an improvement over the supervised clas…
New algorithms recover clusters with minimal queries, connecting margins to recoverability.
Study minimax rates for binary classifier estimation with margin conditions.
This paper tackles multi-marginal optimal transport problems using DC programming.
Existence proved for -Bass martingales with specific marginals.
Recent research has used margin theory to analyze the generalization performance for deep neural networks (DNNs). The existed results are almost based on the spectrally-normalized minimum margin. However, optimizing the minimum margin ignores a mass of information about the entire margin distribution, which is crucial …
We present new differentially private algorithms for learning a large-margin halfspace. In contrast to previous algorithms, which are based on either differentially private simulations of the statistical query model or on private convex optimization, the sample complexity of our algorithms depends only on the margin of…
A scalable algorithm approximates Wasserstein Barycenters using neural networks.
Algorithm learns dynamics from past observations.
Paper analyzes GMM for separable data with various parameter structures.
A new method for optimal transport using neural ODEs that preserves marginal constraints.
The popular Lasso approach for sparse estimation can be derived via marginalization of a joint density associated with a particular stochastic model. A different marginalization of the same probabilistic model leads to a different non-convex estimator where hyperparameters are optimized. Extending these arguments to pr…
We address the problem of aggregating an ensemble of predictors with known loss bounds in a semi-supervised binary classification setting, to minimize prediction loss incurred on the unlabeled data. We find the minimax optimal predictions for a very general class of loss functions including all convex and many non-conv…
We propose in this paper a general framework for deriving loss functions for structured prediction. In our framework, the user chooses a convex set including the output space and provides an oracle for projecting onto that set. Given that oracle, our framework automatically generates a corresponding convex and smooth l…
It is well known that a random vector with given marginal distributions is comonotonic if and only if it has the largest sum with respect to the convex order [ Kaas, Dhaene, Vyncke, Goovaerts, Denuit (2002), A simple geometric proof that comonotonic risks have the convex-largest sum, ASTIN Bulletin 32, 71-80. Cheung (2…
Marginal MAP inference involves making MAP predictions in systems defined with latent variables or missing information. It is significantly more difficult than pure marginalization and MAP tasks, for which a large class of efficient and convergent variational algorithms, such as dual decomposition, exist. In this work,…
A new relaxed framework for pricing illiquid derivatives using bid-ask spreads.
This paper characterizes the equilibrium in a continuous time financial market populated by heterogeneous agents who differ in their rate of relative risk aversion and face convex portfolio constraints. The model is studied in an application to margin constraints and found to match real world observations about financi…
Proposes MGCE for improved classification performance.
Researchers find a way to price American options without relying on specific asset price models.
We carefully study how well minimizing convex surrogate loss functions, corresponds to minimizing the misclassification error rate for the problem of binary classification with linear predictors. In particular, we show that amongst all convex surrogate losses, the hinge loss gives essentially the best possible bound, o…
The Skorokhod embedding problem aims to represent a given probability measure on the real line as the distribution of Brownian motion stopped at a chosen stopping time. In this paper, we consider an extension of the optimal Skorokhod embedding problem to the case of finitely-many marginal constraints. Using the classic…
A fast method for training linear classifiers maximizes margins.
New algorithm learns halfspaces with near-optimal sample complexity in noisy conditions.
Data augmentation (DA) is commonly used during model training, as it significantly improves test error and model robustness. DA artificially expands the training set by applying random noise, rotations, crops, or even adversarial perturbations to the input data. Although DA is widely used, its capacity to provably impr…
Non-convex SGD learns halfspaces with adversarial label noise efficiently.
We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise when margin losses can be proper composite losses, explicitly show how to determ…
Worst-case bounds on the expected shortfall risk given only limited information on the distribution of the random variables has been studied extensively in the literature. In this paper, we develop a new worst-case bound on the expected shortfall when the univariate marginals are known exactly and additional expert inf…
Support vector machine (SVM) has attracted great attentions for the last two decades due to its extensive applications, and thus numerous optimization models have been proposed. To distinguish all of them, in this paper, we introduce a new model equipped with an soft-margin loss (dubbed as -SVM) whic…
GD at EoS edge minimizes logistic loss without monotonic convergence.
We solve the -marginal Skorokhod embedding problem for a continuous local martingale and a sequence of probability measures which are in convex order and satisfy an additional technical assumption. Our construction is explicit and is a multiple marginal generalisation of the Azema and Yor (1979) soluti…
Estimate collapsibility of causal effects in CPDAGs via strong d-convex hulls.
We consider the problem of covariance matrix estimation in the presence of latent variables. Under suitable conditions, it is possible to learn the marginal covariance matrix of the observed variables via a tractable convex program, where the concentration matrix of the observed variables is decomposed into a sparse ma…
This paper studies Fenchel-Young losses, a generic way to construct convex loss functions from a regularization function. We analyze their properties in depth, showing that they unify many well-known loss functions and allow to create useful new ones easily. Fenchel-Young losses constructed from a generalized entropy, …
We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new distances are metrics whenever the ground distance between the marginals is a metric, generalize both…
New study reveals a polynomial penalty for adapting to unknown margin parameters in batched nonparametric bandits.
Study shows uniform-time chaos propagation in mean field Langevin dynamics.
This paper focuses on martingale optimal transport problems when the martingales are assumed to have bounded quadratic variation. First, we give a result that characterizes the existence of a probability measure satisfying some convex transport constraints in addition to having given initial and terminal marginals. Sev…
Recent research has made significant progress on the problem of bounding log partition functions for exponential family graphical models. Such bounds have associated dual parameters that are often used as heuristic estimates of the marginal probabilities required in inference and learning. However these variational est…
Gradient descent finds halfspaces with low error for agnostic learning.
The aims of this study are twofold. First, we consider an optimal risk allocation problem with non-convex preferences. By establishing an infimal representation for distortion risk measures, we give some necessary and sufficient conditions for the existence of optimal and asymptotic optimal allocations. We will show th…
Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…
New method controls linear systems with adversarial disturbances.