We obtain a tight distribution-specific characterization of the sample complexity of large-margin classification with L2 regularization: We introduce the margin-adapted dimension, which is a simple function of the second order statistics of the data distribution, and show distribution-specific upper and lower bounds on…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
New margin-based learning guarantees improve generalization bounds.
New probabilistic complexity measures for linear and kernel methods.
We obtain a tight distribution-specific characterization of the sample complexity of large-margin classification with L_2 regularization: We introduce the γ-adapted-dimension, which is a simple function of the spectrum of a distribution's covariance matrix, and show distribution-specific upper and lower bounds on the s…
New bound on neural network generalization error using geometric complexity.
Efficient algorithms improve learning of large-margin halfspaces.
Recent research has used margin theory to analyze the generalization performance for deep neural networks (DNNs). The existed results are almost based on the spectrally-normalized minimum margin. However, optimizing the minimum margin ignores a mass of information about the entire margin distribution, which is crucial …
3MSBM learns smooth trajectories from multiple snapshots.
This paper presents a margin-based multiclass generalization bound for neural networks that scales with their margin-normalized "spectral complexity": their Lipschitz constant, meaning the product of the spectral norms of the weight matrices, times a certain correction factor. This bound is empirically investigated for…
We consider the problem of adaptation to the margin and to complexity in binary classification. We suggest an exponential weighting aggregation scheme. We use this aggregation procedure to construct classifiers which adapt automatically to margin and complexity. Two main examples are worked out in which adaptivity is a…
We introduce a simple framework for designing private boosting algorithms. We give natural conditions under which these algorithms are differentially private, efficient, and noise-tolerant PAC learners. To demonstrate our framework, we use it to construct noise-tolerant and private PAC learners for large-margin halfspa…
We present new differentially private algorithms for learning a large-margin halfspace. In contrast to previous algorithms, which are based on either differentially private simulations of the statistical query model or on private convex optimization, the sample complexity of our algorithms depends only on the margin of…
This work proposes a new method to estimate joint probability from pairwise marginals, reducing sample complexity.
We derive and analyze a new, efficient, pool-based active learning algorithm for halfspaces, called ALuMA. Most previous algorithms show exponential improvement in the label complexity assuming that the distribution over the instance space is close to uniform. This assumption rarely holds in practical applications. Ins…
Efficiently estimates marginal posteriors for complex simulations.
Researchers expand on best subset selection theory, identifying key complexities.
Study shows a tradeoff between sample complexity and computational efficiency for learning halfspaces with random noise.
One of the main open problems in the theory of multi-category margin classification is the form of the optimal dependency of a guaranteed risk on the number C of categories, the sample size m and the margin parameter gamma. From a practical point of view, the theoretical analysis of generalization performance contribut…
Proposes a new distance metric for multi-marginal optimal transport.
Boosting and other ensemble methods combine a large number of weak classifiers through weighted voting to produce stronger predictive models. To explain the successful performance of boosting algorithms, Schapire et al. (1998) showed that AdaBoost is especially effective at increasing the margins of the training data. …
New method controls error in low-dimensional marginals of spatial models.
For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship is less clear: existing analyses of the output margin give complicated bounds whi…
Improves sampling from complex hierarchical models using HMC and automatic marginalization.
This work connects the Hessian to the decision boundary complexity in neural networks.
Frugal Flows learn complex data and infer marginal causal effects.
New algorithm learns halfspaces with near-optimal sample complexity in noisy conditions.
We consider a problem of risk estimation for large-margin multi-class classifiers. We propose a novel risk bound for the multi-class classification problem. The bound involves the marginal distribution of the classifier and the Rademacher complexity of the hypothesis class. We prove that our bound is tight in the numbe…
Proposes a tabular transformer model to maintain feature effect intelligibility.
A new approach for instance-optimal learning that bypasses impossibility results.
MSBM extends SB for multi-marginal trajectory inference.
With an eye toward understanding complexity control in deep learning, we study how infinitesimal regularization or gradient descent optimization lead to margin maximizing solutions in both homogeneous and non-homogeneous models, extending previous work that focused on infinitesimal regularization only in homogeneous mo…
Develops active learning method for linear optimization with margin-based criterion.
MDMA provides closed-form marginals and conditionals for deep networks.
Neural models improve GLMMs for complex data.
In this paper, we introduce the notion of a quasi-biharmonic submanifold in a pseudo-Riemannian manifold and classify quasi-biharmonic marginally trapped Lagrangian surfaces in Lorentzian complex space forms.
Gradient descent finds halfspaces with low error for agnostic learning.
Combines MCTM and NF for flexible multivariate density regression with interpretable marginals.
The paper improves robust optimization by introducing margin theory.
A classical condition for fast learning rates is the margin condition, first introduced by Mammen and Tsybakov. We tackle in this paper the problem of adaptivity to this condition in the context of model selection, in a general learning framework. Actually, we consider a weaker version of this condition that allows one…
Bayesian approach learns invariances from data alone, but last layer approximation is not always sufficient.
New framework for optimal transport with jumps over intermediate spaces.
Gradient descent and SGD achieve low test error in specific network weight regimes.
Improved exploration in RL with latent state marginalization.
Recurring international financial crises have adverse socioeconomic effects and demand novel regulatory instruments or strategies for risk management and market stabilization. However, the complex web of market interactions often impedes rational decisions that would absolutely minimize the risk. Here we show that, for…
We give polynomial-time algorithms for the exact computation of lowest-energy (ground) states, worst margin violators, log partition functions, and marginal edge probabilities in certain binary undirected graphical models. Our approach provides an interesting alternative to the well-known graph cut paradigm in that it …
Deep neural network (DNN) regression models are widely used in applications requiring state-of-the-art predictive accuracy. However, until recently there has been little work on accurate uncertainty quantification for predictions from such models. We add to this literature by outlining an approach to constructing predi…
Maximum entropy distributions with discrete support in dimensions arise in machine learning, statistics, information theory, and theoretical computer science. While structural and computational properties of max-entropy distributions have been extensively studied, basic questions such as: Do max-entropy distributio…