Batch normalization biases linear models towards uniform margins, improving performance in binary classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SGD converges to critical points of normalized margin in late-stage training for homogeneous neural networks.
We prove that the marginal densities of a global probability mass function in a primal normal factor graph and the corresponding marginal densities in the dual normal factor graph are related via local mappings. The mapping depends on the Fourier transform of the local factors of the models. Details of the mapping, inc…
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
Paper improves normalizing flows to better capture distribution tails.
Local mappings relate dual and primal factor graphs for efficient marginal probability estimation.
A new method estimates marginal likelihood using normalizing flows.
This paper presents a margin-based multiclass generalization bound for neural networks that scales with their margin-normalized "spectral complexity": their Lipschitz constant, meaning the product of the spectral norms of the weight matrices, times a certain correction factor. This bound is empirically investigated for…
COMET Flows model multivariate extremes with heavy tails and asymmetric dependence.
Recent research has used margin theory to analyze the generalization performance for deep neural networks (DNNs). The existed results are almost based on the spectrally-normalized minimum margin. However, optimizing the minimum margin ignores a mass of information about the entire margin distribution, which is crucial …
Combines MCTM and NF for flexible multivariate density regression with interpretable marginals.
A spacelike surface is marginally trapped if its mean curvature vector is lightlike. On any oriented spacelike surface we show that a choice of orientation of the normal bundle determines a smooth map which we call the null Gauss map of…
Principal component analysis (PCA) is arguably the most popular tool in multivariate exploratory data analysis. In this paper, we consider the question of how to handle heterogeneous variables that include continuous, binary, and ordinal. In the probabilistic interpretation of low-rank PCA, the data has a normal multiv…
Study examines liquidation, leverage, and optimal margin requirements in Bitcoin futures markets.
In this paper, we study the implicit regularization of the gradient descent algorithm in homogeneous neural networks, including fully-connected and convolutional neural networks with ReLU or LeakyReLU activations. In particular, we study the gradient descent or gradient flow (i.e., gradient descent with infinitesimal s…
Improved RCPs for MABs using normalized weight functions.
This work analyzes the maximum-margin bias in quasi-homogeneous neural networks.
For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship is less clear: existing analyses of the output margin give complicated bounds whi…
Proposes a new portfolio optimization method considering reward, dispersion, and asymmetry.
Evidential Softmax preserves multimodality in sparse probability distributions for generative models.
GPDFlow models extreme threshold exceedance with flexible dependence using normalizing flows.
New insights into deep learning: reducing training data significantly improves performance.
In this work, we investigate Batch Normalization technique and propose its probabilistic interpretation. We propose a probabilistic model and show that Batch Normalization maximazes the lower bound of its marginalized log-likelihood. Then, according to the new probabilistic model, we design an algorithm which acts cons…
We consider the combined use of resampling and partial rejection control in sequential Monte Carlo methods, also known as particle filters. While the variance reducing properties of rejection control are known, there has not been (to the best of our knowledge) any work on unbiased estimation of the marginal likelihood …
A fast method for training linear classifiers maximizes margins.
Normalized compound random measures are flexible nonparametric priors for related distributions. We consider building general nonparametric regression models using normalized compound random measure mixture models. Posterior inference is made using a novel pseudo-marginal Metropolis-Hastings sampler for normalized comp…
We propose a novel and flexible rank-breaking-then-composite-marginal-likelihood (RBCML) framework for learning random utility models (RUMs), which include the Plackett-Luce model. We characterize conditions for the objective function of RBCML to be strictly log-concave by proving that strict log-concavity is preserved…
The paper analyzes the maximum margin algorithm's performance on noisy data.
A fast method estimates group-adaptive elastic net penalties using co-data.
The paper optimizes portfolios using relative tail risk measures.
Theoretical justification for deep networks' performance with regularization techniques.
This work improves motion planning for quadcopters by learning and reasoning about controller performance.
New bound on neural network generalization error using geometric complexity.
Develops bounds predicting deep learning generalization using optimal transport.
We investigate the class of -stable Poisson-Kingman random probability measures (RPMs) in the context of Bayesian nonparametric mixture modeling. This is a large class of discrete RPMs which encompasses most of the the popular discrete RPMs used in Bayesian nonparametrics, such as the Dirichlet process, Pitman-Yor p…
The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that the normalized maximum likelihood has a Bayes-li…
A new method generates synthetic data with realistic marginal distributions.
As shown in recent research, deep neural networks can perfectly fit randomly labeled data, but with very poor accuracy on held out data. This phenomenon indicates that loss functions such as cross-entropy are not a reliable indicator of generalization. This leads to the crucial question of how generalization gap should…
Normalizing constant (also called partition function, Bayesian evidence, or marginal likelihood) is one of the central goals of Bayesian inference, yet most of the existing methods are both expensive and inaccurate. Here we develop a new approach, starting from posterior samples obtained with a standard Markov Chain Mo…
Develops methods for constructing parameter priors in DAG models.
In a recent paper, Eichmair, Galloway and Pollack have proved a Gannon-Lee-type singularity theorem based on the existence of marginally outer trapped surfaces (MOTS) on noncompact initial data sets for globally hyperbolic spacetimes. However, one might wonder whether the corresponding incomplete geodesics could still …
Unified framework for efficient trans-dimensional Bayesian inference using VI and NFs.
The paper develops a new model-free formula for option initial margins.
A new normalizing flow models continuous stochastic processes efficiently.
The inference of correlated signal fields with unknown correlation structures is of high scientific and technological relevance, but poses significant conceptual and numerical challenges. To address these, we develop the correlated signal inference (CSI) algorithm within information field theory (IFT) and discuss its n…
Proposes logistic-beta process for modeling dependent probabilities with beta marginals.
We consider the problem of learning Bayesian network classifiers that maximize the marginover a set of classification variables. We find that this problem is harder for Bayesian networks than for undirected graphical models like maximum margin Markov networks. The main difficulty is that the parameters in a Bayesian ne…
The generalization error of deep neural networks via their classification margin is studied in this work. Our approach is based on the Jacobian matrix of a deep neural network and can be applied to networks with arbitrary non-linearities and pooling layers, and to networks with different architectures such as feed forw…