Paper improves deep neural networks' generalization by focusing on margin distribution complexity.
problem Improving deep neural networks' generalization performance.
method Proves a generalization upper bound based on margin distribution statistics and optimizes a convex margin distribution loss function.
result Optimizing the ratio of margin standard deviation to expected margin enhances generalization performance.
Deep models maximize minimum margin for high accuracy but decrease average margin, leading to poor robustness.
problem Inadequate balance between accuracy and robustness in deep model training.
method Analyzed the training process of deep models and proposed a new regularizer to promote average margin.
result Demonstrated an intrinsic trade-off between accuracy and robustness, and proposed a regularizer to improve robustness.
A new active learning method that improves model diversity and performance.
problem Efficiently selecting examples to label for training multiple models.
method Trains multiple models on bootstrap samples and selects examples based on minimum margin.
result Min-margin outperforms other methods, especially with larger batch sizes.
Associating distinct groups of objects (clusters) with contiguous regions of high probability density (high-density clusters), is central to many statistical and machine learning approaches to the classification of unlabelled data. We propose a novel hyperplane classifier for clustering and semi-supervised classificati…
We study the problem of identifying the causal relationship between two discrete random variables from observational data. We recently proposed a novel framework called entropic causality that works in a very general functional model but makes the assumption that the unobserved exogenous variable has small entropy in t…
While the channel capacity reflects a theoretical upper bound on the achievable information transmission rate in the limit of infinitely many bits, it does not characterise the information transfer of a given encoding routine with finitely many bits. In this note, we characterise the quality of a code (i. e. a given en…
New method uncovers zero entropy in dependent observations after finite samples.
problem Understanding uncertainty reduction in dependent observations.
method Minimum list entropy coupling, greedy algorithm.
result Zero entropy achieved with O(log(1/P_min)) samples for dependent observations.
A mesh-free method solves continuum-marginal optimal transport problems.
problem Recovering minimum-energy velocity fields from time-continuous probability marginals.
method Embeds weak continuity equation in a reproducing kernel Hilbert space, optimizing with mini-batch stochastic methods.
result Accurately recovers drift and maintains marginal consistency in synthetic experiments.
The paper analyzes boosting and minimum-ℓ1-norm classifiers in high dimensions.
problem Understanding the generalization error and optimal Bayes error in boosting.
method High-dimensional asymptotic theory, Gaussian comparison techniques, uniform deviation argument.
result Precise characterizations of boosting test error and optimal Bayes error.
Enhances ordinal embedding with less data by focusing on margin distribution.
problem Insufficient labeled data for ordinal embedding.
method Proposes Distributional Margin based Ordinal Embedding (DMOE) to improve generalization with less data.
result Demonstrates improved generalization performance with less labeled data.
Large margin approach for deep neural networks.
problem Deep learning's lack of margin enforcement.
method Proposes a novel loss function to enforce margin across layers of deep networks.
result Improved performance on various datasets and tasks.
Let X be a data matrix of rank ρ, whose rows represent n points in d-dimensional space. The linear support vector machine constructs a hyperplane separator that maximizes the 1-norm soft margin. We develop a new oblivious dimension reduction technique which is precomputed and can be applied to any input matrix X. We pr…
Solves a 60-year-old question on agreement measures in statistics.
problem The challenge of measuring agreement between two raters or measures.
method Developed a new algorithm to minimize diagonals in contingency tables, formulated the minimum feasible agreement, and studied the lower limit of maximum feasible agreement.
result Formulated the lower limit of Cohen's kappa and two statistics for agreement analysis.
New method calculates partial information for Gaussian systems based on dependency constraints.
problem Quantifying information sharing in multivariate Gaussian systems.
method Constructing maximum entropy models based on dependency constraints and deriving closed-form solutions.
result Closed-form solutions for Gaussian systems show differences in redundancy and synergy estimates compared to existing methods.
Model examines how margin trading affects stock market stability.
problem Margin trading's impact on stock market stability.
method Cascading failure model based on bipartite graph of investors and shares.
result Margin trading increases share price vulnerability to external shocks.
We consider an on-line system identification setting, in which new data become available at given time steps. In order to meet real-time estimation requirements, we propose a tailored Bayesian system identification procedure, in which the hyper-parameters are still updated through Marginal Likelihood maximization, but …
Sample size determination for a data set is an important statistical process for analyzing the data to an optimum level of accuracy and using minimum computational work. The applications of this process are credible in every domain which deals with large data sets and high computational work. This study uses Bayesian a…
We consider the problem of identifying the causal direction between two discrete random variables using observational data. Unlike previous work, we keep the most general functional model but make an assumption on the unobserved exogenous variable: Inspired by Occam's razor, we assume that the exogenous variable is sim…
Compact neural network reduces portfolio variance by 90% with higher leverage.
problem Minimizing portfolio variance under aggressive leverage constraints.
method Modular neural network with reduced parameters and new moving average.
result Achieves lowest realized portfolio variance with higher leverage.
LoCoV reduces portfolio optimization errors from sample covariance matrices.
problem Large errors in sample covariance matrix for optimal portfolio weights.
method LoCoV (low dimension covariance voting) algorithm to reduce these errors.
result LoCoV outperforms classical methods in portfolio optimization experiments.
A Kronecker product model is the set of visible marginal probability distributions of an exponential family whose sufficient statistics matrix factorizes as a Kronecker product of two matrices, one for the visible variables and one for the hidden variables. We estimate the dimension of these models by the maximum rank …
This work connects the Hessian to the decision boundary complexity in neural networks.
problem Understanding the decision boundary complexity in high-dimensional input space.
method Characterizing the decision boundary using the Hessian top eigenvectors and analyzing the number of outliers.
result The number of outliers in the Hessian spectrum is proportional to the complexity of the decision boundary.
The paper proves a margin inequality for separating hyperplanes, useful for analyzing algorithmic bias.
problem Analyzing the implicit bias of algorithms in machine learning.
method Proves a nonsmooth Kurdyka-Lojasiewicz inequality for margin function.
result The bias of algorithm iterates converges at least as fast as the square-root of the margin convergence rate.
Paper certifies intersection of minimum-volume confidence sets for multinomial outcomes.
problem Certifying intersection of minimum-volume confidence sets for multinomial outcomes.
method Exploits likelihood ordering to induce halfspace constraints, enabling adaptive geometric partitioning and computable bounds on p-values.
result Efficient and provably sound algorithm for certifying intersection, disjointness, or indeterminate result.
Algorithm samples from Wasserstein barycenter of measures.
problem Sampling from Wasserstein barycenter of measures.
method Gradient flow of multimarginal formulation with penalization.
result Algorithm samples close to Wasserstein barycenter.
Study improves speaker verification accuracy using angular based embedding learning.
problem Improving discriminative power of embeddings for open-set speaker verification.
method Optimizes angular distance and adds margin penalty, applying various angular margin embedding strategies and proposing inter-class regularization.
result Achieved impressive results with 16.5% improvement in EER and 18.2% improvement in minimum detection cost function.
We propose a non-parametric anomaly detection algorithm for high dimensional data. We first rank scores derived from nearest neighbor graphs on n-point nominal training data. We then train limited complexity models to imitate these scores based on the max-margin learning-to-rank framework. A test-point is declared as…
Develops a neural network for global minimum variance portfolio optimization.
problem Minimizing portfolio variance for large equity covariance matrices.
method Rotation-invariant neural network that learns lag-transformed returns and covariance regularization.
result End-to-end trained model outperforms competitors in realized volatility and Sharpe ratios.
We address the problem of computing approximate marginals in Gaussian probabilistic models by using mean field and fractional Bethe approximations. As an extension of Welling and Teh (2001), we define the Gaussian fractional Bethe free energy in terms of the moment parameters of the approximate marginals and derive an …
Bounds on the log partition function are important in a variety of contexts, including approximate inference, model fitting, decision theory, and large deviations analysis. We introduce a new class of upper bounds on the log partition function, based on convex combinations of distributions in the exponential domain, th…
We give two provably accurate feature-selection techniques for the linear SVM. The algorithms run in deterministic and randomized time respectively. Our algorithms can be used in an unsupervised or supervised setting. The supervised approach is based on sampling features from support vectors. We prove that the margin i…
Gradient descent dynamics in deep networks leads to optimal margin solutions.
problem Controlling the complexity of deep networks for generalization.
method Gradient descent dynamics on normalized weights.
result Gradient descent dynamics converge to optimal margin solutions.
Max-margin classifiers' behavior is studied in high dimensions with non-Gaussian features.
problem Understanding the role of featurization maps and high-dimensional misclassification error.
method High-dimensional asymptotics, Gaussian model, support vector representation.
result Asymptotic behavior of max-margin classifiers is determined by feature covariance and label covariance.
We provide a set of copulas that can be interpreted as having the negative extreme dependence. This set of copulas is interesting because it coincides with countermonotonic copula for a bivariate case, and more importantly, is shown to be minimal in concordance ordering in the sense that no copula exists which is stric…
A new multi-class active learning method combining informativeness and representativeness.
problem Efficiently labeling large datasets with limited resources.
method A hybrid informative and representative criterion approach for multi-class active learning.
result The proposed method outperforms state-of-the-art methods on multiple UCI datasets.
Hard to learn ReLU with Gaussian data, but can approximate efficiently.
problem Learning a ReLU with Gaussian marginals under arbitrary labels.
method Proved hardness and developed an efficient approximation algorithm.
result Efficient approximation algorithm for best-fitting ReLU with error O(opt2/3). The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that the normalized maximum likelihood has a Bayes-li…
New algorithm finds minimum weight norm solutions in deep neural networks.
problem Training over-parameterized deep neural networks efficiently and improving generalization.
method Minnorm training method that minimizes the sum of weights' norms while fitting training data.
result Faster convergence to minimum-norm solutions and better generalization performance.
Tropical SVM tackles phylogenomics by classifying multi-locus data.
problem Classifying multi-locus data sets for phylogenetic analysis.
method Proposes tropical support vector machines (SVMs) for phylogenomics, formulated as linear programming problems.
result Developed methods for hard and soft margin tropical SVMs, proving necessary and sufficient conditions for separation.
Gradient noise improves privacy-protected optimization performance.
problem Improving privacy in convex optimization while maintaining utility.
method We analyze the effect of gradient perturbation on differentially private convex optimization, focusing on expected curvature.
result Gradient perturbation can achieve a significantly improved utility guarantee for differentially private convex optimization.
Paper analyzes mistake and generalization of MNIC classifiers.
problem Understanding the performance of interpolating classifiers.
method Elementary analyses of MNIC's regret and generalization.
result MNIC generalizes with a rate proportional to the norm of the interpolating solution and inversely proportional to the number of data points.
Improves interpretability of anomaly scores in GBRBM-based detection.
problem Difficulty in setting a proper threshold for anomaly scores.
method Proposes a measure based on cumulative distribution and uses simulated annealing for evaluation.
result Established a guideline for setting the threshold using the interpretable measure.
This paper proposes a new AL method that directly uses geometric sampling over clusters.
problem Performance degeneration in uncertainty evaluation for AL with insufficient labeled data.
method Divide-and-conquer approach to AL, transferring it to geometric sampling over clusters.
result The proposed GAL method significantly outperforms state-of-the-art baselines.
Extends wealth tax neutrality framework to stochastic volatility and non-homothetic preferences.
problem Ensuring wealth taxes are neutral under various economic conditions.
method Extended Frøseth's neutrality framework to stochastic volatility and non-homothetic preferences, identified four channels of non-neutrality, and applied the framework to global minimum wealth taxes.
result Non-uniform assessment, general equilibrium effects, progressive thresholds, and endogenous labour supply can cause non-neutrality under CRRA preferences.
The paper calculates MES bounds for systemic risk contributions under uncertain dependence.
problem Measuring systemic risk contributions of financial firms under uncertainty in dependence structure.
method Derives worst-case and best-case bounds for MES under known individual firm risks and partial dependence information.
result Improved MES bounds derived for various types of dependence models.
LMs perform poorly in true few-shot learning without held-out examples.
problem Evaluating few-shot performance of language models without access to held-out examples.
method Evaluated two model selection criteria (cross-validation and minimum description length) for choosing LM prompts and hyperparameters in true few-shot learning.
result Selection criteria often prefer models that perform worse than random selection, suggesting overestimation of few-shot ability.
We consider the pricing and hedging of exotic options in a model-independent set-up using \emph{shortfall risk and quantiles}. We assume that the marginal distributions at certain times are given. This is tantamount to calibrating the model to call options with discrete set of maturities but a continuum of strikes. In …
New tensor formulation reveals gradient flow's bias in linear neural networks.
problem Understanding implicit bias in linear neural network training.
method Tensor formulation of neural networks, including fully-connected, diagonal, and convolutional networks.
result Gradient flow on linear tensor networks converges to solutions of specific optimization problems.