Efficiently marginalizes over Gaussian Process kernels for better model flexibility and uncertainty.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Novel approach to OT using kernel mean embeddings controls overfitting and achieves dimension-free sample complexity.
New research shows the maximum ℓ1-margin classifier doesn't adapt to sparse ground truths.
Gaffke's bound is optimal for a specific parameter ordering in independent random vectors.
The problem of determining the joint probability distributions for correlated random variables with pre-specified marginals is considered. When the joint distribution satisfying all the required conditions is not unique, the "most unbiased" choice corresponds to the distribution of maximum entropy. The calculation of t…
The paper analyzes the maximum margin algorithm's performance on noisy data.
We study the connections between spectral clustering and the problems of maximum margin clustering, and estimation of the components of level sets of a density function. Specifically, we obtain bounds on the eigenvectors of graph Laplacian matrices in terms of the between cluster separation, and within cluster connecti…
Generative adversarial network for probabilistic forecasting of random systems.
We address the problem of learning the parameters in graphical models when inference is intractable. A common strategy in this case is to replace the partition function with its Bethe approximation. We show that there exists a regime of empirical marginals where such Bethe learning will fail. By failure we mean that th…
We give polynomial-time algorithms for the exact computation of lowest-energy (ground) states, worst margin violators, log partition functions, and marginal edge probabilities in certain binary undirected graphical models. Our approach provides an interesting alternative to the well-known graph cut paradigm in that it …
A new method estimates marginal likelihood using normalizing flows.
We obtain bounds on the distribution of the maximum of a martingale with fixed marginals at finitely many intermediate times. The bounds are sharp and attained by a solution to -marginal Skorokhod embedding problem in Obłój and Spoida [An iterated Azéma-Yor type embedding for finitely many marginals (2013) Preprint]…
In many real-world applications, data is not collected as one batch, but sequentially over time, and often it is not possible or desirable to wait until the data is completely gathered before analyzing it. Thus, we propose a framework to sequentially update a maximum margin classifier by taking advantage of the Maximum…
Bayesian models use hyperparameters to indirectly assign priors, and this work shows how these priors can be derived from maximum entropy principles.
Study shows how over-parameterized classifiers can still perform well on noisy data.
A framework estimates categorical distributions under constraints, ensuring generality and uniqueness.
This work analyzes the maximum-margin bias in quasi-homogeneous neural networks.
A new metric DJP-MMD improves domain adaptation by balancing transferability and discriminability.
Graphical models with bi-directed edges (<->) represent marginal independence: the absence of an edge between two vertices indicates that the corresponding variables are marginally independent. In this paper, we consider maximum likelihood estimation in the case of continuous variables with a Gaussian joint distributio…
Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…
The paper presents a method to estimate joint interventional distributions from marginal interventional data.
This work clarifies the role of inference types in planning.
A fast method for training linear classifiers maximizes margins.
New algorithm improves latent variable model estimation.
Mirror flow optimizes separable data problems, converging to a maximum margin classifier.
Optimizes risk measures given known marginal distributions of two unknown factors.
A new test compares latent variable models with kernel methods.
We consider two connected aspects of maximum likelihood estimation of the parameter for high-dimensional discrete graphical models: the existence of the maximum likelihood estimate (mle) and its computation. When the data is sparse, there are many zeros in the contingency table and the maximum likelihood estimate of th…
New method resolves nonidentifiability in mixture models.
We consider the problem of learning Bayesian network classifiers that maximize the marginover a set of classification variables. We find that this problem is harder for Bayesian networks than for undirected graphical models like maximum margin Markov networks. The main difficulty is that the parameters in a Bayesian ne…
We call a learner super-teachable if a teacher can trim down an iid training set while making the learner learn even better. We provide sharp super-teaching guarantees on two learners: the maximum likelihood estimator for the mean of a Gaussian, and the large margin classifier in 1D. For general learners, we provide a …
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
Novel proof shows continuity of optimal transport feasible set mapping.
The paper analyzes SBL pruning criteria under weakened assumptions.
Deep neural networks can generalize well even with perfect fits to noisy data.
The support vector machine (SVM) is an important class of learning machines for function approach, pattern recognition, and time-serious prediction, etc. It maps samples into the feature space by so-called support vectors of selected samples, and then feature vectors are separated by maximum margin hyperplane. The pres…
We solve the -marginal Skorokhod embedding problem for a continuous local martingale and a sequence of probability measures which are in convex order and satisfy an additional technical assumption. Our construction is explicit and is a multiple marginal generalisation of the Azema and Yor (1979) soluti…
New estimator reduces kernel mean estimation error.
Maximum entropy distributions with discrete support in dimensions arise in machine learning, statistics, information theory, and theoretical computer science. While structural and computational properties of max-entropy distributions have been extensively studied, basic questions such as: Do max-entropy distributio…
NDDV estimates data point value from a single stochastic trajectory.
Graphical models trained using maximum likelihood are a common tool for probabilistic inference of marginal distributions. However, this approach suffers difficulties when either the inference process or the model is approximate. In this paper, the inference process is first defined to be the minimization of a convex f…
New priors for deep neural networks converge to Gaussian processes.
A new method for unsupervised domain adaptation using Gaussian processes.
In binary classification problems, mainly two approaches have been proposed; one is loss function approach and the other is uncertainty set approach. The loss function approach is applied to major learning algorithms such as support vector machine (SVM) and boosting methods. The loss function represents the penalty of …
In this paper, we present a novel and general framework called {\it Maximum Entropy Discrimination Markov Networks} (MaxEnDNet), which integrates the max-margin structured learning and Bayesian-style estimation and combines and extends their merits. Major innovations of this model include: 1) It generalizes the extant …
We introduce a useful tool for analyzing boosting algorithms called the ``smooth margin function,'' a differentiable approximation of the usual margin for boosting algorithms. We present two boosting algorithms based on this smooth margin, ``coordinate ascent boosting'' and ``approximate coordinate ascent boosting,'' w…
The paper presents a new copula based method for measuring dependence between random variables. Our approach extends the Maximum Mean Discrepancy to the copula of the joint distribution. We prove that this approach has several advantageous properties. Similarly to Shannon mutual information, the proposed dependence mea…
The paper explores how benign overfitting occurs in heavy-tailed input distributions.