Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates marginal independence structure of Bayesian networks from data.
Graphical models with bi-directed edges (<->) represent marginal independence: the absence of an edge between two vertices indicates that the corresponding variables are marginally independent. In this paper, we consider maximum likelihood estimation in the case of continuous variables with a Gaussian joint distributio…
DIET tests conditional independence using marginal dependence measures of residual information.
New test for point processes without strong model assumptions.
New DP algorithms with margin guarantees for various hypothesis sets.
CMRFs extend PGMs for topological data, capturing both conditional and marginal dependencies.
In this article we discuss the problem of calculating optimal model-independent (robust) bounds for the price of Asian options with discrete and continuous averaging. We will give geometric characterisations of the maximising and the minimising pricing model for certain types of Asian options in discrete and continuous…
The paper examines how heavy-tailed risks behave under Gaussian copula models.
The study examines how including additional call option prices affects model-independent price bounds for exotic derivatives.
Paper introduces a new test for conditional independence using weighted partial copulas.
We obtain a tight distribution-specific characterization of the sample complexity of large-margin classification with L2 regularization: We introduce the margin-adapted dimension, which is a simple function of the second order statistics of the data distribution, and show distribution-specific upper and lower bounds on…
Reliable measures of statistical dependence could be useful tools for learning independent features and performing tasks like source separation using Independent Component Analysis (ICA). Unfortunately, many of such measures, like the mutual information, are hard to estimate and optimize directly. We propose to learn i…
Ultrahigh-dimensional variable selection plays an increasingly important role in contemporary scientific discoveries and statistical research. Among others, Fan and Lv [J. R. Stat. Soc. Ser. B Stat. Methodol. 70 (2008) 849-911] propose an independent screening framework by ranking the marginal correlations. They showed…
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
A new computationally efficient dependence measure, and an adaptive statistical test of independence, are proposed. The dependence measure is the difference between analytic embeddings of the joint distribution and the product of the marginals, evaluated at a finite set of locations (features). These features are chose…
Optimal transport is #P-hard when components are independent, even with approximate solutions.
New test for conditional independence using GNNs avoids estimating conditional distributions.
We propose a test of independence of two multivariate random vectors, given a sample from the underlying population. Our approach, which we call MINT, is based on the estimation of mutual information, whose decomposition into joint and marginal entropies facilitates the use of recently-developed efficient entropy estim…
New method tests causal association using noise contrastive backdoor adjustment.
We consider estimating the marginal likelihood in settings with independent and identically distributed (i.i.d.) data. We propose estimating the predictive distributions in a sequential factorization of the marginal likelihood in such settings by using stochastic gradient Markov Chain Monte Carlo techniques. This appro…
The paper derives a formula for factorizing categorical data to improve Bayes classifiers.
Study of supervised learning from multiple non-independent sequences.
New method tightens bounds on causation probabilities using independent datasets.
We introduce a simple framework for designing private boosting algorithms. We give natural conditions under which these algorithms are differentially private, efficient, and noise-tolerant PAC learners. To demonstrate our framework, we use it to construct noise-tolerant and private PAC learners for large-margin halfspa…
We present a new family of models that is based on graphs that may have undirected, directed and bidirected edges. We name these new models marginal AMP (MAMP) chain graphs because each of them is Markov equivalent to some AMP chain graph under marginalization of some of its nodes. However, MAMP chain graphs do not onl…
Assessing the quality of discovered results is an important open problem in data mining. Such assessment is particularly vital when mining itemsets, since commonly many of the discovered patterns can be easily explained by background knowledge. The simplest approach to screen uninteresting patterns is to compare the ob…
We present a new replay-based method of continual classification learning that we term "conditional replay" which generates samples and labels together by sampling from a distribution conditioned on the class. We compare conditional replay to another replay-based continual learning paradigm (which we term "marginal rep…
Algorithm learns dynamics from past observations.
New method controls error in low-dimensional marginals of spatial models.
Develops methods for constructing likelihoods and priors for Bayesian networks.
Develops bounds predicting deep learning generalization using optimal transport.
We discuss utility based pricing and hedging of jump diffusion processes with emphasis on the practical applicability of the framework. We point out two difficulties that seem to limit this applicability, namely drift dependence and essential risk aversion independence. We suggest to solve these by a re-interpretation …
As shown in recent research, deep neural networks can perfectly fit randomly labeled data, but with very poor accuracy on held out data. This phenomenon indicates that loss functions such as cross-entropy are not a reliable indicator of generalization. This leads to the crucial question of how generalization gap should…
A new method for binary ICA using non-stationary sources.
Any regular Gaussian probability distribution that can be represented by an AMP chain graph (CG) can be expressed as a system of linear equations with correlated errors whose structure depends on the CG. However, the CG represents the errors implicitly, as no nodes in the CG correspond to the errors. We propose in this…
Study bounds financial path expectations using martingale distributions.
The Freund family of distributions becomes a Riemannian 4-manifold with Fisher information as metric; we derive the induced -geometry, i.e., the -curvature, -Ricci curvature with its eigenvales and eigenvectors, the -scalar curvature etc. We show that the Freund manifold has a positive constant 0-scalar cur…
New method bounds causal effects using local consistency of marginals.
This work develops a non-parametric test for relational independence in non-i.i.d. data.
The paper explores intersectional fairness in machine learning, proving bounds on it.
A new algorithm COVA-FC improves subgroup-fair clustering efficiency.
New method uncovers zero entropy in dependent observations after finite samples.
The Collective Graphical Model (CGM) models a population of independent and identically distributed individuals when only collective statistics (i.e., counts of individuals) are observed. Exact inference in CGMs is intractable, and previous work has explored Markov Chain Monte Carlo (MCMC) and MAP approximations for le…
MDMA provides closed-form marginals and conditionals for deep networks.
We discuss two parameterizations of models for marginal independencies for discrete distributions which are representable by bi-directed graph models, under the global Markov property. Such models are useful data analytic tools especially if used in combination with other graphical models. The first parameterization, i…
GTMs model complex multivariate data with varying conditional independencies.