New pruning rules reduce search space for Bayesian network learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods for scalable causal discovery from complex data.
New metrics improve learning of Gaussian networks.
New approach for learning large Bayesian networks using feature clustering and compression.
New algorithm learns Markov network structures efficiently.
New BIC criterion for clustering unsupervised data.
New bounds improve pruning for BDeu score in BNSL.
A new algorithm learns MAGs from data more efficiently using entropy.
In recent years there has been a flurry of works on learning Bayesian networks from data. One of the hard problems in this area is how to effectively learn the structure of a belief network from incomplete data- that is, in the presence of missing values or hidden variables. In a recent paper, I introduced an algorithm…
The efficacy of family-based approaches to mixture model-based clustering and classification depends on the selection of parsimonious models. Current wisdom suggests the Bayesian information criterion (BIC) for mixture model selection. However, the BIC has well-known limitations, including a tendency to overestimate th…
Combines linear acyclic model with logistic regression for causal structure estimation from mixed data.
Bayesian BIC for multi-trial data improves VAR model order selection.
Reducing volatility proxy improves apparent market correlation dynamics.
A new criterion HBIC improves model selection for factor analysis with missing data.
In a Gaussian graphical model, the conditional independence between two variables are characterized by the corresponding zero entries in the inverse covariance matrix. Maximum likelihood method using the smoothly clipped absolute deviation (SCAD) penalty (Fan and Li, 2001) and the adaptive LASSO penalty (Zou, 2006) hav…
We present and implement two algorithms for analytic asymptotic evaluation of the marginal likelihood of data given a Bayesian network with hidden nodes. As shown by previous work, this evaluation is particularly hard for latent Bayesian network models, namely networks that include hidden variables, where asymptotic ap…
Stochastic blockmodels and variants thereof are among the most widely used approaches to community detection for social networks and relational data. A stochastic blockmodel partitions the nodes of a network into disjoint sets, called communities. The approach is inherently related to clustering with mixture models; an…
Detecting and recovering labels in binomial logistic mixtures is challenging due to an information gap.
Three RFF-based methods for nonlinear causal discovery in mixed data.
Study compares variable selection methods for model evaluation and search.
Bayesian model averaging, model selection and its approximations such as BIC are generally statistically consistent, but sometimes achieve slower rates og convergence than other methods such as AIC and leave-one-out cross-validation. On the other hand, these other methods can br inconsistent. We identify the "catch-up …
We extend the Bayesian Information Criterion (BIC), an asymptotic approximation for the marginal likelihood, to Bayesian networks with hidden variables. This approximation can be used to select models given large samples of data. The standard BIC as well as our extension punishes the complexity of a model according to …
MIC improves VAR order selection accuracy.
Proposes a Bayesian model for variable clustering with Gaussian graphical models to handle noise.
New method corrects Laplace/BIC errors in singular models, revealing effective dimension.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
Mixture model-based clustering has become an increasingly popular data analysis technique since its introduction over fifty years ago, and is now commonly utilized within a family setting. Families of mixture models arise when the component parameters, usually the component covariance (or scale) matrices, are decompose…
We discuss Bayesian methods for learning Bayesian networks when data sets are incomplete. In particular, we examine asymptotic approximations for the marginal likelihood of incomplete data given a Bayesian network. We consider the Laplace approximation and the less accurate but more efficient BIC/MDL approximation. We …
Improved survival analysis using square root Cox's models and neural networks.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
A statistical model or a learning machine is called regular if the map taking a parameter to a probability distribution is one-to-one and if its Fisher information matrix is always positive definite. If otherwise, it is called singular. In regular statistical models, the Bayes free energy, which is defined by the minus…
Regularized MLE improves MoE models for high-dimensional data.
Model selection is indispensable to high-dimensional sparse modeling in selecting the best set of covariates among a sequence of candidate models. Most existing work assumes implicitly that the model is correctly specified or of fixed dimensions. Yet model misspecification and high dimensionality are common in real app…
New method accurately evaluates asset pricing under uncertainty and ambiguity.
A new clustering method for vector time series using autoregressive dynamics.
A contaminated mixture model detects outliers in multivariate functional data.
We propose a parsimonious topic model for text corpora. In related models such as Latent Dirichlet Allocation (LDA), all words are modeled topic-specifically, even though many words occur with similar frequencies across different topics. Our modeling determines salient words for each topic, which have topic-specific pr…
Bayesian network framework assesses urban risks across multiple domains.
A challenging problem in estimating high-dimensional graphical models is to choose the regularization parameter in a data-dependent way. The standard techniques include -fold cross-validation (-CV), Akaike information criterion (AIC), and Bayesian information criterion (BIC). Though these methods work well for lo…
A new framework detects changepoints in complex data.
Bayesian method for knot inference in multivariate spline regression.
This paper presents the R package gRapHD for efficient selection of high-dimensional undirected graphical models. The package provides tools for selecting trees, forests and decomposable models minimizing information criteria such as AIC or BIC, and for displaying the independence graphs of the models. It has also some…
We develop a penalized likelihood estimation framework to estimate the structure of Gaussian Bayesian networks from observational data. In contrast to recent methods which accelerate the learning problem by restricting the search space, our main contribution is a fast algorithm for score-based structure learning which …
An AI approach selects variables in linear models.
We have recently proposed a new information-based approach to model selection, the Frequentist Information Criterion (FIC), that reconciles information-based and frequentist inference. The purpose of this current paper is to provide a simple example of the application of this criterion and a demonstration of the natura…
Defense against spam filter attacks using mixture models.
New method speeds up model selection for complex scientific tasks.
We study the cohomology properties of the singular foliation $\F$ determined by an action where the abelian Lie group preserves a riemannian metric on the compact manifold . More precisely, we prove that the basic intersection cohomology $\lau{\IH}{*}{\per{p}}{\mf}$ is finite dimensiona…