Paper discovers hidden common variables in nonlinear data.
problem Discover hidden common variables in nonlinear high-dimensional observations.
method Local CCA metric integrated with manifold learning.
result Metric discovers hidden common variables without rigid model assumptions.
We extend common entropy concept and propose algorithms to distinguish causation from correlation.
problem Discovering the simplest latent variable for conditional independence of observed variables.
method Renyi common entropy, iterative algorithm, constraint-based methods modification.
result Improved constraint-based methods for causal inference in small samples.
New method measures common information in high-dimensional data.
problem Measuring common information among many variables.
method Information sieve decomposition to formulate a scalable common information problem.
result Scalable approach demonstrates common information's usefulness in high-dimensional learning.
New method identifies common cause in causal insufficiency, revealing complex phase transitions.
problem Identifying common cause in causal insufficiency with observed joint probability.
method Generalized maximum likelihood method, closely related to maximum entropy principle.
result Identifies consistent common cause that aligns with the common cause principle.
We consider the statistical problem of learning common source of variability in data which are synchronously captured by multiple sensors, and demonstrate that Siamese neural networks can be naturally applied to this problem. This approach is useful in particular in exploratory, data-driven applications, where neither …
D-GCCA improves multi-view data analysis by separating common and distinctive components.
problem Analyzing multi-view high-dimensional data with latent factors.
method Decomposes each view's data matrix into common and distinctive sources with orthogonality constraints.
result Consistent estimators with good performance and efficient computation.
Different directed acyclic graphs (DAGs) may be Markov equivalent in the sense that they entail the same conditional independence relations among the observed variables. Meek (1995) characterizes Markov equivalence classes for DAGs (with no latent variables) by presenting a set of orientation rules that can correctly i…
Algorithm estimates common mean from Gaussian variables with unknown variances.
problem Estimating common mean from Gaussian variables with different unknown variances.
method Intuitive and efficient algorithm using Subset-of-Signals model as benchmark.
result Improved estimation error by polynomial factors compared to previous work.
Novel method identifies structural differences between networks using structural equation models.
problem Identifying structural differences between networks characterized by structural equation models.
method Reparameterization and algorithm design with calibration and construction stages to identify differential structures.
result Our method outperformed independently constructed networks on synthetic data and demonstrated applicability on a real data set.
Deep RL agents vary significantly in Atari environments.
problem Challenges in reproducibility due to stochasticity.
method Experiments with OpenAI Baselines agents.
result Variability in agent performance is significant and underreported.
CLOUD method detects causal relationships in various data types without latent variable assumptions.
problem Detecting causal relationships in the presence of unobserved common causes.
method CLOUD method using Normalized Maximum Likelihood (NML) Code for various data types (discrete, mixed, continuous).
result CLOUD method is more effective than existing methods in inferring causal relationships.
MCCA extracts shared structure from multiple tensor datasets.
problem Extracting shared structure from multiple tensor datasets.
method Multilinear common component analysis (MCCA) using Kronecker products of mode-wise covariance matrices.
result MCCA constructs a common basis that retains information from multiple tensor datasets.
In this study, we have investigated factors of determination which can affect the connected structure of a stock network. The representative index for topological properties of a stock network is the number of links with other stocks. We used the multi-factor model, extensively acknowledged in financial literature. In …
Study evaluates progress in common-sense reasoning tasks.
problem Assessing genuine progress in common-sense reasoning systems.
method Case studies of WSC and SWAG, protocol design to clarify results.
result Previous experimental designs had flaws, need for new protocols.
Paper presents a method to estimate mixed-variable distributions.
problem Estimating joint, conditional, and marginal distributions from mixed data.
method Graph representation of data, eigenvector equations for distribution estimation.
result Method successfully estimates distributions for various machine learning tasks.
A game theory study examines gradual concessions in variable contribution games under uncertainty.
problem Gradualism in contribution games due to free rider effect.
method Stochastic game analysis of variable contribution games, extending Nerlove-Arrow model.
result Equilibrium characterized by regular control strategies leading to gradual concession.
The paper develops a method to model high-dimensional data with many variables and weak signals.
problem Modeling high-dimensional dependent data with many explanatory variables and low signal-to-noise ratio.
method Penalized regression for high-dimensional data, factor modeling of residuals, high-dimensional white noise testing, projected Principal Component Analysis.
result Established asymptotic properties of the proposed method for high-dimensional data.
In recent years, there is a growing interest in learning Bayesian networks with continuous variables. Learning the structure of such networks is a computationally expensive procedure, which limits most applications to parameter learning. This problem is even more acute when learning networks with hidden variables. We p…
Binary sequence correlation estimation fails but trinary data succeeds.
problem Estimating correlation in binary sequences generated by thresholding a hidden continuous sequence.
method Formal analysis and numerical experiments on likelihood maximization and discretization effects.
result Consistent estimation of correlation is possible with trinary data but not with binary data.
A new graphical method compares stochastic variables visually.
problem Comparing non-deterministic measurements visually.
method Cumulative distribution function dominance measure and quantile decomposition.
result Additional conclusions missed by other methods can be inferred.
Bayesian model fuses diverse microbiome data types.
problem Challenges in fusing different types of microbiome data.
method Flexible multinomial-Gaussian generative model with variational EM algorithm.
result Inferred latent variables provide common dimensionality reduction and predictive posterior distribution.
This work restricts hidden cardinality in causal models to infer causal relations.
problem Causal relations between variables with a common unobserved cause cannot be directly inferred.
method Derive inequality constraints from d-separation in causal models with known cardinalities of unobserved variables.
result Inference of causal relations is possible with additional assumptions about cardinalities.
Study proposes a method for identifying important variables in multi-class classification problems.
problem Lack of studies on variable selection in nonparametric classification models, especially for multi-class problems.
method Sparse non-parametric density estimation approach for identifying high impacts variables.
result Proposed method identifies important variables for each class in multi-class classification problems.
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a d-dimensional Gaussian latent variable, is extended by incorporating com…
Unified causal models are formed from fragmented data sets.
problem Combining fragmented data sets to form a unified causal explanation is challenging.
method Using conditional independence properties of marginal datasets to reduce the number of possible models.
result Reduces the number of possible models to a unique one in some cases.
New method detects latent common causes from observational data.
problem Detecting latent common causes in observational data.
method Modified causal discovery algorithms to detect latent common causes.
result Successfully detects latent common causes in various noise regimes and real data.
Bayesian method models binary response and covariates for two groups, estimating causal relationships.
problem Estimating causal relationships between binary response and covariates in observational data.
method Gaussian DAG-probit model with MCMC sampling for posterior distribution estimation.
result Validated method on simulated and real datasets, showing value of grouping variable in causality.
Stochastic neural networks with infinite width become deterministic, reducing training variance.
problem Understanding how stochasticity in neural networks affects learning and regularization.
method Theoretical analysis of stochastic neural networks with infinite width.
result As the width of an optimized stochastic neural network increases, its predictive variance on the training set decreases to zero.
Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical domains. Previous work shows that, given a partition of observed variables such …
Review of variable selection methods for model-based clustering.
problem Dealing with high-dimensional data in model-based clustering.
method Variable selection techniques to facilitate interpretation.
result Summary and illustration of existing methods.
New MCMC methods use auxiliary variables to sample from intractable distributions.
problem Sampling from distributions with unknown normalizing constants.
method Unified Markov chain Monte Carlo framework with auxiliary variables.
result New algorithms outperform existing methods on synthetic and real datasets.
The paper derives theoretical foundations for two common machine learning variable importance measures.
problem Understanding variable importance in machine learning problems.
method The paper derives closed-form expressions for Permute-and-Predict (PaP) and Leave-One-Covariate-Out (LOCO) methods.
result Theoretical derivations explain the behavior of PaP and LOCO under collinearity, linking them to coefficients and predictor variability.
Paper extends stochastic dominance for compound binomial distributions.
problem Stochastic dominance for infinite-mean random variables.
method Investigates properties and inclusion relationships of distribution classes, extends results to compound binomial distributions.
result Establishes necessary and sufficient conditions for first-order stochastic dominance preservation.
In a variety of disciplines such as social sciences, psychology, medicine and economics, the recorded data are considered to be noisy measurements of latent variables connected by some causal structure. This corresponds to a family of graphical models known as the structural equation model with latent variables. While …
In a variety of disciplines such as social sciences, psychology, medicine and economics, the recorded data are considered to be noisy measurements of latent variables connected by some causal structure. This corresponds to a family of graphical models known as the structural equation model with latent variables. While …
Whitening, or sphering, is a common preprocessing step in statistical analysis to transform random variables to orthogonality. However, due to rotational freedom there are infinitely many possible whitening procedures. Consequently, there is a diverse range of sphering methods in use, for example based on principal com…
Develops SCMs for latent selection to simplify causal analysis.
problem Latent selection complicates causal analysis.
method Introduces a conditioning operation for SCMs to encode latent selection.
result Conditioning operation preserves simplicity, acyclicity, and linearity of SCMs.
Paper proposes a method to identify key variables in thick data.
problem Selecting significant variables for response prediction.
method Permutation tests on candidate variables for feature selection.
result The approach outperforms Lasso in feature selection.
The paper compares traditional regression with modern neural network methods for financial hedging and risk compression.
problem Finding optimal hedge ratios and managing portfolio risk using traditional regression methods has limitations.
method The paper introduces regularization techniques and common factor analyses using neural networks to improve upon regression methods.
result Neural network methods provide better performance in hedge ratio estimation and risk compression compared to traditional regression.
New method reparameterizes discrete variables to reduce gradient variance.
problem Low variance gradient estimation for discrete variables in neural networks.
method Marginalizing out the variable of interest to bypass discontinuity, resulting in a new reparameterization trick.
result The new reparameterization reduces gradient variance significantly, theoretically not larger than likelihood-ratio method.
Paper shows how SFA fits into FBM framework for time series separation.
problem Identifying time series decomposition in flow-based models.
method Combining SFA and FBM to make time series decomposition identifiable.
result Time series decomposition becomes identifiable using SFA and FBM.
Learning a causal effect from observational data is not straightforward, as this is not possible without further assumptions. If hidden common causes between treatment X and outcome Y cannot be blocked by other measurements, one possibility is to use an instrumental variable. In principle, it is possible under some…
Foster and Hart proposed an operational measure of riskiness for discrete random variables. We show that their defining equation has no solution for many common continuous distributions including many uniform distributions, e.g. We show how to extend consistently the definition of riskiness to continuous random variabl…
Discond-VAE separates continuous and discrete factors in data.
problem Separating shared and class-specific variations in real-world data.
method Introduces private and public latent variables to represent continuous and discrete factors, respectively.
result Discond-VAE successfully disentangles class-dependent continuous factors from discrete factors.
Paper relaxes identifiability conditions for causal models with latent variables.
problem Challenges in identifying causal graphical models with latent variables.
method Proposes a double triangular graphical condition for nonparametric measurement models with binary latent variables.
result Guarantees identifiability of the entire causal graphical model under relaxed conditions.
ARSM estimator improves gradient backpropagation for categorical variables.
problem Improving gradient backpropagation through categorical variables.
method ARSM combines variable augmentation, REINFORCE, Rao-Blackwellization, and variable swapping.
result ARSM outperforms existing estimators and provides variance reduction methods.
Method classifies causal relationship between two discrete variables.
problem Estimating causal direction and confounding in discrete variables.
method Bayesian classifier assuming independence of cause and mechanism.
result Method acknowledges inherent baseline error in classification.
MultiDendrograms is a Java-written application that computes agglomerative hierarchical clusterings of data. Starting from a distances (or weights) matrix, MultiDendrograms is able to calculate its dendrograms using the most common agglomerative hierarchical clustering methods. The application implements a variable-gro…