Derives derivatives and geometric framework for functions with non-independent variables.
problem Characterizing functions with non-independent variables in probabilistic models.
method Derives actual and dependent partial derivatives, dependent Jacobian matrix, and tensor metric.
result Derives gradient, Hessian, and Taylor expansion for functions with non-independent variables.
Proposes a new dependency function for measuring non-linear relationships.
problem Need for a general-purpose measure of dependency between random variables.
method Revision of ideal properties and proposal of a new dependency function.
result Proposes a new dependency function that meets all desired properties.
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are depende…
Study on NNs for forecasting time series with novel control variable combinations.
problem Forecast future time series with novel combinations of control variables.
method Modular NN architecture with inductive bias for independence of control variables.
result Modular NN architecture improves forecasting of dependent variables up to large horizons.
Bayesian test assesses dependence between mixed data types.
problem Assessing dependence between text, image, and sound data.
method Bayesian kernelised correlation test using Dirichlet process model.
result Demonstrated effectiveness compared to other methods.
The causal discovery of Bayesian networks is an active and important research area, and it is based upon searching the space of causal models for those which can best explain a pattern of probabilistic dependencies shown in the data. However, some of those dependencies are generated by causal structures involving varia…
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
We proposed a new statistical dependency measure called Copula Dependency Coefficient(CDC) for two sets of variables based on copula. It is robust to outliers, easy to implement, powerful and appropriate to high-dimensional variables. These properties are important in many applications. Experimental results show that C…
Bayesian networks, and especially their structures, are powerful tools for representing conditional independencies and dependencies between random variables. In applications where related variables form a priori known groups, chosen to represent different "views" to or aspects of the same entities, one may be more inte…
New model tackles causal bandits with dependent variables.
problem Understanding reward-maximizing interventions in causal networks with dependent variables.
method Introduces hierarchical causal bandit model with a contextual variable capturing interactions among variables.
result Derives nearly matching regret bounds for binary context in causal bandits with dependent arms.
Proposes a new VAE model to avoid posterior collapse by modeling latent variable dependencies.
problem Posterior collapse in variational autoencoders due to assumption of factorized variational posterior.
method Introduces Gaussian Copula Variational Autoencoder (GCVAE) to model latent variable dependencies explicitly.
result Empirical results show GCVAE can avoid posterior collapse while maintaining competitive performance.
Solar improves variable selection in high-dimensional data with complicated dependence structures.
problem Variable selection in ultrahigh dimensional data with severe multicollinearity and grouping effect issues.
method Subsample-ordered least angle regression (Solar) for ultrahigh dimensional data.
result Solar yields substantial improvements in sparsity, stability, and accuracy of variable selection compared to traditional methods.
SHAFF estimates Shapley effects efficiently even with dependent variables.
problem Challenges in estimating Shapley effects, especially with dependent variables.
method SHAFF uses random forests to estimate Shapley effects efficiently.
result SHAFF provides a fast and accurate estimate of Shapley effects.
Measuring dependence between two random variables is very important, and critical in many applied areas such as variable selection, brain network analysis. However, we do not know what kind of functional relationship is between two covariates, which requires the dependence measure to be equitable. That is, it gives sim…
A new method selects important variables for clustering from dependency networks.
problem Variable selection for clustering in high-cost data scenarios.
method Create dependency networks, rank variables by centrality, select top-n variables.
result Top-n variables improve clustering performance compared to existing methods.
Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…
Novel framework controls FDR in high-dimensional, dependent data.
problem FDR control failure in high-dimensional, dependent data.
method Dependency-aware T-Rex selector integrating hierarchical graphical models and martingale theory.
result First to control FDR in high-dimensional, dependent data.
Motivation: Algorithms that discover variables which are causally related to a target may inform the design of experiments. With observational gene expression data, many methods discover causal variables by measuring each variable's degree of statistical dependence with the target using dependence measures (DMs). Howev…
Paper models graph edge dependencies using latent variables for community detection.
problem Graphs' edge dependencies not fully explained by community membership.
method Introduces auxiliary latent variables to model edge dependencies and analyzes conditions for exact recovery.
result Exact recovery possible by semidefinite programming down to maximum likelihood threshold.
Paper improves signal proportion estimation by accounting for variable dependence.
problem Traditional estimators assume independence, limiting applicability in real-world scenarios.
method Integrates arbitrary covariance dependence information using principal factor approximation.
result Method outperforms state-of-the-art estimators in accuracy and detection of weaker signals.
The edge structure of the graph defining an undirected graphical model describes precisely the structure of dependence between the variables in the graph. In many applications, the dependence structure is unknown and it is desirable to learn it from data, often because it is a preliminary step to be able to ascertain c…
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
problem Quantifying the joint dependence between continuous random variables.
method Introduces a novel, fully non-parametric measure bounded [0,1] that assesses predictive accuracy loss.
result The measure captures a wide range of relationships, including non-functional ones, and is interpretable.
We propose a methodology to explore and measure the pairwise correlations that exist between variables in a dataset. The methodology leverages copulas for encoding dependence between two variables, state-of-the-art optimal transport for providing a relevant geometry to the copulas, and clustering for summarizing the ma…
We propose a probabilistic graphical model realizing a minimal encoding of real variables dependencies based on possibly incomplete observation and an empirical cumulative distribution function per variable. The target application is a large scale partially observed system, like e.g. a traffic network, where a small pr…
New criteria distinguish cause from effect in data, overcoming statistical limitations.
problem Determining causal direction from statistical dependence alone.
method Intuitive criteria based on simplicity of prediction, tested on synthetic data.
result Criteria accurately distinguish cause from effect in various scenarios.
We propose an approach to the aggregation of risks which is based on estimation of simple quantities (such as covariances) associated to a vector of dependent random variables, and which avoids the use of parametric families of copulae. Our main result demonstrates that the method leads to bounds on the worst case Valu…
We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more dependent on a first target variable or a second. Dependence is measured via the Hilbe…
The paper proposes methods to extract and analyze individual variable information from complex dependencies.
problem Analyzing and understanding complex dependencies between multiple variables.
method Reversible normalization and iterative dependency reduction to extract individual information, and use it for direct mutual information and multi-feature Granger causality analysis.
result Decoupling of variables to analyze their individual information and direct mutual information transfers.
Efficient methods for linear/logistic regression with network-dependent responses.
problem Regression with dependent responses in networked data.
method Projected gradient descent on negative log-likelihood, proving strong convexity and consistency.
result Strong consistency results for vector of coefficients and dependency strength.
Structured Nonparametric Variational Inference for Dependent Latent Modeling
problem Approximating posterior distributions with complex dependencies among latent variables
method Structured Nonparametric Variational Inference (SN-VI)
result Flexible and accurate posterior approximation with arbitrary shapes
The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.
problem Uncertainty in predicting macroeconomic variables like prices and returns.
method Defines theoretical lower bounds of uncertainty and upper limits on forecast accuracy based on statistical moments and trade volumes.
result Accuracy of forecasts of probabilities of macroeconomic variables doesn't exceed Gaussian approximations.
Investigates VaR behavior for sums of one-sided random variables, showing impossibilities and conditions for super-additivity.
problem Investigates the behavior of Value-at-Risk (VaR) for sums of one-sided random variables.
method Analyzes the extremal aggregation behavior of VaR, introduces structural conditions for super-additivity.
result Characterizes when VaR is fully super-additive and provides unified framework for various dependence structures.
New methods needed for accurate feature importance due to feature dependencies.
problem Misleading variable importance measures from PaP methods due to feature dependencies.
method Alternative approaches involving additional modeling to avoid extrapolation.
result PaP metrics can over-emphasize correlated features, requiring more direct methods.
A new method treats all variables equally in fitting data.
problem Fitting relationships to data with multiple variables, especially when dependent and independent variables are not clearly defined.
method A general method treating all variables impartially, using geometric mean functional relationships and correlation.
result The method provides coefficients that are easily calculated from covariances or correlations, making it scale-invariant and applicable to various units.
CMRFs extend PGMs for topological data, capturing both conditional and marginal dependencies.
problem Limited expressiveness of PGMs for topological data.
method Introducing Colored Markov Random Fields (CMRFs) that model Gaussian edge variables on topological spaces.
result CMRFs improve distributed estimation over physical networks compared to baselines.
DynForest R package predicts outcomes with time-dependent predictors.
problem Handling time-dependent predictors in random forest models.
method Random forests with time-dependent predictors summarized using flexible linear mixed models.
result DynForest can predict continuous, categorical, and survival outcomes.
CDPs visualize causal dependencies in AI models.
problem Understanding how AI models depend on data inputs causally.
method Developed Causal Dependence Plots (CDPs) to visualize causal dependencies.
result CDPs show causal changes in predictors and outcomes.
The paper shows how neural networks with less decision boundary variability generalize better.
problem Improving neural network generalizability by reducing decision boundary variability.
method Introduces new measures (algorithm DB variability and (ε,η)-data DB variability) to quantify decision boundary variability and proves theoretical bounds on generalizability. result Neural networks with lower decision boundary variability have better generalizability, as shown by extensive experiments and theoretical bounds.
Paper proposes a new method to identify causal graphs with latent variables using higher-order cumulants.
problem Estimating causal directed acyclic graphs with latent confounders.
method Uses higher-order cumulants to identify causal structures among observed and latent variables.
result Validates the proposed algorithm through simulations and real-world data.
TCMI assesses mutual dependence of continuous variables without parametric assumptions.
problem Estimating mutual information from continuous distributions.
method TCMI extends mutual information to continuous variables using cumulative distributions.
result TCMI facilitates feature selection and ranking of variable sets.
Global sensitivity analysis with variance-based measures suffers from several theoretical and practical limitations, since they focus only on the variance of the output and handle multivariate variables in a limited way. In this paper, we introduce a new class of sensitivity indices based on dependence measures which o…
The study examines how market trade randomness influences price and return volatility.
problem The accuracy of predicting market-based volatilities and macroeconomic variables is limited.
method Analyzes time series of trade values and volumes, and develops econometric methodologies for predicting volatilities.
result Current macroeconomic models underestimate the accuracy of predicting market-based volatilities and macroeconomic variables.
Paper introduces new bounds linking data compressibility to generalization error.
problem Establishing data-dependent generalization bounds.
method Variable-size compressibility framework linking generalization error to compression rate of input data.
result New bounds depend on empirical data measure, subsuming existing PAC-Bayes and intrinsic dimension bounds.
New method improves combinatorial optimization by capturing dependencies among solution variables.
problem Performance limitations in solving combinatorial optimization problems using independent solution variables.
method Subgraph tokenization and variational annealing to capture dependencies and improve learning efficiency.
result Empirical evidence shows superior performance of autoregressive methods with tokenization and annealed entropy regularization.
For the class of systems of PDEs, for which infinitesimal translations (with respect to some (in)dependent variables) possess specific finite-dimensional invariant subspaces of the space of generalized symmetries of the system considered. We establish when there exist generalized symmetries from these subspaces, which …
We investigate the relative information content of six measures of dependence between two random variables X and Y for large or extreme events for several models of interest for financial time series. The six measures of dependence are respectively the linear correlation ρv+ and Spearman's rho ρs(v) conditio…
Unexpectedly, weighted Pareto variables are stochastically dominant.
problem Understanding stochastic dominance in Pareto distributions.
method Analyzing weighted averages of Pareto random variables with infinite mean.
result The weighted average of Pareto variables is stochastically dominant.
Probabilistic linear discriminant analysis (PLDA) is a method used for biometric problems like speaker or face recognition that models the variability of the samples using two latent variables, one that depends on the class of the sample and another one that is assumed independent across samples and models the within-c…