Derives derivatives and geometric framework for functions with non-independent variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a new dependency function for measuring non-linear relationships.
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are depende…
Study on NNs for forecasting time series with novel control variable combinations.
Bayesian test assesses dependence between mixed data types.
The causal discovery of Bayesian networks is an active and important research area, and it is based upon searching the space of causal models for those which can best explain a pattern of probabilistic dependencies shown in the data. However, some of those dependencies are generated by causal structures involving varia…
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
We proposed a new statistical dependency measure called Copula Dependency Coefficient(CDC) for two sets of variables based on copula. It is robust to outliers, easy to implement, powerful and appropriate to high-dimensional variables. These properties are important in many applications. Experimental results show that C…
Bayesian networks, and especially their structures, are powerful tools for representing conditional independencies and dependencies between random variables. In applications where related variables form a priori known groups, chosen to represent different "views" to or aspects of the same entities, one may be more inte…
New model tackles causal bandits with dependent variables.
Solar improves variable selection in high-dimensional data with complicated dependence structures.
SHAFF estimates Shapley effects efficiently even with dependent variables.
Measuring dependence between two random variables is very important, and critical in many applied areas such as variable selection, brain network analysis. However, we do not know what kind of functional relationship is between two covariates, which requires the dependence measure to be equitable. That is, it gives sim…
A new method selects important variables for clustering from dependency networks.
Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…
Novel framework controls FDR in high-dimensional, dependent data.
Motivation: Algorithms that discover variables which are causally related to a target may inform the design of experiments. With observational gene expression data, many methods discover causal variables by measuring each variable's degree of statistical dependence with the target using dependence measures (DMs). Howev…
Paper models graph edge dependencies using latent variables for community detection.
Paper improves signal proportion estimation by accounting for variable dependence.
The edge structure of the graph defining an undirected graphical model describes precisely the structure of dependence between the variables in the graph. In many applications, the dependence structure is unknown and it is desirable to learn it from data, often because it is a preliminary step to be able to ascertain c…
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
We propose a methodology to explore and measure the pairwise correlations that exist between variables in a dataset. The methodology leverages copulas for encoding dependence between two variables, state-of-the-art optimal transport for providing a relevant geometry to the copulas, and clustering for summarizing the ma…
We propose a probabilistic graphical model realizing a minimal encoding of real variables dependencies based on possibly incomplete observation and an empirical cumulative distribution function per variable. The target application is a large scale partially observed system, like e.g. a traffic network, where a small pr…
New criteria distinguish cause from effect in data, overcoming statistical limitations.
We propose an approach to the aggregation of risks which is based on estimation of simple quantities (such as covariances) associated to a vector of dependent random variables, and which avoids the use of parametric families of copulae. Our main result demonstrates that the method leads to bounds on the worst case Valu…
We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more dependent on a first target variable or a second. Dependence is measured via the Hilbe…
The paper proposes methods to extract and analyze individual variable information from complex dependencies.
Structured Nonparametric Variational Inference for Dependent Latent Modeling
The standard linear and logistic regression models assume that the response variables are independent, but share the same linear relationship to their corresponding vectors of covariates. The assumption that the response variables are independent is, however, too strong. In many applications, these responses are collec…
The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.
Investigates VaR behavior for sums of one-sided random variables, showing impossibilities and conditions for super-additivity.
A new method treats all variables equally in fitting data.
CMRFs extend PGMs for topological data, capturing both conditional and marginal dependencies.
DynForest R package predicts outcomes with time-dependent predictors.
The paper shows how neural networks with less decision boundary variability generalize better.
CDPs visualize causal dependencies in AI models.
Paper proposes a new method to identify causal graphs with latent variables using higher-order cumulants.
TCMI assesses mutual dependence of continuous variables without parametric assumptions.
Global sensitivity analysis with variance-based measures suffers from several theoretical and practical limitations, since they focus only on the variance of the output and handle multivariate variables in a limited way. In this paper, we introduce a new class of sensitivity indices based on dependence measures which o…
The study examines how market trade randomness influences price and return volatility.
New method improves combinatorial optimization by capturing dependencies among solution variables.
Paper introduces new bounds linking data compressibility to generalization error.
For the class of systems of PDEs, for which infinitesimal translations (with respect to some (in)dependent variables) possess specific finite-dimensional invariant subspaces of the space of generalized symmetries of the system considered. We establish when there exist generalized symmetries from these subspaces, which …
Variational language models seek to estimate the posterior of latent variables with an approximated variational posterior. The model often assumes the variational posterior to be factorized even when the true posterior is not. The learned variational posterior under this assumption does not capture the dependency relat…
We investigate the relative information content of six measures of dependence between two random variables and for large or extreme events for several models of interest for financial time series. The six measures of dependence are respectively the linear correlation and Spearman's rho conditio…
Unexpectedly, weighted Pareto variables are stochastically dominant.
Probabilistic linear discriminant analysis (PLDA) is a method used for biometric problems like speaker or face recognition that models the variability of the samples using two latent variables, one that depends on the class of the sample and another one that is assumed independent across samples and models the within-c…
A new model integrates LSTM and copulas for high-dimensional financial data.